Models#Open source#Agents#Benchmark#Flagship#Computer use#Chinese models
China’s 1.6T open-source model Nex-N2.5 tops BrowseComp, beating GLM-5.3 and DeepSeek-V4 on agent work
Shanghai Innovation Institute released the Nex-N2.5 family on September 9, 2026 — 35B, 397B and 1.6T models, all under Apache-2.0. On the official table the 1.6T Max tops BrowseComp at 92.6, ahead of Claude Opus 5 and GPT-5.6 Sol, and outscores GLM-5.3 and DeepSeek-V4-Pro on automation and tool-use benchmarks.

Shanghai Innovation Institute announced its Nex-N2.5 family as an open-source release on September 9, 2026, spanning a 35B Mini, a 397B Pro and a 1.6T Max, all under Apache-2.0. Max is post-trained on DeepSeek-V4-Pro-Base, and the team calls it its first complete post-training effort at trillion-parameter scale. On the comparison table published with the release it tops BrowseComp at 92.6 — the best score in the table, ahead of Kimi-K3 at 91.2, Claude Opus 5 at 90.8 and GPT-5.6 Sol at 90.4.
The facts
- The lineup: Mini builds on Qwen3.5-35B-A3B and Pro on Qwen3.5-397B-A17B; both are multimodal MoE models aimed at computer use, web browsing and vision-driven self-correction. Max is a 1.6T-parameter, text-only MoE with 49B active, built on DeepSeek-V4-Pro-Base and aimed at complex reasoning, coding and agent work. All three have a 256K context window.
- Background: Nex-AGI is an open-source ecosystem for agentic models that Shanghai Innovation Institute built with partner research groups and startups; Nex-N2.5 follows the Nex-N2 generation.
- License and access: weights are on Hugging Face and ModelScope under Apache-2.0. Mini and Pro are also hosted on OpenRouter.
- Benchmarks: the comparison table from the official README and site (updated September 8, 2026; bold marks the best score, — means unavailable). Text first:
| Benchmark | Mini | Pro | Max | Claude Opus 5 | GPT-5.6 Sol | Kimi-K3 | GLM-5.3 | DeepSeek-V4-Pro-0813 | Qwen3.8-Max |
|---|---|---|---|---|---|---|---|---|---|
| Terminal-Bench 2.1 | 73.4 | 82.7 | 86.1 | 89.1 | 88.8 | 88.3 | 88.2 | 87.9 | 86.6 |
| SWE-Bench Pro | 43.8 | 61.2 | 65.7 | 79.2 | 64.6 | 63.3 | 64.6 | 55.4 | 67.7 |
| DeepSWE v1.1 | 36.1 | 55.8 | 65.6 | 73.7 | 72.7 | 67.5 | 66.9 | 62.8 | 69.3 |
| AutomationBench v1.0.6 | 32.3 | 44.2 | 50.2 | 50.3 | 45.8 | 46.7 | 48.2 | 43.2 | 39.8 |
| Toolathlon Verified | 54.6 | 68.5 | 74.7 | 76.5 | 74.9 | 76.5 | 73.0 | 74.1 | 72.5 |
| GDPval-AA v2 | 1446 | 1628 | 1713 | 1831 | 1711 | 1675 | 1763 | 1580 | 1717 |
| Job Bench | 28.5 | 41.4 | 53.6 | 65.7 | 45.4 | 52.9 | 58.2 | 54.1 | 53.4 |
| BrowseComp | 83.4 | 89.7 | 92.6 | 90.8 | 90.4 | 91.2 | — | — | — |
And multimodal:
| Benchmark | Mini | Pro | MiniMax-M3 | Claude Opus 5 | GPT-5.6 Sol | Kimi-K3 | GLM-5.3-Flash | DeepSeek-V4-Flash-Vision | Qwen3.8-Max |
|---|---|---|---|---|---|---|---|---|---|
| OSWorld-Verified | 71.2 | 82.2 | 75.2 | 83.4 | 83.2 | 84.8 | 62.3 | 76.7 | 86.1 |
| OSWorld-2 | 30.5 | 56.4 | 22.3 | 68.3 | 62.7 | 58.3 | — | — | 46.7 |
| WebTest | 48.6 | 52.8 | — | — | 54.0 | — | — | — | 52.3 |
| WebArena-Verified | 63.4 | 67.6 | — | — | 69.7 | 71.6 | — | 62.3 | 66.8 |
| OSWorld-G | 82.9 | 87.4 | — | 76.8 | 77.7 | 79.6 | 83.3 | 59.4 | 84.9 |
| Vision2Web | 52.9 | 68.2 | 59.0 | — | 79.8 | — | — | — | 75.1 |
| SWE-MM | 25.5 | 38.2 | — | 59.4 | 40.2 | 37.3 | 20.6 | 39.2 | 39.2 |
| OmniDoc | 89.7 | 92.2 | 91.6 | — | 92.9 | 91.1 | — | — | 92.1 |
- Head to head: on the text table, Max scores higher than both GLM-5.3 and DeepSeek-V4-Pro-0813 on AutomationBench (50.2) and Toolathlon (74.7). It beats DeepSeek but trades wins with GLM on SWE-Bench Pro (65.7), DeepSWE (65.6) and GDPval (1713), and trails both on Terminal-Bench 2.1 (86.1) and Job Bench (53.6). On the multimodal table, Pro’s 87.4 on OSWorld-G is the top score.
- Serving costs: the official Docker image (
nexagi/sglang:v0.5.18-nex-patch) runs Mini on 2×H100, Pro on 8×H100, and Max on two nodes of 16×H200 each. Recommended sampling is temperature 0.7, top_p 0.95, top_k 40;reasoning_effortswitches between no thinking, adaptive thinking and always thinking. - Scoring caveats: scores that have a public source come from model providers’ own reports (Kimi-K3, Qwen3.8-Max, GLM-5.3, HY4); the rest are Nex-AGI’s own runs. Computer-use benchmarks use its NexCUA harness, which is not open-sourced yet.
- Demos: the team shows Max repairing a broken CRC16-CCITT streaming-accelerator RTL and taking it through synthesis and Sky130HD place-and-route (reported Fmax around 386 MHz, zero setup/hold TNS). Pro earns the first gym badge in Pokémon Platinum in 468-plus actions and turns a concept sketch into a 58-part mecha mesh in Blender.
Our take
On agent-style work this is the first Chinese open-weights release to beat every closed frontier model on a hard benchmark — at least on the official table, where BrowseComp 92.6 and OSWorld-G 87.4 are both the top scores including closed models. Read the numbers with care: most are run by the vendor, some competitor scores come from those competitors’ own reports, and none of it is third-party reproduced yet.
Direction by direction it looks narrower. On browser research, computer use and automation, Nex-N2.5 does beat GLM-5.3 and DeepSeek-V4-Pro-0813. On classic coding it trades wins — SWE-Bench Pro at 65.7 is ahead of both, but Terminal-Bench 2.1 and Job Bench lag behind, and Claude Opus 5’s 79.2 on SWE-Bench Pro is still far off.
The practical gate is hardware: Max needs 32 H200s, out of reach for most teams. What you can actually try is Mini or Pro on OpenRouter, or Mini on two H100s — a cheap way to test whether an open model can replace a closed API for computer-use work.
How to try it
- Call
nex-agi/nex-n2.5-proornex-agi/nex-n2.5-minidirectly on OpenRouter - Download weights from Hugging Face or ModelScope (
nex-agi/Nex-N2.5-Max,nex-agi/Nex-N2.5-Pro,nex-agi/Nex-N2.5-mini, all Apache-2.0) - Self-host with the official Docker image
nexagi/sglang:v0.5.18-nex-patch; launch commands are in the README