Models#Open source#Agents#Benchmark#Flagship#Computer use#Chinese models

China’s 1.6T open-source model Nex-N2.5 tops BrowseComp, beating GLM-5.3 and DeepSeek-V4 on agent work

Shanghai Innovation Institute released the Nex-N2.5 family on September 9, 2026 — 35B, 397B and 1.6T models, all under Apache-2.0. On the official table the 1.6T Max tops BrowseComp at 92.6, ahead of Claude Opus 5 and GPT-5.6 Sol, and outscores GLM-5.3 and DeepSeek-V4-Pro on automation and tool-use benchmarks.

Official benchmark chart: bar comparisons across six text benchmarks, with the three Nex-N2.5 sizes alongside Claude, GPT-5.6 Sol, GLM-5.3, Kimi-K3 and DeepSeek-V4-Pro

Shanghai Innovation Institute announced its Nex-N2.5 family as an open-source release on September 9, 2026, spanning a 35B Mini, a 397B Pro and a 1.6T Max, all under Apache-2.0. Max is post-trained on DeepSeek-V4-Pro-Base, and the team calls it its first complete post-training effort at trillion-parameter scale. On the comparison table published with the release it tops BrowseComp at 92.6 — the best score in the table, ahead of Kimi-K3 at 91.2, Claude Opus 5 at 90.8 and GPT-5.6 Sol at 90.4.

The facts

  • The lineup: Mini builds on Qwen3.5-35B-A3B and Pro on Qwen3.5-397B-A17B; both are multimodal MoE models aimed at computer use, web browsing and vision-driven self-correction. Max is a 1.6T-parameter, text-only MoE with 49B active, built on DeepSeek-V4-Pro-Base and aimed at complex reasoning, coding and agent work. All three have a 256K context window.
  • Background: Nex-AGI is an open-source ecosystem for agentic models that Shanghai Innovation Institute built with partner research groups and startups; Nex-N2.5 follows the Nex-N2 generation.
  • License and access: weights are on Hugging Face and ModelScope under Apache-2.0. Mini and Pro are also hosted on OpenRouter.
  • Benchmarks: the comparison table from the official README and site (updated September 8, 2026; bold marks the best score, — means unavailable). Text first:
Benchmark Mini Pro Max Claude Opus 5 GPT-5.6 Sol Kimi-K3 GLM-5.3 DeepSeek-V4-Pro-0813 Qwen3.8-Max
Terminal-Bench 2.1 73.4 82.7 86.1 89.1 88.8 88.3 88.2 87.9 86.6
SWE-Bench Pro 43.8 61.2 65.7 79.2 64.6 63.3 64.6 55.4 67.7
DeepSWE v1.1 36.1 55.8 65.6 73.7 72.7 67.5 66.9 62.8 69.3
AutomationBench v1.0.6 32.3 44.2 50.2 50.3 45.8 46.7 48.2 43.2 39.8
Toolathlon Verified 54.6 68.5 74.7 76.5 74.9 76.5 73.0 74.1 72.5
GDPval-AA v2 1446 1628 1713 1831 1711 1675 1763 1580 1717
Job Bench 28.5 41.4 53.6 65.7 45.4 52.9 58.2 54.1 53.4
BrowseComp 83.4 89.7 92.6 90.8 90.4 91.2 — — —

And multimodal:

Benchmark Mini Pro MiniMax-M3 Claude Opus 5 GPT-5.6 Sol Kimi-K3 GLM-5.3-Flash DeepSeek-V4-Flash-Vision Qwen3.8-Max
OSWorld-Verified 71.2 82.2 75.2 83.4 83.2 84.8 62.3 76.7 86.1
OSWorld-2 30.5 56.4 22.3 68.3 62.7 58.3 — — 46.7
WebTest 48.6 52.8 — — 54.0 — — — 52.3
WebArena-Verified 63.4 67.6 — — 69.7 71.6 — 62.3 66.8
OSWorld-G 82.9 87.4 — 76.8 77.7 79.6 83.3 59.4 84.9
Vision2Web 52.9 68.2 59.0 — 79.8 — — — 75.1
SWE-MM 25.5 38.2 — 59.4 40.2 37.3 20.6 39.2 39.2
OmniDoc 89.7 92.2 91.6 — 92.9 91.1 — — 92.1
  • Head to head: on the text table, Max scores higher than both GLM-5.3 and DeepSeek-V4-Pro-0813 on AutomationBench (50.2) and Toolathlon (74.7). It beats DeepSeek but trades wins with GLM on SWE-Bench Pro (65.7), DeepSWE (65.6) and GDPval (1713), and trails both on Terminal-Bench 2.1 (86.1) and Job Bench (53.6). On the multimodal table, Pro’s 87.4 on OSWorld-G is the top score.
  • Serving costs: the official Docker image (nexagi/sglang:v0.5.18-nex-patch) runs Mini on 2×H100, Pro on 8×H100, and Max on two nodes of 16×H200 each. Recommended sampling is temperature 0.7, top_p 0.95, top_k 40; reasoning_effort switches between no thinking, adaptive thinking and always thinking.
  • Scoring caveats: scores that have a public source come from model providers’ own reports (Kimi-K3, Qwen3.8-Max, GLM-5.3, HY4); the rest are Nex-AGI’s own runs. Computer-use benchmarks use its NexCUA harness, which is not open-sourced yet.
  • Demos: the team shows Max repairing a broken CRC16-CCITT streaming-accelerator RTL and taking it through synthesis and Sky130HD place-and-route (reported Fmax around 386 MHz, zero setup/hold TNS). Pro earns the first gym badge in Pokémon Platinum in 468-plus actions and turns a concept sketch into a 58-part mecha mesh in Blender.

Our take

On agent-style work this is the first Chinese open-weights release to beat every closed frontier model on a hard benchmark — at least on the official table, where BrowseComp 92.6 and OSWorld-G 87.4 are both the top scores including closed models. Read the numbers with care: most are run by the vendor, some competitor scores come from those competitors’ own reports, and none of it is third-party reproduced yet.

Direction by direction it looks narrower. On browser research, computer use and automation, Nex-N2.5 does beat GLM-5.3 and DeepSeek-V4-Pro-0813. On classic coding it trades wins — SWE-Bench Pro at 65.7 is ahead of both, but Terminal-Bench 2.1 and Job Bench lag behind, and Claude Opus 5’s 79.2 on SWE-Bench Pro is still far off.

The practical gate is hardware: Max needs 32 H200s, out of reach for most teams. What you can actually try is Mini or Pro on OpenRouter, or Mini on two H100s — a cheap way to test whether an open model can replace a closed API for computer-use work.

How to try it

  • Call nex-agi/nex-n2.5-pro or nex-agi/nex-n2.5-mini directly on OpenRouter
  • Download weights from Hugging Face or ModelScope (nex-agi/Nex-N2.5-Max, nex-agi/Nex-N2.5-Pro, nex-agi/Nex-N2.5-mini, all Apache-2.0)
  • Self-host with the official Docker image nexagi/sglang:v0.5.18-nex-patch; launch commands are in the README