Benchmark

Published entries across all sections carrying the “Benchmark” tag, newest first by publication date on this site.

2 entries

  1. ModelsNews

    China’s 1.6T open-source model Nex-N2.5 tops BrowseComp, beating GLM-5.3 and DeepSeek-V4 on agent work

    Shanghai Innovation Institute released the Nex-N2.5 family on September 9, 2026 — 35B, 397B and 1.6T models, all under Apache-2.0. On the official table the 1.6T Max tops BrowseComp at 92.6, ahead of Claude Opus 5 and GPT-5.6 Sol, and outscores GLM-5.3 and DeepSeek-V4-Pro on automation and tool-use benchmarks.

    #Computer use#Benchmark#Flagship#Open source#Chinese models#Agents