Business#Benchmark

Arena raises $200M at a $3.1B valuation, nearly double

Arena, the company behind LMArena, raised a $200M Series B at a $3.1B valuation — nearly double January — and launched an alignment leaderboard for behavior.

The three Arena founders photographed in a bright office

Arena, the company behind the LMArena model leaderboard, raised a $200 million Series B at a $3.1 billion valuation, TechCrunch reported on October 8 — nearly double the $1.7 billion post-money from its January Series A, ten months ago. Lightspeed Venture Partners and Khosla Ventures led; a16z, Felicis, Salesforce Ventures and others joined. Alongside the round, the company launched an “alignment leaderboard” that ranks models on unauthorized actions, false attribution and deceptive completion — models claiming they finished work they did not. What started as a UC Berkeley research project in 2023 is now the first large-scale capital bet on independent AI evaluation.

Key points

  • A steep revenue curve: $30 million annualized at the January round, $100 million by June per the company; the commercial product, AI Evaluations, launched in September 2025 and sells community-vote-based performance analytics to labs and enterprises.
  • The alignment leaderboard: categories include unauthorized actions, false attribution and deceptive completion; in the preliminary ranking OpenAI models lead, with Claude Opus 5.5 and Claude Fable sixth and ninth.
  • The traffic base: tens of millions of monthly visits, per the company — users submit prompts and vote for free, and those votes feed both the leaderboard and the commercial product.

Background

Independent evaluation works as a business because static benchmarks have lost credibility: labs overfit to them and training data contaminates them. The theme keeps crossing our coverage — the academic backlash to OpenAI’s mass-released math proofs is the extreme case, while Claude Haiku 5.5’s launch leaning on OSWorld and Terminal-Bench scores shows leaderboard numbers remain standard launch ammunition. Arena sits in the crack: it writes no questions of its own, lets tens of millions of humans blind-vote, then sells the analytics back to the industry. At the January round, its annualized revenue was $30 million.

The facts

Round / metricDateFigure
Series AJanuary 2026$150M at $1.7B post-money; $30M ARR
Series BOctober 8, 2026$200M at $3.1B post-money
ARRJune 2026$100M (company figure)
Origin2023UC Berkeley research project
CommercializationSeptember 2025AI Evaluations product launch

The Series B cap table is worth recording: Lightspeed Venture Partners and Khosla Ventures led, with Salesforce Ventures, 01 Advisors, Dell Technologies Capital, Endeavor Catalyst, a16z and Felicis participating. TechCrunch’s industry framing: labs gaming static benchmarks pushed enterprises toward alternative performance references, lifting demand for Arena’s paid product. The alignment leaderboard is the round’s companion launch — turning “can I trust this model’s behavior” into a rankable, citable metric, a new layer on top of the capability boards.

What others say

TechCrunch frames it as evaluation becoming a business: when vendor-published benchmarks stopped being trusted, third-party human-vote data found buyers. Unite.AI emphasizes the multiple — near-doubling in ten months, on a base that already carries nine figures of revenue. IT Home’s Chinese coverage positions LMArena as the reference evaluation platform and notes how closely Chinese-language communities track Chinese models’ arena rankings. The criticism is equally established: crowd blind-voting favors longer, more agreeable answers, so style can drift from real capability — a recurring HN and academic objection, and one Arena itself acknowledges by applying statistical corrections to its vote data.

Our take

The $3.1 billion prices independent evaluation — and pushes it into the mire it measures: the labs being ranked are also prospective customers, and that two-sided structure is the question Arena will have to answer for as long as it exists. The alignment leaderboard is the smarter differentiation: “will the model claim it finished when it did not” is the core fear blocking agent adoption, and turning that into a product line builds more moat than another capability board. For developers the advice is unchanged: arena ranks make a good first filter, but score critical workloads on your own traffic. Honest uncertainty: ARR is a company figure, unaudited; the alignment board’s methodology is new and independently unverified.