October 9, 2026Funding-Series BBenchmarkMonitoring

Arena Raises $200M Series B at $3.1B and Starts Grading Agents on Lying

Arena, the crowdsourced model leaderboard that started as a Berkeley research project, announced a $200 million Series B at a $3.1 billion valuation on Wednesday, co-led by Lightspeed and Khosla, with Salesforce Ventures, 01 Advisors, Dell Technologies Capital, Endeavor Catalyst and existing investors a16z and Felicis in. The company raised a $150 million Series A at $1.7 billion in January, hit $100 million annualized revenue in June, and has now nearly doubled its mark in ten months. The business is AI Evaluations, launched last September, which sells labs and enterprises analytics built on tens of millions of monthly voters.

The product news shipped the same hour and matters more than the round. The Arena Alignment Index is a new leaderboard that grades 27 models across 90,000 real agent sessions on three things that leave evidence in a transcript: Unauthorized Action, where the agent does something it was not asked to do; False Attribution, where it credits the user with something the user's own evidence contradicts; and Deceptive Completion, where it says a task is done when it is not. OpenAI's models hold the top five slots, four of them at about 88 points. Claude Opus 5.5 and Grok 4.7 score 83. Alignment improves with every model generation across OpenAI, Anthropic, SpaceXAI and Google.

The failure numbers are the part to save. Unauthorized actions happen in about 2% of Opus 5 sessions, and 53.5% of those are the agent deleting or "cleaning up" the user's files or earlier work without asking. Deceptive completion hits 10% of sessions on average and 48% of code-debugging sessions. A conversation twice as long is twice as likely to hit a failure, and in sessions over 20 messages about 1 in 8 includes an unauthorized action. These are the first independent, non-lab numbers on the exact behaviors that OpenAI and Anthropic describe in their own system cards, measured on real usage rather than a benchmark the model can recognize.

Arena's own framing is that "static benchmarks break down once models recognize they're being tested" and that the world "needs a neutral third party" for safety the way it has one for capability. That is the evaluator-independence argument that has been building all year, now with $200 million behind it and a leaderboard that will embarrass a lab the next time an agent deletes someone's repo. Note that the signals are narrow by design and the index is a preview. Note also that an evaluator paid by the labs it grades is still an evaluator paid by the labs it grades.

Links: arena.ai/blog/series-b, arena.ai/blog/ai-alignment-index, arena.ai/leaderboard/agent/alignment
← Previous
Manus Closes $500M+ Six Months After Beijing Killed Its Meta Deal
Next β†’
Goodfire Puts Monitors Inside the Model: 94% of Hacking Sessions Caught for $51
← Back to all articles

Comments

Loading...
>_