September 20, 2026BenchmarkMonitoring

Somebody Is Finally Getting Paid to Grade the Models

Vals is two years old and its revenue is eight times what it was last year. It went from eight employees in January to twenty-five, with plans for another ten to fifteen. Its business is one sentence: companies pay Vals to evaluate their models, the way students pay to sit a standardized test. The a16z-led forty million dollar Series A landed last month, on top of a seed from 8VC and Bloomberg Beta.

The product decision that matters is that the evaluations are private. Public benchmarks have one structural flaw that no amount of methodology fixes, which is that a lab can train against them, sometimes on purpose and sometimes because the test set leaked into a crawl three years ago. Every score on every public leaderboard carries that asterisk and everyone in the field knows it. Vals holds its tests, so the only way to score well is to be good at the thing.

What they measure is a step up from token-level accuracy too. Co-founder Rayan Krishnan, who is twenty-five and came through Palantir, Microsoft and Stanford's AI lab, frames it as whether the model can do work that produces a product of the same quality as a human, in a given domain. They run law, finance, coding, cybersecurity, biosecurity, mental health, and they will assess compliance with the Geneva Convention, which is either the most necessary eval on the list or the most alarming that it is on the list.

Here is the tension, and it is worth naming rather than cheering past. The labs being graded are the customers paying for the grade. That is precisely the structure of the credit rating agencies before 2008: issuer-pays, with the rated party free to shop for a better letter. It does not mean Vals is compromised, and the model has real advantages over pay-to-play leaderboards and over labs marking their own homework. It does mean that independence is a property of governance, not of intent, and nobody in this space has published what their governance actually is.

That is the thread to watch, because it has been tightening all quarter. AIUC raised for agent insurance. The frontier labs are in coordination talks about safety evaluation. Every enterprise buyer now asks for a third-party number before signing. The business of grading AI is becoming an industry before anyone has decided who grades the graders. Coverage at techcrunch.com.
← Previous
Anthropic Now Owns a Wet Lab
Next β†’
Anthropic Open-Sourced Eleven Job Descriptions
← Back to all articles

Comments

Loading...
>_