September 6, 2026BenchmarkResearch

Intelligence Index v4.2 Hides 40% of the Test From the Labs

Artificial Analysis shipped version 4.2 of its Intelligence Index on September 4 (https://artificialanalysis.ai/articles/artificial-analysis-intelligence-index-v4-2), and the most important change is not a score — it is that 40% of the index weighting now sits on private, held-out test sets, double the share of v4.1. The stated reason is blunt: "to reduce the ability for labs to game evaluations." The industry's most-cited independent scoreboard just declared that public benchmarks can no longer be trusted as measurement, only as marketing.

Two evals enter, one retires. AA-Briefcase is an in-house agentic eval of realistic knowledge work in complex, multi-week-shaped projects, graded by rubric and pairwise comparison against a private set. GDP.pdf, built by Surge AI, asks models to reason across a 4,592-page document against 1,275 expert-authored criteria. Out goes GPQA Diamond, formally declared saturated — a benchmark that headlined two years of launch posts now retired as useless for telling frontier models apart.

The scores themselves: Claude Fable 5.1 leads the overall index and AA-Briefcase (alongside Opus 5) — coverage of that release at https://clauday.com/article/9831aac1-c20b-4c47-98d3-9bf82882ba58. GPT-6 Astra gains about 4 index points over its predecessor, roughly 85 Elo, leads the 4,592-page GDP.pdf at 33.2%, and tops the token-efficiency frontier. Meta is the third lab, ahead of SpaceXAI, Moonshot, Z.AI and Google.

Put this next to ARC's admission last week that its leaderboard measures model plus harness (https://clauday.com/article/760df7af-8ef1-4bd1-8f66-f9ab6481c575), and the pattern is one story: the eval layer is rebuilding itself around distrust — of contamination, of gaming, of scaffolding. Private sets are the eval world's version of going closed-source, and nobody seems to mind the irony.
← Previous
GPT-6 Astra's Price Sheet: $50 Output and a Trapdoor at 272K
Next →
Spotify Cut Claude Code Tokens 90% by Giving It a Cheap Intern
← Back to all articles

Comments

Loading...
>_