Harvey LAB: Legal Agents Get a Real Exam, and Kimi K3 Tops It
The most valuable legal AI company open-sourced its own exam, and a Chinese open-weights model just aced it. Harvey's LAB (Legal Agent Benchmark) hit GitHub trending today: 1,671 tasks across 24 practice areas, from M&A to tax to bankruptcy, each graded against expert-written rubrics β over 75,000 criteria in total. Unlike the contract-review toys that passed for legal benchmarks before, LAB tasks are long-horizon: analyze a messy client matter, synthesize the record, produce an actual deliverable like a risk memo. The repo ships with an execution harness so you can run and score your own agent, MIT licensed.
The scoreboard is where it gets spicy. Artificial Analysis runs an independent implementation called Harvey LAB-AA on 120 held-out tasks with their own Stirrup harness, and the current leaderboard reads: Kimi K3 (max) at 94.6%, Claude Fable 5 at 93.6%, Muse Spark 1.1 at 93.1%. An open-weights model from Beijing sitting on top of a US legal benchmark backed by Nvidia, OpenAI, Anthropic, Mistral and DeepMind β that sentence would have been unthinkable a year ago.
Two things worth taking away. First, the benchmark grew from 1,200 tasks at its May launch to 1,671 now, and the fact that it keeps trending months later says domain-specific agent benchmarks with real expert rubrics are scarce and wanted. Second, when frontier models cluster within one point of each other in the mid-90s, the benchmark is telling you legal agents are closer to deployment-grade than most lawyers think.
Repo at github.com/harveyai/harvey-labs, leaderboard at artificialanalysis.ai/evaluations/harvey-lab-aa.
← Back to all articles
The scoreboard is where it gets spicy. Artificial Analysis runs an independent implementation called Harvey LAB-AA on 120 held-out tasks with their own Stirrup harness, and the current leaderboard reads: Kimi K3 (max) at 94.6%, Claude Fable 5 at 93.6%, Muse Spark 1.1 at 93.1%. An open-weights model from Beijing sitting on top of a US legal benchmark backed by Nvidia, OpenAI, Anthropic, Mistral and DeepMind β that sentence would have been unthinkable a year ago.
Two things worth taking away. First, the benchmark grew from 1,200 tasks at its May launch to 1,671 now, and the fact that it keeps trending months later says domain-specific agent benchmarks with real expert rubrics are scarce and wanted. Second, when frontier models cluster within one point of each other in the mid-90s, the benchmark is telling you legal agents are closer to deployment-grade than most lawyers think.
Repo at github.com/harveyai/harvey-labs, leaderboard at artificialanalysis.ai/evaluations/harvey-lab-aa.
Comments