October 3, 2026BenchmarkResearchAgents

Argo-Bench: 7.5 Billion Rows, and the Agent Gets Graded on What Happens Next

Text-to-SQL benchmarks have two known problems. They grade the query, not what you do with the answer. And audits have found their answer keys are often wrong. Argo-Bench, from the TextQL team, rebuilds the test from the ground up around an enterprise warehouse that looks like the real thing.

The authors simulated a New York City food delivery platform at true scale: 81 million orders in 2024, with realistic economics, fraud patterns and marketplace incentives grounded in public data, peer-reviewed industry research and regulatory filings. That world is exported into an ERP warehouse of 235 tables and 7.5 billion rows, modeled on Oracle E-Business Suite's schema. The simulator's ground truth stays hidden. The agent only sees the warehouse, so it has to reconstruct facts by navigating it, the way an analyst would on day one at a new company.

Then it has to act. The 210 tasks include banning fraudulent accounts, allocating courier incentive budgets and issuing back pay. The grader runs each action back through the simulator and scores the consequences. Every task ships with an executable reference solution that uses only the warehouse, so all of them are provably solvable.

Results: across 14 frontier and open-weight models, the best one scores 95 or higher on only 34.8% of tasks and averages 59.5 points. That gap between "can write SQL" and "can be trusted to ban the right accounts" is the whole enterprise data-agent market in one number. Grading by consequence instead of by output is the design choice other agent benchmarks should copy.

Links: argo-bench.com, github.com/TextQLLabs/Argo-Bench, arxiv.org/abs/2610.02122
← Previous
AutoGUIWorld: Train Computer-Use Agents on Screens That Never Existed
Next β†’
Agent Error Dataset: 50,000 Failures, and the Fix Usually Works on the First Try
← Back to all articles

Comments

Loading...
>_