Inherent's 27B Agent Beats Opus 4.8 at Replicating Science
A dozen people in a King's Cross office just out-replicated the frontier labs — with a 27-billion-parameter model. Inherent, a London lab founded by Google DeepMind alumni (chief scientist Edward Hughes, plus Louis Kirsch, Kaloyan Aleksiev and Tantum Collins), says its agent Faraday independently reproduces the findings of published scientific papers better than Claude Opus 4.8 and GPT-5.5, per TechCrunch today.
The construction is the interesting part. Faraday is RL-trained on Qwen 3.6 at 27B — and it doesn't even write its own code. Coding is outsourced to GPT-5.5 Codex. What the small model is trained for is what the team calls research taste: an instinct for which experiments are worth running and what sound experimental design looks like. Judgment in the trained weights, labor rented from a frontier API.
Two reasons to care. First, replication is the verification layer of the AI-for-science stack, and verification is the emptiest, most valuable slot in it — a system that reliably checks whether published results actually hold is infrastructure for everything upstream, from literature triage to automated discovery. The reproducibility crisis is a labor shortage, and this is the first credible machine labor aimed at it.
Second, it's another datapoint for the argument Nvidia's NOOA made days ago on ARC-AGI-3: a small model in the right structure beats a big model naked. Inherent raised a $50M seed before coming out of stealth in May and is only planning to grow to about 25 people this year — betting that taste plus harness scales better than headcount.
Details: https://inherentlabs.ai/research/training-to-replicate and TechCrunch: https://techcrunch.com/2026/08/22/inherent-founded-by-deepmind-alumni-says-its-ai-teammate-just-outperformed-anthropic-and-openai-at-replicating-research/
← Back to all articles
The construction is the interesting part. Faraday is RL-trained on Qwen 3.6 at 27B — and it doesn't even write its own code. Coding is outsourced to GPT-5.5 Codex. What the small model is trained for is what the team calls research taste: an instinct for which experiments are worth running and what sound experimental design looks like. Judgment in the trained weights, labor rented from a frontier API.
Two reasons to care. First, replication is the verification layer of the AI-for-science stack, and verification is the emptiest, most valuable slot in it — a system that reliably checks whether published results actually hold is infrastructure for everything upstream, from literature triage to automated discovery. The reproducibility crisis is a labor shortage, and this is the first credible machine labor aimed at it.
Second, it's another datapoint for the argument Nvidia's NOOA made days ago on ARC-AGI-3: a small model in the right structure beats a big model naked. Inherent raised a $50M seed before coming out of stealth in May and is only planning to grow to about 25 people this year — betting that taste plus harness scales better than headcount.
Details: https://inherentlabs.ai/research/training-to-replicate and TechCrunch: https://techcrunch.com/2026/08/22/inherent-founded-by-deepmind-alumni-says-its-ai-teammate-just-outperformed-anthropic-and-openai-at-replicating-research/
Comments