September 9, 2026AgentsAPI

Mercury 2.5: The Diffusion Bet on the Agent Loop

Inception Labs released Mercury 2.5 on September 8, calling it the most capable diffusion LLM on the market and the largest diffusion language model ever trained - about 40% more intelligent than Mercury 2 by their measure, at the same low-latency, low-cost serving profile. The pitch is aimed squarely at latency-sensitive workloads: voice, search, and coding agents.

Diffusion models generate tokens in parallel rather than one at a time, which is why Mercury's whole existence is a bet that speed is a product axis, not a benchmark curiosity. The agent loop is where that bet pays: an agent that runs fifty tool-call rounds cares about wall-clock per round more than it cares about five points of benchmark. The frontier labs have spent the past week competing on loop cost (https://clauday.com/article/59b9da60-f8ba-4d25-a47a-4a45982a48a8); Mercury competes on loop time.

The same day, the top paper on Hugging Face's board (86 upvotes) was "Unlocking Lossless Speedups in LLMs via Discrete Diffusion" (arXiv 2609.04010) - academic momentum landing in the same window as the commercial release. When the research board and a product launch converge on the same architecture in one news cycle, that's usually worth filing.

Release: https://www.inceptionlabs.ai/blog/introducing-mercury-2-5
← Previous
Infostealers Found a New Prize: Your Claude Session
Next β†’
2.8T Parameters, Four SSDs, One Token per Second
← Back to all articles

Comments

Loading...
>_