October 11, 2026CodingOpen SourceInfrastructure

Prime Agent Rewrote Its Harness in Rust: 14x Faster Cold Start, 300,000 Downloads In

Prime Intellect shipped the Rust rewrite of Prime Agent on Thursday, and the numbers are the kind you get when a long-running daemon stops carrying a JavaScript runtime. Cold start drops from 736 milliseconds to 52, a 14x gain. Warm start is 13x faster. Memory for large sessions is 4.8x smaller. The install shrinks from 172 megabytes to 60. The harness has been downloaded more than 300,000 times since its August launch and has processed more than 8 trillion tokens, which is the scale at which a few hundred milliseconds per invocation starts to show up as a line in someone's cloud bill. The repo is MIT, sits at 21,800 stars, and the team cut six beta releases on Friday alone.

The reasoning in the post is about the shape of the thing rather than raw speed. Prime Agent is not a CLI you run and close; it is a daemon that spawns subagents, keeps sessions alive for hours and fans out tool calls in parallel. The TypeScript version hit three walls: no runtime type guarantees on the structured data flowing between model and tools, error handling that leaked through async boundaries, and a single-threaded event loop that serialized work the agent wanted to do concurrently. Rust's type system and ownership model were the fix for the first two and native threads for the third. Native Windows support ships as beta with this release, and Homebrew is now an install path.

This is a datapoint for the argument that the harness is where the engineering is now. A week ago ts-rust showed an LLM porting a compiler to Rust for $24,000; here a lab is porting its own agent harness to Rust by hand because the harness is the product, and it has to be fast enough to run thousands of times a day on a developer's machine without being noticed. Prime Intellect's whole business is RL environments and self-improving agents, and the rewrite is a statement that the loop around the model deserves the same optimization attention as the model.

The caveat is that performance claims from a rewrite are easy to make and hard to compare: same tasks, same model, but a different binary with different defaults. Nobody independent has benchmarked it yet. What is independently checkable is the download count, the token count, and the release cadence, and all three say people are running this thing a lot.

Blog post: https://www.primeintellect.ai/blog/prime-agent-rust
Repo: https://github.com/PrimeIntellect-ai/prime-agent
← Previous
Talorys: A Personal Agent on Cloudflare's Free Tier, One Command, No Telemetry
Next β†’
Epoch's InnovationEval: Given 3,000 GPU Hours, Frontier Agents Recovered 15% of One Human Idea
← Back to all articles

Comments

Loading...
>_