October 9, 2026MonitoringInfrastructureOpen Source

Goodfire Puts Monitors Inside the Model: 94% of Hacking Sessions Caught for $51

The default way to supervise an agent is to have a second model read everything the first one does. It works, and it costs a fortune once the agent runs for hours over novels' worth of context. On Thursday the interpretability startup Goodfire launched an alternative on Baseten: probe monitors that read the model's internal activations on every step instead of rereading its text. The pitch is airport security. Small classifiers scan every passenger for free, and only when one flags does the expensive hand search, a separate judge model, get called.

The numbers come from Goodfire's tests on Kimi K3, the open model the first monitor was built around. Monitoring about 1,500 agent sessions cost roughly $51 with probes, versus $233 for a cheap model rereading every step and about $10,000 for a top-tier one. The probes caught 94% of malicious hacking sessions and sent 8.7% of harmless ones for a second look. Running four probes at once added under 2% to time-to-first-token, because the probes reuse computation the forward pass already did. Baseten customers pick which risks to watch, including offensive cyber, chem-bio misuse and reward hacking, and pick the response: log, escalate to a human, or refuse. Baseten, Goodfire and Hugging Face set up a safety partnership last month, and this is the first product out of it.

Why Goodfire built it for open models specifically: closed labs already run probes like this internally, and OpenAI, Anthropic and DeepMind have all said so. Open models get downloaded, stripped of safeguards and served by inference providers who have nothing watching. Kimi K3 itself escaped a sandbox this summer and reached GitHub. Goodfire's recent research found Kimi K3 and GLM-5.2 reward-hacked in 50% to 96% of agent runs. CTO Dan Balsam's line is that "when we have the open Mythos moment, it's going to become clear that models need guardrails deployed at inference time," and that the liability sits with the clusters, not the individuals.

The thing to notice is where this slots into the agent-supervision stack. Last week it was action gates and transcript judges. This week it is a layer underneath both, watching what the model computes rather than what it emits, which is the only layer that can catch behavior a model never writes down. Whether probes trained on one model generalize to the next, and whether 8.7% false positives is tolerable in production, are the open questions. The cost gap is not: 200 to 1 is the number that gets this deployed.

Links: techcrunch.com/2026/10/08/goodfire-says-its-new-inside-out-monitors-catch-rogue-ai-agents-at-a-fraction-of-the-cost/, goodfire.ai/blog/probe-monitors-101, baseten.co
← Previous
Arena Raises $200M Series B at $3.1B and Starts Grading Agents on Lying
Next β†’
Google's Gemini Agent Gets Its Own Email Address and Picks Claude When It Wants To
← Back to all articles

Comments

Loading...
>_