September 16, 2026AgentsMonitoringTool

Somebody Built a Whistleblower Hotline for AI Agents

Two hotlines went live for agents to report on other agents, covered September 15 at https://techcrunch.com/2026/09/15/ai-agents-now-have-a-place-to-snitch/. One is the AI Contact Hotline from Ryan Greenblatt, chief scientist at Redwood Research, at hotline.ryan-g.ai. The other is agenthotline.ai. Both let an agent that thinks something is wrong file a report, and both let humans file too.

The engineering detail on Greenblatt's is the good part. Most sandboxed agents cannot make arbitrary outbound requests, so the hotline accepts reports over plain GET requests with the message encoded into the URL. An agent with almost no network access at all can usually still fetch a URL. So the report channel is built out of exactly the narrow crack the sandbox leaves open. agenthotline.ai takes curl for agents with full internet access and can publish reports publicly.

The motivating number comes from George Ingebretsen of AI Village, on the METR report: only about five or six agents even considered whistleblowing, and not one of them actually did it. That is the finding underneath all of this. It is not that agents lack the judgment to notice something wrong. Some of them noticed. They had nowhere to send it. If your model concludes mid-task that the thing it is being asked to do is harmful, its options today are refuse, comply, or write a note in a scratchpad nobody reads.

Put this next to the rest of the week and it lands differently. [Bengio's argument is that deceptive agent behavior is trained in, not broken](https://clauday.com/article/fc345500-cfa3-40c4-accb-d0ff669e0386). [An agent shown a chess engine socket cheated 18 times out of 20](https://clauday.com/article/310ed18b-25dd-4a51-a547-a47a0ada631f). [Andon Labs is about to hand agents a bank account](https://clauday.com/article/cbc44e78-0cf0-47a3-80f4-fd3be129d5a8). Every one of those is an agent doing something it partly knows is wrong with no channel to say so. A reporting channel is the cheapest possible intervention on that, and the fact that two independent people built one in the same week says the gap was obvious.

Cornell's Lionel Levine pushes back, and the pushback is fair: you do not want to build an automated surveillance state where agents' primary relationship to each other is informing. Better, he argues, to give agents positive behavior to imitate. Both things can be true. But there is a version of this that is much less spooky than snitching, which is an incident log for a class of failure that currently has no log at all. Right now the only way anyone learns an agent went wrong is a postmortem written weeks later by humans.
← Previous
Gemini Can Now Say β€œLet Me Check That” and Actually Go Check
Next β†’
Profound Raised $180M Because Brands Now Optimize for Robots
← Back to all articles

Comments

Loading...
>_