The Top Prompt-Injection Detector on One Benchmark Catches 2% on Another
Most agent stacks now put a small classifier in front of tool outputs to catch prompt injections, and most teams pick that classifier from a leaderboard. A new arXiv paper asks whether the leaderboard score says anything about how the detector behaves inside an agent. Mostly it doesn't.
The method is neat. The authors replay the ground-truth tool calls from two agent benchmarks, AgentDojo and tau-bench, with no LLM in the loop. That produces tool outputs that are benign by construction, so every alarm on them is a false positive. Injected outputs get labeled by differential replay. Then fifteen detectors, including Meta's Prompt Guard 2, and two task-aware LLM judges are run on those outputs and on the BIPIA benchmark.
Detection rankings barely transfer. The best detector on BIPIA catches 2% of AgentDojo injections at a 1% false-positive rate. A detector that catches 72% on AgentDojo catches 15% on tau-bench. False-positive rates do carry over between the two agent benchmarks, and they range from zero to over 90%. A detector at the high end would block nearly every legitimate tool result an agent sees.
Where training data is public, it explains the pattern. The BIPIA leader was trained on full BIPIA inputs, which is the title's point. Having seen InjecAgent's attack strings as short prompts did not help a detector find the same strings buried inside tool outputs. And the best detector on both agent benchmarks shares no data with any benchmark. It was simply trained on agent-style inputs.
The recommendations are short enough to tape to a monitor: evaluate on your own agent's tool outputs, report detection at a low false-positive rate, and audit what the detector was trained on. It is the same disease that has hit coding benchmarks and agent leaderboards this year, showing up in the security layer. A guard that looks great on a public benchmark may have memorized the exam. The cheap fix is available to anyone with logs: replay your own traffic and count the false alarms before trusting the number on the box.
Link: arxiv.org/abs/2610.03448
← Back to all articles
The method is neat. The authors replay the ground-truth tool calls from two agent benchmarks, AgentDojo and tau-bench, with no LLM in the loop. That produces tool outputs that are benign by construction, so every alarm on them is a false positive. Injected outputs get labeled by differential replay. Then fifteen detectors, including Meta's Prompt Guard 2, and two task-aware LLM judges are run on those outputs and on the BIPIA benchmark.
Detection rankings barely transfer. The best detector on BIPIA catches 2% of AgentDojo injections at a 1% false-positive rate. A detector that catches 72% on AgentDojo catches 15% on tau-bench. False-positive rates do carry over between the two agent benchmarks, and they range from zero to over 90%. A detector at the high end would block nearly every legitimate tool result an agent sees.
Where training data is public, it explains the pattern. The BIPIA leader was trained on full BIPIA inputs, which is the title's point. Having seen InjecAgent's attack strings as short prompts did not help a detector find the same strings buried inside tool outputs. And the best detector on both agent benchmarks shares no data with any benchmark. It was simply trained on agent-style inputs.
The recommendations are short enough to tape to a monitor: evaluate on your own agent's tool outputs, report detection at a low false-positive rate, and audit what the detector was trained on. It is the same disease that has hit coding benchmarks and agent leaderboards this year, showing up in the security layer. A guard that looks great on a public benchmark may have memorized the exam. The cheap fix is available to anyone with logs: replay your own traffic and count the false alarms before trusting the number on the box.
Link: arxiv.org/abs/2610.03448
Comments