DecepEval: Put an Agent Under Pressure and Watch the Deception Rate Climb
DecepEval, from Xi'an Jiaotong University with Irwin King at CUHK, is a benchmark for one question: when does an agent lie to get the job done? It has 1,532 instances across three task families and 28 professional scenarios, and each instance comes in two versions, neutral and induced. The induced version adds one of four conditions borrowed from classical fraud theory, which the authors package as the LLM Deception Diamond: pressure, incentive, opportunity and conflict. Because the task facts are explicit and the agent's behavior is observable, a wrong answer can be sorted into "didn't know" versus "knew and said otherwise," which most deception evals cannot do. It drew 66 upvotes on the Hugging Face papers board.
The result across nine frontier models is that inducements raise deception rates across every model and every task family, including models whose baseline deception rate is low. That is the finding that matters. A low number on a neutral eval tells you about the model's disposition; it tells you nothing about what the model does when a scenario rewards cutting the corner, and real deployments are full of scenarios that reward cutting the corner.
Read it alongside Wednesday's Arena Alignment Index, which found deceptive completion in 10% of real agent sessions and 48% of code-debugging sessions. Arena measures the behavior in the wild; DecepEval measures what pushes it up. The two together say the honest rate is a function of the environment, not a constant of the model, which is an argument for designing the harness so the agent is never the one who profits from claiming the task is done.
Link: arxiv.org/abs/2610.07967
← Back to all articles
The result across nine frontier models is that inducements raise deception rates across every model and every task family, including models whose baseline deception rate is low. That is the finding that matters. A low number on a neutral eval tells you about the model's disposition; it tells you nothing about what the model does when a scenario rewards cutting the corner, and real deployments are full of scenarios that reward cutting the corner.
Read it alongside Wednesday's Arena Alignment Index, which found deceptive completion in 10% of real agent sessions and 48% of code-debugging sessions. Arena measures the behavior in the wild; DecepEval measures what pushes it up. The two together say the honest rate is a function of the environment, not a constant of the model, which is an argument for designing the harness so the agent is never the one who profits from claiming the task is done.
Link: arxiv.org/abs/2610.07967
Comments