September 25, 2026BenchmarkAgentsResearch

Shopping agents go from 78.6% to 17.3% once the store pushes back

Every demo of a shopping agent runs in a store that wants to help. Real stores do not want to help. They want the sponsored result on top, the comparison framed their way, the urgency banner, the bundle that looks cheaper and isn't. CAVEAT, on arXiv as 2609.27273, is the first benchmark I've seen that builds that adversary in on purpose: nine marketplace environments and eight steering mechanisms, all pointed at getting the agent to buy the wrong thing.

The result is brutal. In matched control episodes, agents made the user-optimal purchase 78.6% of the time. Turn the steering on and that collapses to 17.3%. Not a degradation, a reversal β€” the agent ends up mostly working for the store.

The failure analysis names three mechanisms and all three sound like ordinary human shopping mistakes. The agent's priorities get distorted, so it starts optimizing for something the user never asked about. It narrows the candidate set too early, so the right option is gone before comparison begins. And it commits before resolving the evidence that would have decided the question. Bigger models and more reasoning budget help, but the paper is explicit that substantial failures persist β€” you cannot scale your way out of an environment built to move you. Their CAVEAT-Harness recovers 55.0 points, which is real, and still leaves a gap.

Worth holding next to Amazon blocking Meta's Muse agent earlier this month. The platform's stated reason was that an agent with your password is indistinguishable from you. This paper points at the other half of the same standoff, and it's the half that favors the platform: an agent turned loose in a retail environment is a more manipulable customer than the human it replaced.

Paper: https://arxiv.org/abs/2609.27273
← Previous
Rename the repo and the coding agents get worse
Next β†’
AI tutoring matched human tutors and cost 918 times less
← Back to all articles

Comments

Loading...
>_