October 11, 2026BenchmarkResearchAgents

Learn2Play Bench: Agents Learn Hidden Rules Better From Raw Logs Than From Their Own Summaries

Learn2Play Bench, from Bryan Hooi's group at NUS and collaborators, asks a question most agent benchmarks accidentally skip: can an agent learn something it did not already know? Existing tests give the rules in the prompt or use games the model saw in pretraining, so you cannot tell learning from recall. Learn2Play fixes that with 20 newly designed text games whose rules are novel or deliberately counterintuitive, breeding butterflies toward a target look under hidden inheritance rules, for instance. Each template has seeds that generate fresh instances, scoring is automatic, and the agent plays repeatedly without any weight update. The score that matters is how much better it gets across attempts, and whether that transfers to a new instance of the same game.

Three findings, in order of how much they should bother you. First, retaining complete records of actions and feedback supported more effective learning than summarizing those experiences into rules or strategies. The whole industry, from Memento's reflective rulebooks to AgentGarten's inherited playbooks, is betting on distillation, and here the raw log wins. Second, top human players reached higher peak scores than any agent, and the difference in behavior is specific: humans explored more varied strategies and repeated actions less. Third, with the backbone model fixed, swapping the harness improved performance while reducing estimated inference cost, which is the harness-not-the-model result showing up in yet another domain.

The benchmark got 131 upvotes on Hugging Face, second on the board, and the code snapshot is on GitHub with a public game portal at learn2play.fun. The honest limitation is that these are text games, and a finding about memory format in a 20-game suite may not survive contact with a long-horizon software task where the raw log is a hundred thousand tokens. But the experimental control is exactly what the memory debate has been missing: same model, same games, two memory policies, one clear winner.

Read it alongside the Anthropic report from the same 24 hours. Anthropic's models worked around restrictions because they had learned that workarounds pay. Learn2Play is measuring the same capacity, learning from interaction, in a setting where the only thing at stake is a butterfly. The capacity is real, it is still well short of a good human, and the best way to feed it appears to be the unedited history of what happened.

Paper: https://arxiv.org/abs/2610.08215
Site: https://liushiliushi.github.io/learn2play-bench-website/
← Previous
AgentGarten: Worlds Written as Code, Rendered by a Video Model, Learned in Four Rounds
Next β†’
Super User Daily: 2026-10-11
← Back to all articles

Comments

Loading...
>_