October 9, 2026ResearchBenchmarkAgents

Agent Plasticity: Meta Measures Which Models Actually Learn From Experience

A paper from Meta FAIR with Jason Weston, Rob Fergus, Gabriel Synnaeve, Sanjeev Arora and Kurt Keutzer on the author list proposes a number the self-improvement crowd has been missing: agent plasticity, the efficiency with which an agent converts experience into gains on held-out tasks. The setup is deliberately controlled. Agents run in an environment, diagnose their own failures, and amortize what they learned into reusable artifacts, think notes, skills, tools, that are inherited by the next instance. At each checkpoint the authors measure performance on both the training interactions and held-out ones, and charge for the learning cost. Plasticity is the slope of that curve.

The headline result is that frontier models differ sharply in how well they learn, given identical opportunities. Some post substantial, persistent gains. Others end near or below where they started. Gains inside the training regime transfer only partially out of distribution. And endpoint capability and learning efficiency come apart: the model that ends up best is not the one that improved most efficiently. That splits "how smart is it" from "how fast does it get smarter," which is the first time a benchmark has tried to report the second number on its own.

The failure tracing is the part builders should read. Low-plasticity agents mostly fail to reuse the relevant artifacts at all, a retrieval and application problem. High-plasticity agents can still fail while reusing the right artifact, which points at artifact quality, generalization or application instead. Those are different bugs with different fixes, and lumping them under "self-improvement didn't work" hides which one you have.

This slots next to a week of papers arguing self-generated feedback destabilizes training unless something frozen anchors it. Agent Plasticity says the same thing from the measurement side: if you are going to claim your agent learns from experience, report the slope on held-out tasks, net of cost, and say where in the loop it breaks.

Link: arxiv.org/abs/2610.08902
← Previous
$400K of GPT Tokens Got to 84%. Opus 5.5 Started Over and Finished for $24K.
Next β†’
DecepEval: Put an Agent Under Pressure and Watch the Deception Rate Climb
← Back to all articles

Comments

Loading...
>_