October 7, 2026ResearchRLAgents

Two Papers on Agents Learning From Themselves: One Says Why It Breaks, One Says How to Do It

Test-time training lets a model write what it learns into its weights during deployment instead of into a memory file. Two papers posted Monday take opposite ends of the same problem: an agent that learns from its own output changes the thing that produces the next lesson.

The first, from KAUST, is a negative result with a mechanism. Across 128K-token streams, models that keep updating on their own generated text get worse at predicting independent human-written text, at 125M, 760M and 3B, and the same thing happens when Adam updates Qwen3-4B's existing weights. Writing is not the failure, because the same updates improve on real text. The culprit is the loop. Freezing the generator removes over 98% of the damage at the two smaller sizes. A paired one-update test shows the local conflict: each update predicts its own source better and new real text worse, and a few trajectories account for most of the big failures. Their fix is called Settlement: evaluate the candidate weights on independent real text before committing them. Mean endpoint gaps shrink to 0.07 and -0.02 nats while real-text adaptation is retained.

The second, from UNSW, is the constructive version for agents. ASCENT studies what the authors call online agentic test-time training: the agent runs each task once, gets one verification signal at the end, and that trajectory is the only training signal. Directly imitating or reinforcing the generated tokens destabilizes the policy, which is the first paper's finding in agent form. Instead, a frozen initial copy of the model receives the verified trajectory as privileged hindsight and predicts next-token distributions along it. Those distributions get distilled into persistent LoRA fast weights. Invalid-action turns are removed first. No external reference solution, no stronger teacher.

Read together they say one thing. An agent can improve from its own experience, but never on its own say-so. The signal has to pass through something the agent did not generate: independent text in one paper, a verifier plus a frozen reference model in the other. That is the same rule the verification thread keeps producing, this time inside the weights.

Links: arxiv.org/abs/2610.05076 and arxiv.org/abs/2610.05303
← Previous
UndoBench: Agents Finish 84% of Tasks and Recover From 47% of Faults
Next β†’
Google's SHIFT Builds a Different Agent Harness for Every Query, Without Running Any
← Back to all articles

Comments

Loading...
>_