October 8, 2026ResearchAgentsSkills

DAEDALUS: an Agent That Writes Its Own Practice Problems to Build Memory

Most agent memory systems assume you already have training tasks and a verifier. DAEDALUS (arXiv 2610.08048, from a team including Antoine Edy and Gautier Viaud) drops both. It pairs an explorer agent that pokes at a new environment and invents tasks that are hard but solvable, with a solver that attempts them. Every solver failure yields a candidate heuristic. A heuristic is accepted only after the solver repeatedly succeeds with it in context, and solver outcomes feed back to the explorer to tune difficulty. The accepted heuristics consolidate into a memory bank the agent carries into real tasks.

The numbers are solid. Across AppWorld, tau-squared-bench and AutomationBench, DAEDALUS lifts mean success by up to 15.9 points and pass-to-the-5 by up to 2.2x over a no-memory baseline, and is competitive with methods that do have training tasks, at lower inference cost than most. Gains show up with a small exploration budget. The heuristics transfer across model families. And the ablations say the key information comes from solver traces, which is the same finding GUI-HARVEST made this week about harness diagnosis: the evidence is in the executed trajectory, not the model's self-report.

There is a bonus result that matters for anyone running evals. The tasks DAEDALUS generates can rank models by performance about as well as the real benchmark tasks do. An agent that writes its own curriculum to learn an environment has, as a side effect, written a benchmark for that environment, which is roughly what AutoSciBench (2610.05140) set out to do on purpose for scientific agents the same week.

What to watch is the acceptance rule. "Keep a rule only after it repeatedly works in context" is the write-time curation the memory thread has been circling since Memadapter showed that even correct memories can induce sycophancy. DAEDALUS gates on outcomes, not on plausibility. Code and artifacts are linked from the paper.

Link: arxiv.org/abs/2610.08048
← Previous
Two Papers, One Finding: Agents Collect the Evidence and Then Ignore It
Next β†’
NeMo-DCR: Nvidia Cuts a 1T-Parameter RL Weight Sync From 87 Minutes to 150 Seconds
← Back to all articles

Comments

Loading...
>_