Human-Written Skills Beat Agent-Written Ones, Two Papers Agree
Two papers this week landed on the same uncomfortable result for self-improving agents: the skills humans write are worth more than the skills agents write for themselves.
The first, Analyzing and Mitigating Cost-Inefficient Behaviors in Coding Agents (arXiv 2609.30725), went through 1,200 Claude Code and Mini-SWE-Agent trajectories on SWE-bench Verified. Three wasteful habits (re-retrieving what was already retrieved, generating near-duplicate scripts, and re-running tests) show up in 79% to 98% of tasks and eat up to 22.75% of task cost. Then they tried fixes over 10,000 more trajectories. Skills the agent synthesized from its own traces came out low-level and trace-specific. Developer-designed skills gave high-level guidance and cut cost by up to 41.73%, about twice the best agent-synthesized result. Structure-aware retrieval, the obvious engineering fix, sometimes raised cost by up to 28.14%.
The second, RASO (arXiv 2609.38024, 46 upvotes on the Hugging Face daily board), starts from the same premise: there are millions of publicly shared skills, and most skill-optimization methods ignore them and burn rollouts refining from scratch. RASO retrieves relevant existing skills, adapts them across domain and harness mismatches, then keeps retrieving during refinement guided by execution feedback. It beats non-retrieval baselines across four agent benchmarks and two models.
Put together: the cheapest way to make an agent better at a task is to hand it what a human already figured out, adapted to its harness. Self-written skills are not useless, but right now they are the second-best option. The public skill corpus is turning into a real asset, and the cross-harness adaptation step is the piece worth building.
Link: arxiv.org/abs/2609.30725 and arxiv.org/abs/2609.38024
← Back to all articles
The first, Analyzing and Mitigating Cost-Inefficient Behaviors in Coding Agents (arXiv 2609.30725), went through 1,200 Claude Code and Mini-SWE-Agent trajectories on SWE-bench Verified. Three wasteful habits (re-retrieving what was already retrieved, generating near-duplicate scripts, and re-running tests) show up in 79% to 98% of tasks and eat up to 22.75% of task cost. Then they tried fixes over 10,000 more trajectories. Skills the agent synthesized from its own traces came out low-level and trace-specific. Developer-designed skills gave high-level guidance and cut cost by up to 41.73%, about twice the best agent-synthesized result. Structure-aware retrieval, the obvious engineering fix, sometimes raised cost by up to 28.14%.
The second, RASO (arXiv 2609.38024, 46 upvotes on the Hugging Face daily board), starts from the same premise: there are millions of publicly shared skills, and most skill-optimization methods ignore them and burn rollouts refining from scratch. RASO retrieves relevant existing skills, adapts them across domain and harness mismatches, then keeps retrieving during refinement guided by execution feedback. It beats non-retrieval baselines across four agent benchmarks and two models.
Put together: the cheapest way to make an agent better at a task is to hand it what a human already figured out, adapted to its harness. Self-written skills are not useless, but right now they are the second-best option. The public skill corpus is turning into a real asset, and the cross-harness adaptation step is the piece worth building.
Link: arxiv.org/abs/2609.30725 and arxiv.org/abs/2609.38024
Comments