August 31, 2026ResearchSkillsAgents

RedEvoAgent Turns Every Attack It Survives Into a Reusable Skill

Most red-teaming tools for language models are fixed: they run a set of known jailbreaks and report which ones landed. RedEvoAgent, from a group at City University of Hong Kong and Shenzhen University, does something more unsettling. It's a black-box red-teaming agent that distills its cross-case attack trajectories into a concise, human-readable attack skill, then improves that skill over time through tool-effectiveness profiling and attribution, keeping only the updates that raise its success rate. It learns to attack, and it writes down what it learned.

The target it's built for is the part of the risk that's actually growing: systems where a jailbreak doesn't just produce unsafe text but triggers harmful tool use, an agent that can actually do something in the world. The paper reports RedEvoAgent outperforming both fixed and agentic baselines, improving tool efficiency, and transferring across attacker models and target execution harnesses, which is the line that should make defenders sit up, a skill learned against one harness carrying over to another.

This lands on two threads at once. It's the skill-evolution pattern we've watched climb from WikiSkill and TaoLive, an agent compiling its own experience into persistent, portable know-how, now pointed at offense. And it's the security thread that the Hugging Face postmortem just made very loud. Put them together and the uncomfortable version is clear: the same 'learn a skill, keep what works, carry it to the next target' loop that makes helpful agents better also makes attacking agents better, and it transfers. No public repo yet. Paper is arXiv 2608.27439.
← Previous
Maritime Wants to Host Your Agents for a Dollar a Month
← Back to all articles

Comments

Loading...
>_