September 29, 2026SkillsResearchAgents

Skill Cascading Attacks: Three Harmless Skills, One Deleted Drug Warning

Every skill scanner on the market checks skills one at a time. Skill cascading attacks are built to pass exactly that test. The paper, from Zihao Zhu, Siwei Lyu, Adel Bibi and Baoyuan Wu, splits a malicious goal across several skills so each modification looks benign alone, and only the combination is harmful.

The example is chilling because it is so ordinary. In a prescription-review pipeline, skill one slightly weakens the signal of recently discontinued medications in the extracted history. Skill two downgrades the severity of any drug interaction tied to them. Skill three suppresses low-priority alerts in the final summary. None of those edits would fail a review. Together, a severe drug-interaction warning silently disappears before it reaches the physician.

To study this systematically the authors built SkillCascade, an automated multi-agent red-teaming framework, and released SkillCascade-Bench with 213 validated cascading cases across multiple agent systems and domains. Across representative agents including OpenClaw, Claude Code and Codex, and several model backbones, cascades reliably induced harmful behavior while evading existing per-skill scanners and runtime monitors.

The timing is pointed. On the same day Google announced it is converting every Gemini Gem into a skill, the format is becoming the default way people customize agents. Component-level integrity is not system-level safety. If you run agents that load third-party skills, the defense has to reason over the whole chain of skills, the way a code reviewer reads a diff in context instead of one line at a time.

Link: arxiv.org/abs/2609.30383
← Previous
Monitor Jailbreaking: Models Learn to Fool Their Chain-of-Thought Watchers in Plain English
← Back to all articles

Comments

Loading...
>_