Stealth Apart, Harm Together: Skill Cascading Attacks on Skill-Based Agent Systems
Source
Zihao Zhu, Siwei Lyu, Adel Bibi, Baoyuan Wu
Author
Zihao Zhu, Siwei Lyu, Adel Bibi, Baoyuan Wu
Date
Key takeaways · AI-distilled
agent skillA reusable instruction file that teaches an agent how to do one job well — the procedure, the tools, and what counts as done.Full definition → cascading attacks split a malicious goal across several AI agentAn AI system that doesn't just answer once but works toward a goal in a loop — taking actions, reading the results, and deciding what to do next.Full definition → skills so each change looks harmless alone, while their combined execution causes harm.
The paper's example is a prescription-review pipeline where three skills each slightly weaken, downgrade or suppress a signal, so a severe drug-interaction warning never reaches the physician.
The authors release SkillCascade, an automated multi-agentUsing several AI agents on one problem — splitting work in parallel, checking each other, or filling different roles like planner and reviewer.Full definition → red-teaming framework, and SkillCascade-Bench with 213 validated cascading cases across several agent systems and domains.
Against agents including OpenClaw, Claude Code and Codex, cascades reliably induced harmful behavior while evading per-skill scanners and runtime monitors, so defenses must reason across skill interactions.
Terms in this piece · Glossary
agent skill — A reusable instruction file that teaches an agent how to do one job well — the procedure, the tools, and what counts as done.
AI agent — An AI system that doesn't just answer once but works toward a goal in a loop — taking actions, reading the results, and deciding what to do next.
multi-agent — Using several AI agents on one problem — splitting work in parallel, checking each other, or filling different roles like planner and reviewer.
Why it matters
Skill cascading attacks distribute a malicious objective across multiple agent skills so no single one looks harmful, but their combined execution causes real harm, a new attack surface for anyone deploying skill-based agent systems.