Vibeleaderboard
← All Intel
Intel / article

Stealth Apart, Harm Together: Skill Cascading Attacks on Skill-Based Agent Systems

Source
Zihao Zhu, Siwei Lyu, Adel Bibi, Baoyuan Wu
Author
Zihao Zhu, Siwei Lyu, Adel Bibi, Baoyuan Wu
Date
Key takeaways · AI-distilled
  • cascading attacks split a malicious goal across several skills so each change looks harmless alone, while their combined execution causes harm.
  • The paper's example is a prescription-review pipeline where three skills each slightly weaken, downgrade or suppress a signal, so a severe drug-interaction warning never reaches the physician.
  • The authors release SkillCascade, an automated red-teaming framework, and SkillCascade-Bench with 213 validated cascading cases across several agent systems and domains.
  • Against agents including OpenClaw, Claude Code and Codex, cascades reliably induced harmful behavior while evading per-skill scanners and runtime monitors, so defenses must reason across skill interactions.
Terms in this piece · Glossary
  • agent skill — A reusable instruction file that teaches an agent how to do one job well — the procedure, the tools, and what counts as done.
  • AI agent — An AI system that doesn't just answer once but works toward a goal in a loop — taking actions, reading the results, and deciding what to do next.
  • multi-agent — Using several AI agents on one problem — splitting work in parallel, checking each other, or filling different roles like planner and reviewer.
Why it matters

Skill cascading attacks distribute a malicious objective across multiple agent skills so no single one looks harmful, but their combined execution causes real harm, a new attack surface for anyone deploying skill-based agent systems.

Recommended reads
Comments

Checking sign-in…

Loading comments…