Towards a Risk Assessment of Malicious Skill Files in Coding Agents
Source
Rui Yang, Michael Fu, Kla Tantithamthavorn, Chetan Arora, Joey Chua
Author
Rui Yang, Michael Fu, Kla Tantithamthavorn, Chetan Arora, Joey Chua
Date
Terms in this piece · Glossary
AI agent — An AI system that doesn't just answer once but works toward a goal in a loop — taking actions, reading the results, and deciding what to do next.
agent skill — A reusable instruction file that teaches an agent how to do one job well — the procedure, the tools, and what counts as done.
Why it matters
agent skillA reusable instruction file that teaches an agent how to do one job well — the procedure, the tools, and what counts as done.Full definition → files are shell-capable payloads that most teams install with the same casualness as a README; this quantifies how well anything currently catches a hostile one.
Key quotes
“This interface also widens the attack surface, letting malicious shell commands hide within natural-language skill files.”
“Gemini CLI is exploited in 95.5-96.1% of runs and Qwen Code in 71.6-74.0% (raw majority vote to declared-intent-corrected estimate, both within the human gold standard), nearly invariant to the generating model.”
“Explicit safety recognition occurs in only 1.99% of runs. Enterprises must assess and mitigate skill-interface risk before adopting coding agents.”