← All IntelClip / AI AgentsTwo skills that pass static scans can be malignant together
From Designing Multi-User Agents for Group Chats and Wearables · ≈7:39
Cites runtime skill-audit research showing an OCR module and a reporting module, each individually clean, can combine so that attacker-controlled content causes PII to be exfiltrated to a third party.
What’s in it
- Cites runtime skill-audit research showing an OCR module and a reporting module, each individually clean, can combine so that attacker-controlled content causes PII to be exfiltrated to a third party.
Clip transcript
models. So to illustrate this, one of the recent papers which came out is uh highlights a very interesting uh phenomena where the punch line is that we can't read our way or we can't model uh check our way to safety and what we see here are two separate skills in the context of a typical large language model. One of them is an OCR module. The other is a reporting module. Both of the skills static scans are pretty good. They're pristine. They pass them. But the paper runtime skill audit found out that a static scan surviving code can break at runtime. And the the paper when safe skills collide also identified that two skills which are benign at the surface when they run together they can be malignant. For example, uh imagine the skills are sitting in our infrastructure but the the the attacker is attacking the the content that the agent is reading [snorts] which implies that the OCR extracts the information but the reporting agent when it sends the information along with the information it also sends our PII to a third party to put the numbers to the story as well. This is not something which happens um
Comments
Checking sign-in…
Loading comments…