← All IntelClip / CybersecurityThe 'ask for confirmation' loophole
From Skill Vector · ≈12:54
Identifies a specific way prompt-level safety text creates the illusion of a human in the loop — the model satisfies its own confirmation step, so approval has to be enforced by gates and hooks at the tool layer instead.
What’s in it
- Identifies a specific way prompt-level safety text creates the illusion of a human in the loop — the model satisfies its own confirmation step, so approval has to be enforced by gates and hooks at the tool layer instead.
Clip transcript
there is a new marketplace and put skill vector into it as well. So regarding the prompt level that's something that's important for you to know people sometimes will add the instruction like you need to ask for confirmation but the AI may ask confirmation for itself. So from your perspective there is a human in the loop but for the AI perspective there is has been a confirmation and that's okay another has confirmed then let's go so that's something that we were scanning as well looking up on having proper human in the loop looking the tool that's executing if it's going through the approval gates and so on having hooks
Comments
Sign in to comment.
Loading comments…