
A rigorous, research- map of how jailbreaks and adversarial prompts actually work against aligned LLMs, giving engineers the vocabulary and threat models needed to red-team and harden their own AI products.
“However, adversarial attacks or jailbreak prompts could potentially trigger the model to output something undesired.”
“A large body of ground work on adversarial attacks is on images, and differently it operates in the continuous, high-dimensional space.”
“Attacks for discrete data like text have been considered to be a lot more challenging, due to lack of direct gradient signals.”
Checking sign-in…
Loading comments…