Adversarial Attacks on LLMs
lilianweng.github.io- Category
- Other
- Type
- ARTICLE
- Builder
- @lilianweng
- Added
- Jul 21, 2026
About
The use of large language models in the real world has strongly accelerated by the launch of ChatGPT. We (including my team at OpenAI, shoutout to them) have invested a lot of effort to build default safe behavior into the model during the alignment process (e.g. via RLHF ). However, adversarial attacks or jailbreak prompts could potentially trigger the model to output something undesired. A large body of ground work on adversarial attacks is on images, and differently it operates in the continu
Why it made the leaderboard
A rigorous, research-grounded map of how jailbreaks and adversarial prompts actually work against aligned LLMs, giving engineers the vocabulary and threat models needed to red-team and harden their own AI products.
Comments (0)
No comments yet
Indexed by a proprietary survey. Corrections welcome.