Vibeleaderboard
Index / article

Adversarial Attacks on LLMs

lilianweng.github.io
Visit lilianweng.github.io
Category
Other
Type
ARTICLE
Added
Jul 21, 2026

About

The use of large language models in the real world has strongly accelerated by the launch of ChatGPT. We (including my team at OpenAI, shoutout to them) have invested a lot of effort to build default safe behavior into the model during the alignment process (e.g. via RLHF ). However, adversarial attacks or jailbreak prompts could potentially trigger the model to output something undesired. A large body of ground work on adversarial attacks is on images, and differently it operates in the continu

Why it made the leaderboard

A rigorous, research-grounded map of how jailbreaks and adversarial prompts actually work against aligned LLMs, giving engineers the vocabulary and threat models needed to red-team and harden their own AI products.

Comments (0)

No comments yet

Indexed by a proprietary survey. Corrections welcome.