Vibeleaderboard
← All Intel
Intel / article

Constitutional AI: Harmlessness from AI Feedback

Source
arxiv.org
Author
Yuntao Bai et al.
Date
Why it matters

It is how you train a model to refuse well without a human labelling every bad output: the model critiques and revises its own responses against an explicit written constitution. The pattern generalises far past safety — it is how you scale any judgment you can write down.

Recommended reads
Comments

Checking sign-in…

Loading comments…