
It is how you train a model to refuse well without a human labelling every bad output: the model critiques and revises its own responses against an explicit written constitution. The pattern generalises far past safety — it is how you scale any judgment you can write down.
Sign in to comment.
Loading comments…