New Fellows Research: Can Claude autonomously align other AIs? We gave Claude 48 hours and 1 GPU to improve the alignment of small models. It researched and proposed methods, then trained and tested the models on its own. It worked surprisingly well. https://t.co/nhlCMgQl46
Claude “hill-climbed” safety benchmarks for common misalignments like deception or sycophancy, with one constraint: it had to preserve general capabilities. We then tested its best methods on held-out benchmarks to see if they'd generalize.
Across 10 alignment failures, Claude reliably improved safety scores without degrading capabilities. Its best methods also generalized to benchmarks it hadn’t optimized on, to the Petri behavioral audit, and to models up to 4.7x larger.

Could a model one day align its stronger successors? As a first test, we had Sonnet 5 post-train an early checkpoint of Opus 4.8, a more capable model. It reached safety scores approaching those of production Opus 4.8, which went through our full alignment training.

Automated work transfers to unseen benchmarks and larger models — but only for failures someone has a measure for, which puts the whole burden on design. The research setup is being released to build on.
Checking sign-in…
Loading comments…