Can AI automate Epoch? We're introducing Epoch Automation Reports to evaluate frontier models on realistic, open-ended tasks drawn from our own work.
Claude Fable 5.1 and GPT-6 Astra lead, yet they are far from fully automating Epoch’s work.
Shows where frontier models still fail on real research work: they miss house style and report experiment-setup errors as findings. Useful for judging how far to trust agents on open-ended tasks.