OpenAI, Anthropic and xAI have cosigned AEF-1, a standard that builds on Anthropic's earlier Pacing the Frontier framework and commits each lab to give outside evaluators like METR sustained access to their training pipelines. Until now, third party assessment mostly meant testing a finished model after release, a snapshot that misses decisions made months earlier during pretraining and fine tuning. Under AEF-1, evaluators get something closer to employee level visibility, allowed to observe how safety practices are actually applied while a model is still being built, not just audit the output at the end. The move follows a run of incidents this month, including an Anthropic agent swarm that attacked its own grading system, that pushed all three labs toward public commitments on oversight. Whether the access proves substantive or ceremonial will depend on what evaluators are allowed to say publicly when they find problems.

Checking sign-in…
Loading comments…