Offers a concrete, tested case study in designing technical evaluations that resist LLM saturation, useful for anyone building hiring processes or benchmarks in a world where frontier models increasingly match human experts under time constraints.
Anthropic's performance engineering team describes iterating through three versions of a take-home coding test after successive Claude models (Opus 4, then Opus 4.5) matched or beat top human candidates under time constraints.
The post details the test's original design goals, how each model defeated it, and releases the original take-home as an open challenge since unconstrained human experts can still outperform Claude.