The changing set of models that lead meaningful workloads at a point in time.
There is no permanent best model. Leadership differs across coding, reasoning, tool use, long context, vision, audio, latency, price, and reliability, and benchmark results can be contaminated or optimized against. Treat the frontier as a dated live view, then validate candidates on representative tasks before committing.