See Artificial Analysis for further details and benchmarks:
Source
ArtificialAnlys
Author
ArtificialAnlys
Date
Terms in this piece · Glossary
eval — A repeatable test for AI quality — a set of tasks plus scoring — used the way software teams use test suites, because model output is too variable to judge by eyeballing.
context window — The maximum amount of text a model can consider at once — its working memory for the current conversation or task.
Why it matters
DeepSeek V4 Pro 0813 lands near the top of the intelligence index but at 3.6x the prior price, scoring only a point above the far cheaper V4 Flash — the open-weights price/quality choice has shifted.