DeepSeek V4.1 Flash overtakes DeepSeek V4 Pro 0813 as DeepSeek’s new flagship model with a score of 40 on Artificial Analysis Intelligence Index. At just 552B parameters, it outperforms the Pro (1.6T) model while costing ~4x less per token, placing it just short of the Intelligence vs. Cost Pareto frontier because of its verbosity @deepseek_ai has released DeepSeek V4.1 Flash, the successor to DeepSeek V4 Flash 0731. This is a 552B model features a new causal Encoder–Decoder architecture, allowing it to have just 8B active parameters for input and 16B active parameters for output. On the first-party API, DeepSeek V4.1 Flash is priced at $0.30 per 1M input tokens and $1.20 per 1M output tokens, with cached input tokens priced at just $0.006 per 1M tokens, a 98% discount. Off-peak pricing provides a further 50% discount across input, cached input, and output tokens. V4.1 Flash is ~20% cheaper than DeepSeek V4 Flash 0731 and ~4x cheaper than DeepSeek V4 Pro 0813 while delivering a higher performance. Key results: ➤ DeepSeek V4.1 Flash makes gains in agentic capabilities and long context reasoning. DeepSeek V4.1 Flash scores 27% in Terminal-Bench v4.0, more than double DeepSeek…

A smaller, cheaper model outperforming DeepSeek's own flagship on agentic and long- benchmarks resets the cost/performance baseline engineers should compare other models against.
“DeepSeek V4.1 Flash is one of the most verbose models we’ve measured at 89k Tokens per Intelligence Index Task.”
“Despite such verbosity, DeepSeek V4.1 Flash still costs just $0.27 per Intelligence Index task.”
“DeepSeek V4.1 Flash takes first place on AutomationBench-AA with 69%, equal to GPT-6 Astra (69%) and slightly above Grok 4.6 (67%).”
postArtificial Analysis Adds Per-Effort-Level Model Comparison PagesChecking sign-in…
Loading comments…