Vibeleaderboard
← All Intel
Intel / post

DeepSeek V4.1 Flash Beats Its Own Flagship on Price and Performance

Source
ArtificialAnlys
Date
From the Daily Brief

Artificial Analysis benchmarked DeepSeek V4.1 Flash against the company's own V4 Pro flagship and found the smaller model wins on its Intelligence Index despite using only 8 billion to 16 billion active parameters against V4 Pro's 1.6 trillion. It also costs roughly four times less per token, a gap wide enough to change which model a team defaults to for agentic and long context workloads. V4.1 Flash shipped yesterday with native vision support and a smaller KV cache, and this result is the first independent evidence that those architectural changes translate into a real performance gain rather than just efficiency. It resets the baseline other vendors get compared against: a model an order of magnitude smaller is now the one to beat on both price and reasoning benchmarks, not the flagship it was meant to complement.

Read the 2026-09-11 Brief →
ArtificialAnlys@ArtificialAnlys

DeepSeek V4.1 Flash overtakes DeepSeek V4 Pro 0813 as DeepSeek’s new flagship model with a score of 40 on Artificial Analysis Intelligence Index. At just 552B parameters, it outperforms the Pro (1.6T) model while costing ~4x less per token, placing it just short of the Intelligence vs. Cost Pareto frontier because of its verbosity @deepseek_ai has released DeepSeek V4.1 Flash, the successor to DeepSeek V4 Flash 0731. This is a 552B model features a new causal Encoder–Decoder architecture, allowing it to have just 8B active parameters for input and 16B active parameters for output. On the first-party API, DeepSeek V4.1 Flash is priced at $0.30 per 1M input tokens and $1.20 per 1M output tokens, with cached input tokens priced at just $0.006 per 1M tokens, a 98% discount. Off-peak pricing provides a further 50% discount across input, cached input, and output tokens. V4.1 Flash is ~20% cheaper than DeepSeek V4 Flash 0731 and ~4x cheaper than DeepSeek V4 Pro 0813 while delivering a higher performance. Key results: ➤ DeepSeek V4.1 Flash makes gains in agentic capabilities and long context reasoning. DeepSeek V4.1 Flash scores 27% in Terminal-Bench v4.0, more than double DeepSeek…

Read the full post on X

Context

DeepSeek's V4.1 Flash overtakes V4 Pro 0813 as its flagship model, scoring 40 on the Intelligence Index, Artificial Analysis reports. The thread describes a 552 billion parameter model with a causal encoder-decoder design, using 8 billion active parameters for input and 16 billion for output, against 1.6 trillion parameters for V4 Pro. On DeepSeek's API it costs $0.30 per million input and $1.20 per million output tokens, about four times less per token than V4 Pro.

It reports 69% on AutomationBench-AA, a 657-task test of whether an completes business-app tasks in tools like Salesforce and Jira without tripping a guardrail, equal to GPT-6 Astra and above V4 Pro's 57%. The catch is verbosity. At about 89,000 tokens per Intelligence Index task, it uses more output tokens than even the frontier models, including Anthropic's Claude Opus 5, and Artificial Analysis says that verbosity is why it sits just short of the Intelligence versus Cost Pareto frontier, despite a cost of $0.27 per task. DeepSeek's own announcement emphasizes a smaller instead.

Terms in this piece · Glossary
  • token — The chunk of text a model reads and writes in — roughly three-quarters of a word — and the unit AI usage is billed in.
  • AI agent — An AI system that doesn't just answer once but works toward a goal in a loop — taking actions, reading the results, and deciding what to do next.
  • KV cache — The memory a model keeps about text it has already read, so generating each new token doesn't require reprocessing the whole conversation.
Key quotes

“DeepSeek V4.1 Flash is one of the most verbose models we’ve measured at 89k Tokens per Intelligence Index Task.”

“Despite such verbosity, DeepSeek V4.1 Flash still costs just $0.27 per Intelligence Index task.”

“DeepSeek V4.1 Flash takes first place on AutomationBench-AA with 69%, equal to GPT-6 Astra (69%) and slightly above Grok 4.6 (67%).”

More from ArtificialAnlys
Recommended reads
Comments

Checking sign-in…

Loading comments…