Solar Mini 4: cheap tokens, expensive tasks, weak agentic coding
- Source
- ArtificialAnlys
- Date
Korean AI Lab 🇰🇷 Upstage has released Solar Mini 4 which scores 24 on the Artificial Analysis Intelligence Index, but costs ~5x as much per task as GPT-6 Luna (max) despite similar per-token prices Upstage has released Solar Mini 4, a new proprietary reasoning model. Upstage reports 35B total and 3B active parameters, setting a new Pareto optimal point on Intelligence Index vs. Active Parameters for models under 3B active parameters. It also scores 16 points higher than Upstage's previous-generation flagship, Solar Pro 3 (8), and cuts per-token pricing by a third to $0.10/$0.40 per 1M input/output tokens. Key results: ➤ Strong for its reported active parameter size: Solar Mini 4 scores 6 points higher than Qwen3.6 35B A3B (Reasoning), which has the same 3B active parameters, and 1 point higher than Nemotron 3 Ultra, which has 55B active. As a proprietary model, its size cannot be independently verified. ➤ Long context reasoning is a relative strength: It scores 83% on AA-LCR v1.1, matching MiniMax-M3 and GPT-6 Luna (max), and ahead of Gemini 3.8 Flash (high) and GPT-6 Astra (max) at 81%. It also scores 48% on SciCode, ahead of MiniMax-M3 and Inkling (xhigh) at 47%. ➤ Fast…

- Upstage cut per-token pricing by a third to $0.10/$0.40 per 1M input/output tokens ($0.01 for cache hits), and Solar Mini 4 scores 16 points above its previous flagship, Solar Pro 3, according to Artificial Analysis.
- Output speed and task time diverge: Solar Mini 4 generates 208 tokens/s, faster than GPT-6 Luna (max) at 152, but it averages 7.1 minutes per Intelligence Index task because it emits so many tokens.
- Long- reasoning is its strongest area at 83% on AA-LCR v1.1, matching MiniMax-M3 and GPT-6 Luna (max), plus 48% on SciCode. Agentic coding is weak, with 1% on Terminal-Bench 4.0 and 22% on AutomationBench-AA.
- It answers few knowledge questions correctly (18% accuracy, -11 on AA-Omniscience) but abstains on about half, giving a 64% non- rate versus 32% for Inkling (xhigh) and 23% for GPT-6 Luna (max).
- Specs: 1M-token context, 262k max output, text only, February 2026 , weights not released. Artificial Analysis notes the 35B total / 3B active parameter count is Upstage's claim and cannot be independently verified.
- reasoning model — A model trained to think — generating extended internal reasoning before answering — trading time and tokens for accuracy on hard problems.
- context window — The maximum amount of text a model can consider at once — its working memory for the current conversation or task.
- hallucination — When a model states something false with full confidence — inventing facts, citations, or APIs that don't exist.
- knowledge cutoff — The date after which a model saw no training data — everything later has to reach it through search, tools, or the context window.
Solar Mini 4 has cheap per-token pricing but uses 88k output tokens per task, so real cost per task is about 5x GPT-6 Luna. Check per-task cost, not token price, before choosing it.
postSafety blocks hit defensive cyber work: CyberGym results and costs
postArtificial Analysis launches Cyber Index for AI cyber defense evaluation
postGPT-6.1 Sol reaches near-Astra intelligence at a quarter of the cost per task
postSonnet 5.5 nears Opus 5.5 on agentic evals but burns record output tokens
Checking sign-in…
Loading comments…

