Vibeleaderboard
← All Intel
Intel / post

Solar Mini 4: cheap tokens, expensive tasks, weak agentic coding

Source
ArtificialAnlys
Date
ArtificialAnlys@ArtificialAnlys

Korean AI Lab 🇰🇷 Upstage has released Solar Mini 4 which scores 24 on the Artificial Analysis Intelligence Index, but costs ~5x as much per task as GPT-6 Luna (max) despite similar per-token prices Upstage has released Solar Mini 4, a new proprietary reasoning model. Upstage reports 35B total and 3B active parameters, setting a new Pareto optimal point on Intelligence Index vs. Active Parameters for models under 3B active parameters. It also scores 16 points higher than Upstage's previous-generation flagship, Solar Pro 3 (8), and cuts per-token pricing by a third to $0.10/$0.40 per 1M input/output tokens. Key results: ➤ Strong for its reported active parameter size: Solar Mini 4 scores 6 points higher than Qwen3.6 35B A3B (Reasoning), which has the same 3B active parameters, and 1 point higher than Nemotron 3 Ultra, which has 55B active. As a proprietary model, its size cannot be independently verified. ➤ Long context reasoning is a relative strength: It scores 83% on AA-LCR v1.1, matching MiniMax-M3 and GPT-6 Luna (max), and ahead of Gemini 3.8 Flash (high) and GPT-6 Astra (max) at 81%. It also scores 48% on SciCode, ahead of MiniMax-M3 and Inkling (xhigh) at 47%. ➤ Fast…

Read the full post on X
Key takeaways · AI-distilled
  • Upstage cut per-token pricing by a third to $0.10/$0.40 per 1M input/output tokens ($0.01 for cache hits), and Solar Mini 4 scores 16 points above its previous flagship, Solar Pro 3, according to Artificial Analysis.
  • Output speed and task time diverge: Solar Mini 4 generates 208 tokens/s, faster than GPT-6 Luna (max) at 152, but it averages 7.1 minutes per Intelligence Index task because it emits so many tokens.
  • Long- reasoning is its strongest area at 83% on AA-LCR v1.1, matching MiniMax-M3 and GPT-6 Luna (max), plus 48% on SciCode. Agentic coding is weak, with 1% on Terminal-Bench 4.0 and 22% on AutomationBench-AA.
  • It answers few knowledge questions correctly (18% accuracy, -11 on AA-Omniscience) but abstains on about half, giving a 64% non- rate versus 32% for Inkling (xhigh) and 23% for GPT-6 Luna (max).
  • Specs: 1M-token context, 262k max output, text only, February 2026 , weights not released. Artificial Analysis notes the 35B total / 3B active parameter count is Upstage's claim and cannot be independently verified.
Terms in this piece · Glossary
  • reasoning model — A model trained to think — generating extended internal reasoning before answering — trading time and tokens for accuracy on hard problems.
  • context window — The maximum amount of text a model can consider at once — its working memory for the current conversation or task.
  • hallucination — When a model states something false with full confidence — inventing facts, citations, or APIs that don't exist.
  • knowledge cutoff — The date after which a model saw no training data — everything later has to reach it through search, tools, or the context window.
Why it matters

Solar Mini 4 has cheap per-token pricing but uses 88k output tokens per task, so real cost per task is about 5x GPT-6 Luna. Check per-task cost, not token price, before choosing it.

More from ArtificialAnlys
Recommended reads
Comments

Checking sign-in…

Loading comments…