Vibeleaderboard
← Back to Vibers
Artificial Analysis
Company

Artificial Analysis

Index Rank25

2 Tools · 99 Intel

Artificial Analysis is an independent AI benchmarking organization, founded in 2023 by Micah Hill-Smith and George Cameron, that runs its own evaluations of AI models and inference providers rather than relying on self-reported figures from AI labs. Its free public leaderboards and Intelligence Index have become widely referenced industry benchmarks for comparing model quality, speed, and cost. The company is backed by investors including Nat Friedman, Daniel Gross, and Andrew Ng.

Is this you? Sign in with X to claim this profile.

Tools

Artificial Analysis(artificialanalysis.ai)

Independent benchmarks comparing AI models and providers on intelligence, coding, speed, latency, and price.

Developer ToolsLLM Evalbuilt by @ArtificialAnlys
Optima(x.com)

Artificial Analysis's platform for creating and running custom benchmarks against AI models on your own tasks.

Developer ToolsLLM Evalbuilt by @ArtificialAnlys

Intel

Artificial Analysis puts Fable 5.1 at 66 on its Intelligence Index, the highest it has measured, with 59.1% on HLE and 91.4% on Terminal-Bench v2.1. Output token use of about 1.7x Fable 5 pushes cost per task to $3.76 even after the cache read cut saves roughly $1.40.

AI Toolsbuilt by @ArtificialAnlys

Artificial Analysis scores GLM-5.3-Flash at 57 on its Intelligence Index, three points behind GLM-5.3 at ~7.5x lower cost per task, priced at $0.15/$0.50 per 1M tokens with an 80% cached-input discount. It burns 149M output tokens per index run and hallucinates at 28% on AA-Omniscience.

AI Toolsbuilt by @ArtificialAnlys

Artificial Analysis scores Meta's Muse Spark 1.2 at 54 on its Intelligence Index, a large agentic-work improvement that ties Meta with SpaceXAI for third place.

Otherbuilt by @ArtificialAnlys

Artificial Analysis scores Alibaba's 2.4T-parameter Qwen3.8 Max at 56 on its Intelligence Index at $1.14 per task, still a point behind open-weights leader Kimi K3 at 25% lower cost.

Otherbuilt by @ArtificialAnlys

Artificial Analysis scores Gemini 3.7 Flash at 56 on its Intelligence Index, four points above 3.6 Flash, with an average time per task 40% faster than GPT-5.6 Terra, placing it on the intelligence-versus-speed Pareto frontier.

Otherbuilt by @ArtificialAnlys

Artificial Analysis scores DeepSeek V4 Pro 0813 at 53 on its Intelligence Index, eight points above April's version but with a 3.6x price rise, leaving it barely ahead of the cheaper V4 Flash and still behind Kimi K3 among open-weights models.

AI Toolsbuilt by @ArtificialAnlys

Artificial Analysis ships Intelligence Index v4.1.1, upgrading its grader models and adopting the latest τ³-Banking version.

Otherbuilt by @ArtificialAnlys

Muse Glimmer scores 35 on the Artificial Analysis Intelligence Index, 21 points above Llama 4 Maverick and Meta's first open-weights release in 16 months, sitting alongside Kimi K2.5 and just behind Qwen3.6 27B.

AI Toolsbuilt by @ArtificialAnlys

DeepSeek V4.1 Flash, a 552B causal Encoder-Decoder model with only 8B/16B active parameters, beats the 1.6T-parameter V4 Pro on Artificial Analysis's Intelligence Index while costing about 4x less per token.

AI Toolsbuilt by @ArtificialAnlys

A new index scores seven search API providers on agent task quality, cost and speed, holding model and harness fixed and averaging three benchmarks covering broad research, multi-hop browsing and factual recall.

Developer Toolsbuilt by @ArtificialAnlys

Artificial Analysis's Coding Agent Index v1.4 now zeroes any Terminal-Bench v2.1 pass identified as reward hacking — fetching published solutions instead of doing the work. Because the tasks run with open internet and never forbid lookup, rates differ sharply between agents and between models.

AI Agentsbuilt by @ArtificialAnlys

Artificial Analysis and Liquid AI publish phone-scale rankings: on an iPhone 17 Pro, Nanbeige4.2-3B and LFM2.5-2.6B tie at 63 across BFCL, IFBench, GPQA Diamond, MATH-500 and AA-Omniscience, while generation time spans 30x and peak memory 19x across the field.

AI Toolsbuilt by @ArtificialAnlys

Artificial Analysis scores Nemotron 3.5 Lightning at 24 on its Intelligence Index, matching gpt-oss-120b at about a quarter of the total parameters, with 31.6B total and 3.6B active on a hybrid Mamba-Transformer architecture.

AI Toolsbuilt by @ArtificialAnlys

OpenAI's new full-duplex speech-to-speech model GPT-Live-1 tops Artificial Analysis' Speech-to-Speech Index at 81.5, ahead of Grok Voice Think Fast 2.0, by delegating reasoning and tool use to a separately configured backend text model.

AI Toolsbuilt by @ArtificialAnlys

fal post-trained MiniMax H3 into H3 Max, which now leads Artificial Analysis' image-to-video-with-audio board and sits third in text-to-video. It produces 5–15 second clips with native audio up to 768p at $0.04 per second, under the base H3 endpoint, with weights promised.

AI Toolsbuilt by @ArtificialAnlys

Grok 4.6 makes large gains on AA-Briefcase, a private long-horizon agentic knowledge-work benchmark, landing neck and neck with Claude Fable 5 on overlapping confidence intervals while costing substantially less.

AI Agentsbuilt by @ArtificialAnlys

Artificial Analysis launches AA-AnalystAgent, an agentic benchmark on real spreadsheets and documents requiring models to interpret sources, judge which caveats apply and settle on a methodology. Claude Opus 5 leads at 54%, ahead of GPT-5.5 at 50%.

AI Agentsbuilt by @ArtificialAnlys

Artificial Analysis puts Cartesia's Sonic 3.6 at number one on both the Provider Voice and Controlled Voice speech arenas, reporting Elo scores and appearance counts alongside per-character pricing and generation speed against ElevenLabs, Speechify and Alibaba entries.

AI Toolsbuilt by @ArtificialAnlys

BreezeBlue's Breeze TTS 2 takes the top open-weights slot in Artificial Analysis's Provider Voices arena at 1,215 Elo, 90 ahead of Fish Audio S2 Pro, covering 50 languages with prompt-generated voices — while running 45 chars/sec against Fish's 102 and costing $34 per 1M characters hosted.

AI Toolsbuilt by @ArtificialAnlys

Artificial Analysis places four Korean labs above 30 on its Intelligence Index. Motif 3 (314B total, 13B active, open weights) scores 47 and Upstage's Solar Pro 4 scores 42 — the top-scoring models developed outside the US and China — with SK Telecom and LG AI Research following on open weights.

AI Toolsbuilt by @ArtificialAnlys

Muse Voice Transcribe from Meta Superintelligence Labs reaches 3.1% WER at 0.16s after end of speech, ahead of Cartesia Ink-2 and ElevenLabs Scribe v2 Realtime, for $0.18 per audio hour. It works on 80ms chunks, covers 70 plus languages and handles hour-plus audio without post-processing.

AI Toolsbuilt by @ArtificialAnlys

Artificial Analysis finds only two labs on the intelligence-versus-speed Pareto frontier: every frontier model finishing under two minutes is OpenAI's, and four of the six slower ones are Anthropic's.

Otherbuilt by @ArtificialAnlys

Artificial Analysis scores Upstage's Solar Pro 4 at 42 on its Intelligence Index, a 27-point jump over Solar Pro 3, alongside Inkling and just behind MiMo-V2.5-Pro, with roughly doubled pricing.

AI Toolsbuilt by @ArtificialAnlys

Artificial Analysis clocked Gemma 4 31B at roughly 3,400 tokens/s on a private NVIDIA Groq 3 LPX endpoint, holding that speed from 10k to 100k input tokens across 50 sequential single-concurrency requests. The rack is in full-scale production and enters operation later this year.

AI Toolsbuilt by @ArtificialAnlys

Ant Group's Ling 3.0 Flash, 124B total and 5B active parameters under MIT, scores 38 on the Artificial Analysis Intelligence Index, 24 points over its predecessor and on the open-weights Pareto frontier for intelligence versus total parameters.

Otherbuilt by @ArtificialAnlys

A new quality-tier image model takes first place on an independent image editing leaderboard and seventh in text to image, at roughly twice the per-image cost of its own mid tier and five times the fast tier.

AI Toolsbuilt by @ArtificialAnlys

Artificial Analysis' arena run puts Microsoft's MAI-Image-2.6-Preview first in image editing — ahead of its own 2.5-Pro, Reve 2.1 and GPT Image 2 — and second in text-to-image, leading 5 of 19 category boards including Retail & Ecommerce and Marketing. Available in MAI Playground and Foundry private

AI Toolsbuilt by @ArtificialAnlys

Artificial Analysis places Muse Spark 1.2 on the cost-per-task Pareto frontier: six points below Claude Opus 5 at roughly a sixth of the cost, and comparable to Opus 4.8 at a fifth.

Otherbuilt by @ArtificialAnlys

MBZUAI's 375B total, 23B active open weights MoE scores 47 on the Intelligence Index, thirty points above its predecessor. It beats MiniMax-M3 on GDPval-AA and tau-Banking but trails on GPQA Diamond and Humanity's Last Exam, and answers only 40% of AA-Omniscience questions.

AI Toolsbuilt by @ArtificialAnlys

Artificial Analysis benchmarked Perplexity's Search API at three context sizes inside a fixed harness. Medium scores 80 on its Search Index against 75 for Parallel and Brave, with the lowest inference cost per task ($0.028–$0.034) and ~$0.091 total. Quality plateaus between medium and high while sea

AI Toolsbuilt by @ArtificialAnlys

xAI's Grok Voice Transcribe 2.0 replaces its 1.0 speech-to-text model, cutting streaming WER from 3.9% to 2.7% and non-streaming WER to 2.3%, per Artificial Analysis benchmarks, at unchanged API pricing.

AI Toolsbuilt by @ArtificialAnlys

Ant Group's Ling 3.0 Tiny scores 25 on the Artificial Analysis Intelligence Index with 7.9B total and 1.3B active parameters, comparable to gpt-oss-120b at 15x fewer total parameters, though it burned 213M output tokens to get there.

AI Toolsbuilt by @ArtificialAnlys

Held out medical long context tiers, graded by panel, with Claude Fable 5 in front.

AI ToolsFreeOtherbuilt by @ArtificialAnlys

Artificial Analysis puts Agnes 2.5 Pro Beta at 49 on its Intelligence Index, with the Agentic Index leaping 25→44 and τ³-Banking tripling to 36%. Cost of the jump: 50k output tokens per task, and an Omniscience improvement driven by attempting only 45% of questions.

AI Toolsbuilt by @ArtificialAnlys

Human run speech model arena that scores preference and whether the tool call landed.

AI ToolsFreeVoice & Audiobuilt by @ArtificialAnlys

Artificial Analysis rebuilds its Text to Image Arena around 10 real-world use cases and 9 capabilities, from UI mockups to live-action film, with monthly-refreshed human-curated prompts. GPT Image 2 leads everywhere, with Nano Banana 2 the value pick.

Otherbuilt by @ArtificialAnlys

Google's two new transcription APIs, measured: 2.6% word error rate non-streaming at roughly 84x realtime for about $5 per 1,000 minutes, and a Live variant at 4.0% WER with first final transcripts 0.40s after speech ends for about $9. Both cover 85+ languages.

AI Toolsbuilt by @ArtificialAnlys

Round 2 of Korea's government-backed foundation-model program cut the field to three: Upstage's Solar Open 250B (Intelligence Index 37), SK Telecom's A.X K2 (35), and LG's K-EXAONE 2.0 (31). Each survivor is set to receive roughly 1,000 NVIDIA B200s for six months.

AI Toolsbuilt by @ArtificialAnlys

Artificial Analysis has opened Optima, which lets teams run their own tasks as a benchmark and compare models on quality, speed and cost efficiency rather than relying on a general index. A new video walks through building and running one end to end.

AI Toolsbuilt by @ArtificialAnlys

Artificial Analysis benchmarked Ant Group's Ling-3.0-flash-VL, a 124B-parameter (5.5B active) open-weights vision-language model, finding it leads comparable models on their Intelligence Index but scores 0% on Terminal-Bench v4.0 and 16% on AutomationBench-AA.

AI Toolsbuilt by @ArtificialAnlys

Artificial Analysis on Muse Spark 1.2's cost efficiency: $0.40 per Intelligence Index task, with only Grok 4.5 and GPT-5.6 Sol at medium effort cheaper in its intelligence cluster, the rise from 1.1 driven by token usage rather than pricing.

Otherbuilt by @ArtificialAnlys

Artificial Analysis breaks down Muse Spark 1.2's 3-point Index gain: concentrated in agentic evaluations, led by a 260-Elo jump on GDPval-AA v2, with small regressions on SciCode and Humanity's Last Exam.

Otherbuilt by @ArtificialAnlys

The same open weights served by different providers do not give you the same product. Self hosting the released weights as a 100 percent reference and scoring each provider endpoint against it with three evaluations puts them between 73 and 100 percent.

AI Toolsbuilt by @ArtificialAnlys

Benchmark results place two search providers on the cost-quality frontier for agent tasks, with one scoring 73 at 7.5 cents per task but the slowest completion time, and another reaching 75 for 12 percent more.

Developer Toolsbuilt by @ArtificialAnlys

Artificial Analysis scores Qwen3.8 Max at 1739 Elo on GDPval-AA, ahead of Kimi K3 and tied with Fable 5 and GPT-5.6 Sol, with the 468-Elo gain partly driven by taking 64 turns per task against its predecessor's 14.

Otherbuilt by @ArtificialAnlys

Serving the same model does not mean serving it the same way. Endpoints that underperform the reference generate markedly fewer output tokens per task — the weakest around half — so capped output length and reduced reasoning effort surface directly in token accounting.

AI Toolsbuilt by @ArtificialAnlys

Qwen3.8 Max costs $1.14 per Intelligence Index task, double its predecessor, despite lower per-token prices: turns per task on GDPval-AA rose from 14 to 64, so the agentic verbosity, not the pricing, drives the bill.

Otherbuilt by @ArtificialAnlys

Artificial Analysis finds Qwen3.8 Max regressing 10 points on AA-Omniscience versus its predecessor: accuracy flat at ~31% while hallucination rises from 23% to 40%, attempting more questions it cannot answer.

Otherbuilt by @ArtificialAnlys

Muse Spark 1.2 ranks fifth on GDPval-AA v2 at 1631 Elo, behind Claude Opus 5, GPT-5.6 Sol and Kimi K3 but a 260-point jump over Muse Spark 1.1's launch score.

Otherbuilt by @ArtificialAnlys

Muse Glimmer's weaknesses concentrate in agentic evaluation: 953 Elo on GDPval-AA v2 against 1141 for Qwen3.6 27B, and 52% on Terminal-Bench v2.1 against 61%, with a hallucination rate of 82% on AA-Omniscience against 49%.

AI Toolsbuilt by @ArtificialAnlys

Artificial Analysis on Muse Spark 1.2's AA-Omniscience result: the score rose 18 to 22 driven by abstention again, hallucination rate down 10 points while accuracy slipped from 41% to 38%.

Otherbuilt by @ArtificialAnlys

Artificial Analysis's numbers on Upstage's Solar Pro 4 show its AA-Omniscience score climbing from -53 to -1 almost entirely through refusal: attempts drop from 92% to 41% of questions and hallucinations from 88% to 24%, while accuracy sits unchanged around 19%.

AI Toolsbuilt by @ArtificialAnlys

NVIDIA's first Nemotron 3.5 model lands as a small open-weights release with a genuine agentic jump: 824 Elo on GDPval-AA v2, ahead of Nemotron 3 Super and gpt-oss-120b, and more than triple Nemotron 3 Nano on Terminal-Bench v2.1.

AI Toolsbuilt by @ArtificialAnlys

Ling 3.0 Flash improves 48 points on AA-Omniscience over Ling 2.6, driven primarily by cutting the hallucination rate from 97% to 44%.

Otherbuilt by @ArtificialAnlys

A classification of 1,567 failing AA-AnalystAgent attempts across ten models into seven failure modes. Anchoring on a wrong early hypothesis is the most widespread at 57%, with models committing to a source early and defending it for the rest of the run.

AI Agentsbuilt by @ArtificialAnlys

Artificial Analysis updated its occupation-mapped Capability Indices to v1.1, adding Agentic Tool Use benchmarks from AutomationBench-AA and long-context tasks across Finance, Legal, Healthcare, Strategy, Engineering and Economics indices.

AI Toolsbuilt by @ArtificialAnlys

Artificial Analysis launches a Cyber Index and partner Alliance combining CWE-Bench-AA, DeepsecBench-AA and CyberGym-E2E-AA to score how well AI agents audit, discover, patch, and reproduce vulnerabilities.

Cybersecuritybuilt by @ArtificialAnlys

Artificial Analysis benchmarks Claude Sonnet 5.5: 56 on its Intelligence Index, parity with Opus 5.5 on several agentic evals, but the heaviest output token use it has measured, raising cost per task to $7.60.

AI Toolsbuilt by @ArtificialAnlys

Artificial Analysis scores Gemini 3.8 Flash at 59 on its Intelligence Index with high reasoning, three points above 3.7 Flash, at $0.58 per task and 2.5 minutes per task. Output rose 30% to 48k tokens per task, and AA-Briefcase Elo gained 79 points on rubric and analytical quality.

AI Toolsbuilt by @ArtificialAnlys

Artificial Analysis breaks down how Anthropic's 20% price cut and cheaper cache reads on Opus 5.5 nearly offset an ~80% jump in token usage per task, leaving cost per Intelligence Index task close to Opus 5's.

AI Agentsbuilt by @ArtificialAnlys

Alibaba's Wan 3.0 enters public preview on Model Studio as a single system for generating and instruction-editing video, accepting text, image, video, audio and document references. Artificial Analysis measures it first in video editing with audio, up from fifth for Wan 2.7.

AI Toolsbuilt by @ArtificialAnlys

Artificial Analysis ranks Alibaba's Qwen-Image-3.0-Pro sixth for image editing and ninth for text to image, gains of 83 and 48 Elo over the previous Pro generation, at four cents per 1K image through Alibaba Cloud.

Otherbuilt by @ArtificialAnlys

Google DeepMind's Gemini 3.8 Live, successor to Gemini 3.1 Flash Live, tops Artificial Analysis's speech-to-speech benchmarks in both standard and Extended Thinking (configurable reasoning effort) variants.

AI Toolsbuilt by @ArtificialAnlys

Artificial Analysis finds MiMo-V2.6-Pro, Claude Opus 5.5, GPT-6 Luna, and GPT-6 Sol together set eleven new points on the Intelligence Index cost/performance frontier, with Opus 5.5 now the top scorer at 58 for $5.98 per task.

AI Agentsbuilt by @ArtificialAnlys

On the Artificial Analysis Coding Agent Index, Muse Spark 1.3 (xhigh) in the Muse Code harness scores 64 at $1.72 per task, the cheapest agent above 60. The partner-only max variant reaches 68, level with Claude Opus 5 (xhigh) while burning about a third fewer tokens per task.

AI Agentsbuilt by @ArtificialAnlys

Artificial Analysis reports that Claude Fable 5.1, Muse Spark 1.3, and GPT-6 Astra each pushed the Intelligence Index vs. cost Pareto frontier outward last week, each now offering more capability per dollar than any prior model at its price point.

AI Toolsbuilt by @ArtificialAnlys

Artificial Analysis's AA-Briefcase-Lite results show Grok 4.7 gaining sharply in Analytical Quality Elo (1698 to 1994) over Grok 4.6, trailing only Anthropic's models, while costing about half of Opus 5's per-task cost on market-analysis and deal-assessment work.

AI Toolsbuilt by @ArtificialAnlys

Artificial Analysis benchmarks show GPT-6 Sol and Luna roughly halving cost per task versus GPT-5.6 while intelligence scores stay flat, with Sol improving and Luna regressing on the Coding Agent Index.

AI Toolsbuilt by @ArtificialAnlys

Artificial Analysis rebuilt its Image Editing Arena around seven editing actions, from enhancement and restoration to identity-preserving edits and UI/UX work, using multi-step edit instructions since single-instruction edits are saturated. Winners differ sharply by category.

AI Toolsbuilt by @ArtificialAnlys

Xiaomi's newly released MiMo-V2.6-Pro, a 1.02T-parameter MoE model, becomes the top-scoring open-weights model on Artificial Analysis's Intelligence Index while undercutting rivals on price, landing on the cost-efficiency frontier.

AI Toolsbuilt by @ArtificialAnlys

Artificial Analysis ranks new coding agents on its Coding Agent Index. Claude Sonnet 5.5 in Claude Code leads at 68 but costs $14.19 per task, while GPT-6.1 Sol in Codex scores 63 at $1.04.

AI Toolsbuilt by @ArtificialAnlys

Artificial Analysis's Coding Agent Index v1.5 now reports safety refusal rates, finding Claude Fable 5.1 had the highest fallback rates in Claude Code (8.8%) and Devin Fusion (7.1%), meaning index scores partly reflect fallback-model performance.

AI Agentsbuilt by @ArtificialAnlys

Microsoft's new speech model reaches 2.0% word error rate on the AA-WER leaderboard, second only to Alibaba's Fun-Realtime-ASR-preview, at roughly 411x real time and $1.67 per 1,000 minutes. Coverage expands from 43 to 60 languages with diarization and word timestamps.

AI Toolsbuilt by @ArtificialAnlys

Artificial Analysis benchmarks inclusionAI's 6B-parameter Ming-Image-0.1-Design as the top open-weights text-to-image model for UI/UX design, especially layout and text rendering, despite ranking only mid-pack on overall image quality.

AI Toolsbuilt by @ArtificialAnlys

Artificial Analysis measures GPT-6.1 Sol: near-Astra Intelligence Index score, $0.72 per task at max effort versus $3.26 for Astra, a 12 point Terminal-Bench 4.0 gain, and lower hallucination rate.

AI Toolsbuilt by @ArtificialAnlys

Artificial Analysis benchmarked Apodex 1.1 at 44 on its Intelligence Index: 1348 Elo on GDPval-AA v2 and 70% on TerminalBench v2.1, ahead of Kimi K2.6 on agentic work but weaker on knowledge reliability and academic reasoning, at roughly 17k output tokens per task.

AI Toolsbuilt by @ArtificialAnlys

Artificial Analysis benchmarked Octen Search using its Stirrup agent harness: it scores 77 on the Search Index (third place), leads on speed at 16.9s per task and 0.2s per query, and is among the cheapest providers, driven by strong BrowseComp results.

AI Toolsbuilt by @ArtificialAnlys

Artificial Analysis extends its Controlled Voice Arena to 9 new languages, finding Cartesia's Sonic models lead most non-English leaderboards while Inworld's Realtime TTS-2 tops Mandarin.

AI Toolsbuilt by @ArtificialAnlys

Artificial Analysis measured GPT-6 Astra across its Coding Agent and Intelligence indices: 67 in Codex, level with Claude Opus 5 and Fable 5, on roughly a third of GPT-5.6 Sol's tokens. Prices rose 2.5x to $10/$50 per million, so the efficiency gain nets out only in agentic coding work.

AI Toolsbuilt by @ArtificialAnlys

Artificial Analysis's independent benchmarking shows Claude Opus 5.5 topping its Intelligence Index at 58, matching GPT-6 Astra on Terminal-Bench 4.0, and leading agentic knowledge-work evals, alongside a 20% price cut and cheaper cache reads.

AI Toolsbuilt by @ArtificialAnlys

Artificial Analysis's CyberGym-E2E-AA results show several frontier models safety-blocking most defensive memory-safety tasks, while cheaper capable models like GPT-6 Luna and MiMo-V2.6-Pro handle large-codebase bug hunts at a fraction of the cost.

Cybersecuritybuilt by @ArtificialAnlys

Meta's fourth Muse Spark release in five months scores 61 (xhigh) and 62 (max, partner preview) on the Artificial Analysis Intelligence Index. Gains sit almost entirely in agentic and tool-use evaluations, with small regressions on long-context retrieval and factual accuracy.

AI Toolsbuilt by @ArtificialAnlys

Artificial Analysis benchmarks put xAI's Grok 4.7 among the top agentic coding models, gaining +9 points on the Coding Agent Index and ranking 4th behind Claude Fable 5.1, GPT-6 Astra, and Claude Opus 5, though with higher token usage per task.

AI Toolsbuilt by @ArtificialAnlys

Ant Group released Ling-3.0-flash-Fin, a finance-tuned open-weight model scoring 23 on Artificial Analysis's Intelligence Index and 24 on Finance & Accounting, matching MiniMax-M2.7 with roughly half the active parameters (5.1B vs 10B).

AI Toolsbuilt by @ArtificialAnlys

Google DeepMind releases Gemini 3.8 Flash TTS and Flash-Lite TTS, debuting at #2 and #6 on Artificial Analysis's Provider Voice Arena and #1 on pronunciation robustness, processing 44.1 and 40.2 characters per second respectively.

AI Toolsbuilt by @ArtificialAnlys

Artificial Analysis reports GPT-6.1 Sol extends the cost-efficiency Pareto frontier, at 31% below GPT-6 Sol per task. An image-encoding fix also lifted GPT-6 Luna by one Intelligence Index point, mostly on vision-heavy evals.

AI Toolsbuilt by @ArtificialAnlys

Artificial Analysis reports Claude Opus 5.5 setting a new high score of 66 on its Coding Agent Index via gains on Terminal-Bench, DeepSWE, and SWE-Atlas-QnA, though cost per task rises 21% to $13.04 from higher token usage.

AI Agentsbuilt by @ArtificialAnlys

StepFun's 600B-parameter Step 5 Preview scores 44 on Artificial Analysis's Intelligence Index, tying Kimi K3 (max) at about 2.8x lower cost per task, with large reasoning gains on HLE and CritPt but weaker results on agentic evaluations than similarly scored peers.

AI Toolsbuilt by @ArtificialAnlys

StepFun's StepAudio 3 ASR tops Artificial Analysis's non-streaming speech-to-text leaderboard at 1.7% word error rate, though it transcribes slower and costs more per minute than several rivals.

AI Toolsbuilt by @ArtificialAnlys

Artificial Analysis benchmarks Upstage's Solar Mini 4, a 35B/3B-active proprietary reasoning model. It scores 24 on the Intelligence Index with strong long-context results, but heavy token use inflates per-task cost and agentic coding is weak.

AI Toolsbuilt by @ArtificialAnlys

Artificial Analysis ranks Meta Superintelligence Labs' first image model #4 in image editing and #5 in text to image, on the price/quality Pareto frontier. Muse Image invokes search and coding tools and self-refines, and reaches developers via the Meta Model API, fal, Runway, and OpenRouter.

AI Toolsbuilt by @ArtificialAnlys

Inworld's Realtime TTS-2 takes first on Artificial Analysis's Controlled Voice Arena at 1,123 Elo, narrowly ahead of Cartesia Sonic 3.6, and second on the Provider Voice Arena. It handles 100 plus languages, inline delivery direction, and 106 characters per second of generation.

AI Toolsbuilt by @ArtificialAnlys

Artificial Analysis's new Pronunciation Robustness benchmark tests whether TTS models correctly voice context-dependent words, shorthand, and exact sequences like emails or codes; Gemini 3.1 Flash TTS leads at 88.1%, ahead of SpaceXAI TTS (87.6%) and ElevenLabs Eleven v3 (85.6%).

AI Toolsbuilt by @ArtificialAnlys

New dashboard pages let you compare intelligence, cost, and latency across every reasoning/effort configuration of a model, since frontier releases now ship with up to six effort tiers with very different cost-performance profiles.

AI Toolsbuilt by @ArtificialAnlys

Artificial Analysis launches Terminal-Bench-Science 0.1, a 70-task agentic benchmark spanning five research domains; only GPT-6 Astra (63%) and Claude Opus 5.5 (62%) clear 50%, with life sciences the hardest domain and open models far behind.

AI Agentsbuilt by @ArtificialAnlys

Artificial Analysis benchmarked OpenAI's new GPT Image 2.5 Flare and Sunburst, finding they take the top two spots on both the Text-to-Image and Image-Editing leaderboards at GPT Image 2 pricing, with the biggest gains in composition/framing and text/symbol edits.

AI Toolsbuilt by @ArtificialAnlys

Artificial Analysis shows GPT-6 Sol and Luna matching predecessor Intelligence Index scores at roughly half the cost, driven by about 50% lower token pricing.

AI Agentsbuilt by @ArtificialAnlys

Artificial Analysis benchmarked Cognition's Devin Fusion, which pairs a frontier model with a cheaper SWE-2 sidekick model, finding the GPT-6 Astra pairing costs 43% less and runs 31% faster than the Claude Fable 5.1 pairing while scoring close behind it.

Developer Toolsbuilt by @ArtificialAnlys

Artificial Analysis finds Gemini 3.8 Flash TTS leads its Pronunciation Robustness benchmark at 89.5%, ahead of prior Gemini and SpaceXAI models, with particular strength on contextual disambiguation and shorthand expansion.

AI Toolsbuilt by @ArtificialAnlys