The day's most useful number was not a leaderboard rank but a denominator. Artificial Analysis showed Qwen3.8 Max cutting per-token price while more than doubling cost per Intelligence Index task, winning GDPval-AA Elo largely by taking roughly four times as many turns, and regressing ten points on AA-Omniscience as it stopped abstaining — three different ways a headline score can hide latency, spend and hallucination risk, arriving alongside an Intelligence Index grader change that makes cross-version comparisons invalid. Against that, Meta's Muse Spark 1.2 landed on the cost-per-task frontier and Kimi K3 reached GA in Copilot before a rollout pause, giving builders cheaper defaults whose real economics still need per-workload measurement. Underneath the model news, the plumbing converged: Agent Plugins shipped as a cross-client standard from OpenAI and Cursor on the same day the MCP ecosystem moved toward stateless servers and web-side tool interfaces.
Qwen3.8 Max is the cleanest case yet that per-token pricing misleads: cost per task more than doubled versus Qwen3.7 Max even as token rates fell, and its GDPval-AA Elo lead over Kimi K3 came from taking about four times more turns per task.
A dated brief from the vibe-coding frontier. Today’s Intel.
Artificial Analysis patched the Intelligence Index to v4.1.1 with changed graders, so scores cited across index versions are no longer directly comparable — check the version before quoting a number in a design doc.
Cheaper defaults arrived from two directions: Muse Spark 1.2 puts Meta on the cost-per-task Pareto frontier, while open-weight Kimi K3 hit GA in Copilot and is callable through the LangSmith gateway — though GitHub has paused the K3 rollout while it mitigates an incident.
Agent Plugins landed as an open standard from OpenAI Devs and Cursor on the same day, letting a skill or MCP configuration authored once ship to Codex, Cursor, Copilot and VS Code without per-client repackaging.
The agent-facing web stack advanced in parallel — a stateless next-generation MCP that runs as ordinary request handlers, WebMCP interfaces for existing sites, an agent-first browser in V8 isolates, managed retrieval via Cloudflare AI Search, and a protocol sketch for a readable, callable, payable agentic internet.
Two papers push back on summarization as a default context strategy: FinPerMA finds summarization drops the preference signal personalization depends on, with plain retrieval outperforming it, while SONAR argues code summaries for agent consumption should optimize correctness and abstraction level rather than polish.
OpenAI moved GPT-5.6 Sol behind all paid chats with a quantified factuality gain, expanded GPT-5.6 Luna to free users and removed free-tier chat limits, changing what you can assume about the model your users are actually on.