Vibeleaderboard
Index — Latest Intelligence

Intel

Page 02

Together Link: open models in the harness you already use. Start with one command today.

Together AI's Link connects Claude Code, Claude Desktop, Codex, OpenCode and Pi to open models hosted on Together via a one-line install.

Why it mattersLets teams point existing coding agent harnesses at open models such as Kimi K3 and GLM 5.3 with one command and revert easily. Relevant for cutting coding-agent spend without changing workflow.

articlewww.together.ai

Qwen3.8 27B addition in words

Simon Willison reruns a GPT-4o-era experiment on local Qwen3.8-27B, testing integer addition expressed in English words across 5,070 cases.

Why it mattersShows where a local 27B model's arithmetic breaks down, and how reasoning mode changes it, with exact accuracy figures by operand length. Useful when deciding if a local model needs a calculator tool.

articlesimonwillison.net

Agents That Write Their Own Tools at Runtime — Sandhya Subramani, AWS

AWS's Sandhya Subramani demonstrates meta-tooling in Strands Agents, where agents write, load and use their own tools and sub-agents at runtime.

Why it mattersShows that a system prompt plus editor, shell and load_tool lets an agent author tools and sub-agents at runtime. It also lists the evals and sandboxing needed before trusting a self-modifying agent.

videoAI Engineer

AI Coding Agents Are Breaking Big Codebases — Dan Adler, Sourcegraph

Sourcegraph's CEO makes the case that large-scale code visibility is infrastructure for coding agents, describing codebase decay from AI output and introducing Agentic Batch Changes for cross-repo edits.

Why it mattersArgues that agents cannot act on code they cannot search, and that AI-generated code accelerates duplication and drift. Sourcegraph's Agentic Batch Changes targets multi-repo changes from a single prompt.

videoAI Engineer

Keyword Search Is Dying. Is Your Catalog Ready for AI Agents? — PayPal

PayPal's agentic commerce lead explains why keyword-optimized catalogs fail for agents and shares experiment results on product data enrichment, hybrid search.

Why it mattersPayPal's experiments found that enriching merchant product data improves agent recommendations, but unstructured boilerplate dilutes the signal and increases hallucination. Useful if you are making catalogs or retrieval readable to agents.

videoAI Engineer

The Human Is an Async API — Melanie Warrick, Temporal

Temporal's Melanie Warrick shows how to build human-in-the-loop agents with durable execution, using wait conditions and signals with Google ADK and LangGraph.

Why it mattersA blocking ask-the-human call breaks in production. Durable workflows with a wait condition and signal let an agent pause for minutes or weeks and resume after a worker crash without losing state.

videoAI Engineer

We Built an AI Support Agent That Resolves 80% of Tickets — AssemblyAI

AssemblyAI's Joey support agent, built on the Claude Agent SDK with checked-out docs, embeddings, agentic search and one CLAUDE.md, resolves 80% of tickets for about $700 a month, then gains voice via a websocket.

Why it mattersGives a concrete architecture and cost figure for a support agent, including turning escalations into roadmap items and a deploy loop that ships fixes in about 30 seconds.

videoAI Engineer
Lenny Rachitsky

OpenAI’s Head of ChatGPT: We’re entering a new era of AI (again) | Tibo Sottiaux

podcast

Dashboards Are Dead — Sarah Simionescu, Composio

Composio's Sarah Simionescu argues dashboards are obsolete, explains why a dozen MCP servers alone fail.

Why it mattersNames concrete reasons raw MCP wiring breaks down and one approach, tool search with execution plans, for agents working across apps with less context load.

videoAI Engineer

Stop Fine-Tuning to Fix Retrieval Problems — Anant Srivastava

Oracle's Anant Srivastava argues prompt, memory and weights are three tools for three jobs, not a ladder.

Why it mattersGives a diagnostic for choosing between prompt edits, retrieval and fine-tuning. It explains why fine-tuning to fix retrieval problems backfires.

videoAI Engineer

We're going to need default hard budget caps on pretty much everything

Simon Willison argues that usage-billed APIs and hosts need hard, default-on spending caps because coding agents lower the friction of spinning up costly services.

Why it mattersAgents can create paid resources faster than anyone watches them. The argument gives a concrete design rule (hard caps by default, opt-out with explicit consent) to apply to your own services and agent budgets.

The Agent Said It Was Done. The Database Disagreed.

Microsoft's ThinkingBox benchmark grades agents on the records they leave behind across 507 stateful MCP workflows, repeated 20 times to measure consistency and its cost.

Why it mattersAgents can make well-formed tool calls and still leave the database wrong. Grading terminal state across 20 repeated runs measures reliability, which tool-call and final-answer checks miss.

articlehuggingface.co

What Makes Open Models Fast in Production — Sujee Maniyam, Nebius

Nebius engineers cover production serving of open LLMs.

Why it mattersLists the specific serving layers that determine cost and latency for self-hosted open models, which helps when weighing a closed API against running open weights.

videoAI Engineer

An Interaction Is All You Need — Ivan Leo, Google DeepMind

Google DeepMind presents the Gemini Interactions API and Managed Agents.

Why it mattersGemini's Interactions API moves conversation state server-side and Managed Agents run in persistent remote sandboxes with credential injection. Agent builders can drop manual thought-signature handling and hosted-harness plumbing.

videoAI Engineer

Automating my 35mm film scanning pipeline

A photographer lets Claude automate a 35mm scanning pipeline, then recounts how it damaged the scanner by misdriving the carriage, how the fault was traced.

Why it mattersClaude drove a scanner through SANE and nearly broke it after a blind reverse move at the end of each scan. A concrete case for sandboxing and confirming hardware actions before granting agents device access.

articleshannadige

Agents don't need memory, they need documentation

An essay arguing that RAG-based agent memory plugins are structurally flawed because similarity search can't tell current from stale information.

Why it mattersThe piece argues RAG-style memory plugins fail because similarity search can't distinguish correct from stale information.

I Built a Personal AI Agent on a Raspberry Pi — Jeremy Adams, Neo4j

A live build of a wearable personal agent on a Raspberry Pi 4B.

Why it mattersDemonstrates graph-based agent memory (POLE schema) over MCP on cheap, isolated hardware, with offline capture that syncs later. Gives builders a concrete pattern for persistent memory beyond chat history.

videoAI Engineer

How VS Code Went from Monthly to Weekly Releases with AI — Harald Kirschner

A Microsoft VS Code product lead details how the team moved to weekly releases with agents.

Why it mattersShows how a 50M-user product restructured CI, review, triage and evals around coding agents, with measurable results such as code survival rising from 55% to 86%. Useful as a template for agent-ready engineering workflows.

videoAI Engineer

Why Specialized AI Could Beat The God Model

a16z talks with OpenRouter's Alex Atallah and Replit's Amjad Masad about specialized models over a single god model.

Why it mattersArgues that routing and combining specialized models can match frontier performance at lower cost and reduce single-provider dependence. Relevant when designing model selection and agent architecture.

videoa16z

Crusoe CEO: Why Everyone Gets GPU Depreciation & AI Energy Costs Wrong

Crusoe's CEO discusses AI data center bottlenecks.

Why it mattersCompute supply, energy limits and GPU lifespan assumptions shape future inference cost and availability. The talk offers an operator's view on those constraints.

video20VC with Harry Stebbings

Germany's new sovereign AI model Kolibri

Walkthrough of Aleph Alpha's Kolibri, an open-weight 78B mixture-of-experts model (3.5B active per token) for German and English under Apache 2.0.

Why it mattersA new Apache 2.0 open-weight MoE with 3.5B active parameters and 262k context is a self-hostable option for German and English workloads with data-residency needs. The post covers when it fits and where it falls short.

articletejaskumar__

Kolibri: A Sovereign Open-Weight Model

Aleph Alpha releases Kolibri, an English-German MoE model (78B total, 3B active, up to 1M context) with Apache 2.0 weights, built through a reusable training pipeline for sovereign, regulated deployments.

Why it mattersAn Apache 2.0 English-German MoE with only 3B active parameters and 1M context can run on-premise for regulated work. It gives teams with sovereignty or compliance needs another open-weight option.

articlebastitx

80s Choplifter game ported to Vision Pro as an agent eval (WIP)

A developer building a Vision Pro Choplifter-style game uses Astra and Claude Opus 5.5 with Blender MCP, assigning them different roles and cross-checking output.

Why it mattersDescribes a working split between two coding agents (planning and review versus focused implementation) and the handoff cost. Notes that comfort testing in VR still needs a human.

articleboundsj

Your LLM App Returned 200 OK. It Was Still Wrong. — Marina Petzel, Datadog

Datadog's Marina Petzel explains why latency and error golden signals miss GenAI failures.

Why it mattersAn LLM app can return 200 OK and still be wrong or expensive. This lists the cost, safety and quality signals to add to standard monitoring.

videoAI Engineer

YOLO Mode, Safely: MicroVM Sandboxes for Any Agent — Rowan Christmas, Docker

Docker's Rowan Christmas shows a coding agent finding browser history and bank data in five prompts, then demos Docker Sandboxes.

Why it mattersIt shows concretely why running agents in YOLO mode on a laptop is risky, and how a microVM sandbox gives isolation that works with Claude Code, Codex or other agents.

videoAI Engineer
OpenAIDevs@OpenAIDevs

OpenAI's Agents API gained computer use that spins up a browser and agent in a single call, hosted environment sizes, a Bedrock Managed Agents path on AWS, and GPT-6.1 Sol support, according to the developer account's weekly roundup.

Why it mattersThe Agents API can now start a browser plus agent with one call and offers light or large hosted environments. Bedrock Managed Agents makes the same API reachable on AWS.

ArtificialAnlys@ArtificialAnlys

Artificial Analysis places Ideogram 4.5 at #23 on its image editing leaderboard and #34 on text to image, with pricing from $0.008 to $0.22 per image by quality tier. It is live in the API, with open weights promised.

Why it mattersIdeogram 4.5 is ranked #23 for image editing and #34 for text to image, with a $0.22 per edit price at high quality and open weights coming. Useful if you pick image models for edit pipelines.

ArtificialAnlys@ArtificialAnlys

Artificial Analysis launched a side-by-side model comparison tool covering its Intelligence Index, individual benchmarks, a domain Capability Index (finance, legal, healthcare, engineering, economics), itemized token pricing, output speed and latency.

Why it mattersLets you compare models side by side on benchmark scores, domain capability, cached/input/output token cost, and speed in one place, which helps with model selection for agent workloads.

The 5 Levels of Self-Driving Production — Eric Schwartz, Traversal

Traversal's Eric Schwartz lays out five levels of self-driving production, arguing root cause analysis is a causal problem rather than an observability one.

Why it mattersIt gives a framework for judging AI SRE tools and explains why dashboards show what broke but not why. That matters as agent-written code increases production debugging load.

videoAI Engineer

Lessons from Generating 12 Trillion Synthetic Tokens — Bogdan Gaza, DatologyAI

DatologyAI's CTO details running synthetic data generation at about 12 trillion tokens.

Why it mattersShows how a seeded-rephrasing synthetic data pipeline was scaled to about 12 trillion tokens on Ray, KubeRay and vLLM, with specific fixes for S3 metadata, GPU failures and inference throughput.

videoAI Engineer
OpenAIDevs@OpenAIDevs

OpenAI's developer team summarizes its September releases: GPT-6.1 Sol for agentic coding, always-on dots agents, Codex cloud environments and code review, new CLI commands, plugin extensions, a Decisions API preview and an Ultrafast mode.

Why it mattersLists what changed in OpenAI's developer stack: GPT-6.1 Sol with a 95% cached-input discount, a Decisions API preview, Ultrafast mode up to 8x faster, and new Codex CLI commands. Useful for deciding what to adopt.

New: Model Router Benchmarks Compare 7 routers side by side on quality, speed,…

OpenRouter published Model Router Benchmarks comparing seven routers, including NVIDIA Switchyard and Unbiased Pareto, on quality, speed and cost across six benchmarks.

Why it mattersRouters that pick models automatically are hard to compare. This gives side-by-side quality, speed and cost data across seven routers so you can evaluate them before adopting one.

articleOpenRouter
Vals AI@ValsAI

Vals AI launched the Vals Web Search Index, a third-party benchmark run with the same settings for every provider, with Exa, Keenable, Parallel and Tavily as partners. It aims to fix contamination, realism and answer-leakage problems.

Why it mattersSearch tool vendors have been judged by their own benchmarks. A single third-party index with identical settings lets agent builders compare web search providers fairly.

Why AI Didn't Actually Make You Ship Faster — Gabriel Spencer-Harper, Meticulous

Meticulous CEO Gabriel Spencer-Harper argues verification is the new bottleneck.

Why it mattersAs agents write code faster than people can review it, assertion-based tests fall behind. This describes a way to run visual regression checks on every pull request without hand-written tests.

videoAI Engineer

Why 99% Accurate Browser Agents Still Fail — Derek Meegan, Browserbase

Browserbase's Derek Meegan explains why a 99%-per-step browser agent succeeds about 36% of the time over 100 steps.

Why it mattersIt quantifies why long browser-agent runs fail and shows an architecture of deterministic tools, OCR verification, separated authentication and skills for reliable unattended runs.

videoAI Engineer
ClaudeDevs@ClaudeDevs

Claude Code adds a built-in 'You should know' plugin, implemented as a mod that spins off a side agent to watch Claude's output and surface easily missed information. Enable it with the plugin command; Anthropic also published a getting-started post on mods.

Why it mattersClaude Code now ships a built-in plugin that uses a side agent to flag important details in output you might miss, enabled with one command.

ArtificialAnlys@ArtificialAnlys

Artificial Analysis places Ideogram 4.5 at #23 in image editing and #34 in text-to-image. High quality costs $0.22 per edit and $0.10 per generation, lower tiers start at $0.008, and it is on the Ideogram API with open weights promised.

Why it mattersGives ranked position and per-image API pricing across quality tiers for Ideogram 4.5, useful when choosing an image editing model for a product.

Stop Rationing Tokens: Let the Harness Pick the Model — Kimchi by Cast AI

Cast AI's team explains Kimchi, an open-source coding harness that picks models by task outcome and cost per task.

Why it mattersIt argues token rationing hurts developers and proposes measuring cost per task, with the harness routing between proprietary and open models. That is relevant to anyone managing coding-agent spend.

videoAI Engineer
repoempath75

A lean formalization of From Linearity to Borrowing

A hobbyist reports Claude mechanized the From Linearity to Borrowing paper in Lean over 3-4 weeks, including the Fundamental Property and Adequacy theorems that well-typed programs terminate with an empty heap.

Why it mattersShows a coding agent producing a near-complete Lean formalization of a recent programming-languages paper in weeks, a practical signal for AI-assisted formal verification.

Updates to Full Disk Access in macOS

Apple says it will tighten Full Disk Access in macOS with controls requiring very explicit user action, pointing to the growing risk of capable autonomous AI agents holding sweeping access to files, mail, messages and browsing history.

Why it mattersApple will add stricter, more explicit user consent for Full Disk Access on macOS, explicitly citing AI agents. Agent tools that depend on broad filesystem access should expect a harder permission flow.

articlenotfirstpost
AIatMeta@AIatMeta

Meta shares six papers from mathematicians working with Muse Spark 1.1 and 1.2 in Thinking Mode on open problems. Humans guided the work, a second group reviewed it, and each paper marks human versus AI-drafted passages and credits prior and parallel work.

Why it mattersIt documents frontier models contributing to open research problems through a plain chat interface with no scaffold, along with a human-review protocol that marks AI-drafted passages.

Inception@_inception_ai

Inception launched Mercury Voice, a diffusion LLM for voice agents. A dental receptionist demo shows median end-to-end latency under 300ms with tool calls, and access is via a playground and enterprise sales.

Why it mattersMercury Voice is a diffusion LLM claiming 2x+ lower latency than GPT-6 Luna, with median end-to-end under 300ms including tool calls. This matters when latency is the bottleneck in voice agents.

StanfordAILab@StanfordAILab

SWE-chat v2 from Stanford's SALT group grows the in-the-wild dataset of coding agent interactions to 230K prompts across 18K sessions from real users, now on HuggingFace, giving researchers grounded data on how developers work with coding agents.

Why it mattersA large dataset of real developer interactions with coding agents lets you study how people actually prompt and where agents fail, instead of relying on synthetic benchmarks.

ALoDLM 1.7B released

Amazon released ALoDLM-1.7B, a diffusion language model initialized from Qwen3 that adaptively loops shared transformer layers, keeps latent states for unresolved tokens.

Why it mattersShows a diffusion LM that spends variable compute per token via recurrent loops and a learned halting policy, with tunable thresholds trading refinement against parallel decoding speed.

articleAmazon model releases

ALoDLM 8B released

Amazon's ALoDLM-8B is a diffusion language model built on a Qwen3 backbone that refines unresolved tokens through adaptive recurrent passes and commits confident tokens in parallel.

Why it mattersOffers an 8B diffusion LM that allocates compute per token and decodes in parallel, with reported accuracy and throughput trade-offs on a B200.

articleAmazon model releases
EpochAIResearch@EpochAIResearch

Epoch AI estimates that AI infrastructure could soon run hundreds of millions to billions of agents, rivaling global human working hours at the high end. The thread walks through scenarios driven by model efficiency and whether AI demand keeps climbing.

Why it mattersEpoch AI sizes how many concurrent agents planned compute could run under different efficiency and demand scenarios, which helps engineers reason about future agent availability and cost.

GPT-6 Astra and Claude 5.5 Opus race to create the best StarCraft bot

StarSkirmish Hillclimb tests frontier LLMs on writing Protoss bots in C++ against BWAPI, climbing five opponent tiers up to top human-written bots.

Why it mattersRemoves the one-hour reasoning cap from an agentic coding benchmark, showing how frontier models perform on long-horizon implementation tasks against strong human-written bots.

article__cayenne__

GitHub September ship log: HydraFusion, new Copilot models, star history API

GitHub's September 2026 ship log covers Project HydraFusion, a multi-model orchestration preview now in the Copilot app and VS Code, new Copilot models.

Why it mattersLists what changed in Copilot this month: multi-model HydraFusion preview, three new models in the picker, and a privacy-safe star history API.

articleGitHub
Factory@FactoryAI

Factory launched revamped Analytics for its coding agent, breaking down consumption, model efficiency and adoption by model and user across all organization sessions, aimed at cost visibility for agentic engineering teams.

Why it mattersReports agent consumption by model and by user across every session instead of a single abstract compute number. Helps engineering leaders see where agent spend goes.

allen_ai@allen_ai

Ai2 releases AstaBrief 8B, a Qwen3-8B fine-tune (SFT plus DPO) that turns research questions and literature excerpts into cited reports. Weights and training data are open, and it powers Fast mode in Asta, running about 3.5x faster than the Claude-based mode.

Why it mattersAn open 8B model produces cited research reports in about 51 seconds versus 178 for a Claude-powered pipeline, and runs locally on your hardware. Weights and data are open, so you can reproduce or adapt the recipe.

An index of the vibe-coding frontier. Corrections welcome.