Intel
Page 02
Together Link: open models in the harness you already use. Start with one command today.
Together AI's Link connects Claude Code, Claude Desktop, Codex, OpenCode and Pi to open models hosted on Together via a one-line install.
Why it mattersLets teams point existing coding agent harnesses at open models such as Kimi K3 and GLM 5.3 with one command and revert easily. Relevant for cutting coding-agent spend without changing workflow.

Qwen3.8 27B addition in words
Simon Willison reruns a GPT-4o-era experiment on local Qwen3.8-27B, testing integer addition expressed in English words across 5,070 cases.
Why it mattersShows where a local 27B model's arithmetic breaks down, and how reasoning mode changes it, with exact accuracy figures by operand length. Useful when deciding if a local model needs a calculator tool.

Agents That Write Their Own Tools at Runtime — Sandhya Subramani, AWS
AWS's Sandhya Subramani demonstrates meta-tooling in Strands Agents, where agents write, load and use their own tools and sub-agents at runtime.
Why it mattersShows that a system prompt plus editor, shell and load_tool lets an agent author tools and sub-agents at runtime. It also lists the evals and sandboxing needed before trusting a self-modifying agent.

AI Coding Agents Are Breaking Big Codebases — Dan Adler, Sourcegraph
Sourcegraph's CEO makes the case that large-scale code visibility is infrastructure for coding agents, describing codebase decay from AI output and introducing Agentic Batch Changes for cross-repo edits.
Why it mattersArgues that agents cannot act on code they cannot search, and that AI-generated code accelerates duplication and drift. Sourcegraph's Agentic Batch Changes targets multi-repo changes from a single prompt.

Keyword Search Is Dying. Is Your Catalog Ready for AI Agents? — PayPal
PayPal's agentic commerce lead explains why keyword-optimized catalogs fail for agents and shares experiment results on product data enrichment, hybrid search.
Why it mattersPayPal's experiments found that enriching merchant product data improves agent recommendations, but unstructured boilerplate dilutes the signal and increases hallucination. Useful if you are making catalogs or retrieval readable to agents.

The Human Is an Async API — Melanie Warrick, Temporal
Temporal's Melanie Warrick shows how to build human-in-the-loop agents with durable execution, using wait conditions and signals with Google ADK and LangGraph.
Why it mattersA blocking ask-the-human call breaks in production. Durable workflows with a wait condition and signal let an agent pause for minutes or weeks and resume after a worker crash without losing state.

We Built an AI Support Agent That Resolves 80% of Tickets — AssemblyAI
AssemblyAI's Joey support agent, built on the Claude Agent SDK with checked-out docs, embeddings, agentic search and one CLAUDE.md, resolves 80% of tickets for about $700 a month, then gains voice via a websocket.
Why it mattersGives a concrete architecture and cost figure for a support agent, including turning escalations into roadmap items and a deploy loop that ships fixes in about 30 seconds.

OpenAI’s Head of ChatGPT: We’re entering a new era of AI (again) | Tibo Sottiaux

Dashboards Are Dead — Sarah Simionescu, Composio
Composio's Sarah Simionescu argues dashboards are obsolete, explains why a dozen MCP servers alone fail.
Why it mattersNames concrete reasons raw MCP wiring breaks down and one approach, tool search with execution plans, for agents working across apps with less context load.

Stop Fine-Tuning to Fix Retrieval Problems — Anant Srivastava
Oracle's Anant Srivastava argues prompt, memory and weights are three tools for three jobs, not a ladder.
Why it mattersGives a diagnostic for choosing between prompt edits, retrieval and fine-tuning. It explains why fine-tuning to fix retrieval problems backfires.
We're going to need default hard budget caps on pretty much everything
Simon Willison argues that usage-billed APIs and hosts need hard, default-on spending caps because coding agents lower the friction of spinning up costly services.
Why it mattersAgents can create paid resources faster than anyone watches them. The argument gives a concrete design rule (hard caps by default, opt-out with explicit consent) to apply to your own services and agent budgets.

The Agent Said It Was Done. The Database Disagreed.
Microsoft's ThinkingBox benchmark grades agents on the records they leave behind across 507 stateful MCP workflows, repeated 20 times to measure consistency and its cost.
Why it mattersAgents can make well-formed tool calls and still leave the database wrong. Grading terminal state across 20 repeated runs measures reliability, which tool-call and final-answer checks miss.

What Makes Open Models Fast in Production — Sujee Maniyam, Nebius
Nebius engineers cover production serving of open LLMs.
Why it mattersLists the specific serving layers that determine cost and latency for self-hosted open models, which helps when weighing a closed API against running open weights.

An Interaction Is All You Need — Ivan Leo, Google DeepMind
Google DeepMind presents the Gemini Interactions API and Managed Agents.
Why it mattersGemini's Interactions API moves conversation state server-side and Managed Agents run in persistent remote sandboxes with credential injection. Agent builders can drop manual thought-signature handling and hosted-harness plumbing.

Automating my 35mm film scanning pipeline
A photographer lets Claude automate a 35mm scanning pipeline, then recounts how it damaged the scanner by misdriving the carriage, how the fault was traced.
Why it mattersClaude drove a scanner through SANE and nearly broke it after a blind reverse move at the end of each scan. A concrete case for sandboxing and confirming hardware actions before granting agents device access.
Agents don't need memory, they need documentation
An essay arguing that RAG-based agent memory plugins are structurally flawed because similarity search can't tell current from stale information.
Why it mattersThe piece argues RAG-style memory plugins fail because similarity search can't distinguish correct from stale information.

I Built a Personal AI Agent on a Raspberry Pi — Jeremy Adams, Neo4j
A live build of a wearable personal agent on a Raspberry Pi 4B.
Why it mattersDemonstrates graph-based agent memory (POLE schema) over MCP on cheap, isolated hardware, with offline capture that syncs later. Gives builders a concrete pattern for persistent memory beyond chat history.

How VS Code Went from Monthly to Weekly Releases with AI — Harald Kirschner
A Microsoft VS Code product lead details how the team moved to weekly releases with agents.
Why it mattersShows how a 50M-user product restructured CI, review, triage and evals around coding agents, with measurable results such as code survival rising from 55% to 86%. Useful as a template for agent-ready engineering workflows.

Why Specialized AI Could Beat The God Model
a16z talks with OpenRouter's Alex Atallah and Replit's Amjad Masad about specialized models over a single god model.
Why it mattersArgues that routing and combining specialized models can match frontier performance at lower cost and reduce single-provider dependence. Relevant when designing model selection and agent architecture.

Crusoe CEO: Why Everyone Gets GPU Depreciation & AI Energy Costs Wrong
Crusoe's CEO discusses AI data center bottlenecks.
Why it mattersCompute supply, energy limits and GPU lifespan assumptions shape future inference cost and availability. The talk offers an operator's view on those constraints.

Germany's new sovereign AI model Kolibri
Walkthrough of Aleph Alpha's Kolibri, an open-weight 78B mixture-of-experts model (3.5B active per token) for German and English under Apache 2.0.
Why it mattersA new Apache 2.0 open-weight MoE with 3.5B active parameters and 262k context is a self-hostable option for German and English workloads with data-residency needs. The post covers when it fits and where it falls short.

Kolibri: A Sovereign Open-Weight Model
Aleph Alpha releases Kolibri, an English-German MoE model (78B total, 3B active, up to 1M context) with Apache 2.0 weights, built through a reusable training pipeline for sovereign, regulated deployments.
Why it mattersAn Apache 2.0 English-German MoE with only 3B active parameters and 1M context can run on-premise for regulated work. It gives teams with sovereignty or compliance needs another open-weight option.

80s Choplifter game ported to Vision Pro as an agent eval (WIP)
A developer building a Vision Pro Choplifter-style game uses Astra and Claude Opus 5.5 with Blender MCP, assigning them different roles and cross-checking output.
Why it mattersDescribes a working split between two coding agents (planning and review versus focused implementation) and the handoff cost. Notes that comfort testing in VR still needs a human.

Your LLM App Returned 200 OK. It Was Still Wrong. — Marina Petzel, Datadog
Datadog's Marina Petzel explains why latency and error golden signals miss GenAI failures.
Why it mattersAn LLM app can return 200 OK and still be wrong or expensive. This lists the cost, safety and quality signals to add to standard monitoring.

YOLO Mode, Safely: MicroVM Sandboxes for Any Agent — Rowan Christmas, Docker
Docker's Rowan Christmas shows a coding agent finding browser history and bank data in five prompts, then demos Docker Sandboxes.
Why it mattersIt shows concretely why running agents in YOLO mode on a laptop is risky, and how a microVM sandbox gives isolation that works with Claude Code, Codex or other agents.
OpenAI's Agents API gained computer use that spins up a browser and agent in a single call, hosted environment sizes, a Bedrock Managed Agents path on AWS, and GPT-6.1 Sol support, according to the developer account's weekly roundup.
Why it mattersThe Agents API can now start a browser plus agent with one call and offers light or large hosted environments. Bedrock Managed Agents makes the same API reachable on AWS.
Artificial Analysis places Ideogram 4.5 at #23 on its image editing leaderboard and #34 on text to image, with pricing from $0.008 to $0.22 per image by quality tier. It is live in the API, with open weights promised.
Why it mattersIdeogram 4.5 is ranked #23 for image editing and #34 for text to image, with a $0.22 per edit price at high quality and open weights coming. Useful if you pick image models for edit pipelines.
Artificial Analysis launched a side-by-side model comparison tool covering its Intelligence Index, individual benchmarks, a domain Capability Index (finance, legal, healthcare, engineering, economics), itemized token pricing, output speed and latency.
Why it mattersLets you compare models side by side on benchmark scores, domain capability, cached/input/output token cost, and speed in one place, which helps with model selection for agent workloads.

The 5 Levels of Self-Driving Production — Eric Schwartz, Traversal
Traversal's Eric Schwartz lays out five levels of self-driving production, arguing root cause analysis is a causal problem rather than an observability one.
Why it mattersIt gives a framework for judging AI SRE tools and explains why dashboards show what broke but not why. That matters as agent-written code increases production debugging load.

Lessons from Generating 12 Trillion Synthetic Tokens — Bogdan Gaza, DatologyAI
DatologyAI's CTO details running synthetic data generation at about 12 trillion tokens.
Why it mattersShows how a seeded-rephrasing synthetic data pipeline was scaled to about 12 trillion tokens on Ray, KubeRay and vLLM, with specific fixes for S3 metadata, GPU failures and inference throughput.
OpenAI's developer team summarizes its September releases: GPT-6.1 Sol for agentic coding, always-on dots agents, Codex cloud environments and code review, new CLI commands, plugin extensions, a Decisions API preview and an Ultrafast mode.
Why it mattersLists what changed in OpenAI's developer stack: GPT-6.1 Sol with a 95% cached-input discount, a Decisions API preview, Ultrafast mode up to 8x faster, and new Codex CLI commands. Useful for deciding what to adopt.

New: Model Router Benchmarks Compare 7 routers side by side on quality, speed,…
OpenRouter published Model Router Benchmarks comparing seven routers, including NVIDIA Switchyard and Unbiased Pareto, on quality, speed and cost across six benchmarks.
Why it mattersRouters that pick models automatically are hard to compare. This gives side-by-side quality, speed and cost data across seven routers so you can evaluate them before adopting one.

Vals AI launched the Vals Web Search Index, a third-party benchmark run with the same settings for every provider, with Exa, Keenable, Parallel and Tavily as partners. It aims to fix contamination, realism and answer-leakage problems.
Why it mattersSearch tool vendors have been judged by their own benchmarks. A single third-party index with identical settings lets agent builders compare web search providers fairly.

Why AI Didn't Actually Make You Ship Faster — Gabriel Spencer-Harper, Meticulous
Meticulous CEO Gabriel Spencer-Harper argues verification is the new bottleneck.
Why it mattersAs agents write code faster than people can review it, assertion-based tests fall behind. This describes a way to run visual regression checks on every pull request without hand-written tests.

Why 99% Accurate Browser Agents Still Fail — Derek Meegan, Browserbase
Browserbase's Derek Meegan explains why a 99%-per-step browser agent succeeds about 36% of the time over 100 steps.
Why it mattersIt quantifies why long browser-agent runs fail and shows an architecture of deterministic tools, OCR verification, separated authentication and skills for reliable unattended runs.

Claude Code adds a built-in 'You should know' plugin, implemented as a mod that spins off a side agent to watch Claude's output and surface easily missed information. Enable it with the plugin command; Anthropic also published a getting-started post on mods.
Why it mattersClaude Code now ships a built-in plugin that uses a side agent to flag important details in output you might miss, enabled with one command.
Artificial Analysis places Ideogram 4.5 at #23 in image editing and #34 in text-to-image. High quality costs $0.22 per edit and $0.10 per generation, lower tiers start at $0.008, and it is on the Ideogram API with open weights promised.
Why it mattersGives ranked position and per-image API pricing across quality tiers for Ideogram 4.5, useful when choosing an image editing model for a product.

Stop Rationing Tokens: Let the Harness Pick the Model — Kimchi by Cast AI
Cast AI's team explains Kimchi, an open-source coding harness that picks models by task outcome and cost per task.
Why it mattersIt argues token rationing hurts developers and proposes measuring cost per task, with the harness routing between proprietary and open models. That is relevant to anyone managing coding-agent spend.
A lean formalization of From Linearity to Borrowing
A hobbyist reports Claude mechanized the From Linearity to Borrowing paper in Lean over 3-4 weeks, including the Fundamental Property and Adequacy theorems that well-typed programs terminate with an empty heap.
Why it mattersShows a coding agent producing a near-complete Lean formalization of a recent programming-languages paper in weeks, a practical signal for AI-assisted formal verification.

Updates to Full Disk Access in macOS
Apple says it will tighten Full Disk Access in macOS with controls requiring very explicit user action, pointing to the growing risk of capable autonomous AI agents holding sweeping access to files, mail, messages and browsing history.
Why it mattersApple will add stricter, more explicit user consent for Full Disk Access on macOS, explicitly citing AI agents. Agent tools that depend on broad filesystem access should expect a harder permission flow.
Meta shares six papers from mathematicians working with Muse Spark 1.1 and 1.2 in Thinking Mode on open problems. Humans guided the work, a second group reviewed it, and each paper marks human versus AI-drafted passages and credits prior and parallel work.
Why it mattersIt documents frontier models contributing to open research problems through a plain chat interface with no scaffold, along with a human-review protocol that marks AI-drafted passages.

Inception launched Mercury Voice, a diffusion LLM for voice agents. A dental receptionist demo shows median end-to-end latency under 300ms with tool calls, and access is via a playground and enterprise sales.
Why it mattersMercury Voice is a diffusion LLM claiming 2x+ lower latency than GPT-6 Luna, with median end-to-end under 300ms including tool calls. This matters when latency is the bottleneck in voice agents.
SWE-chat v2 from Stanford's SALT group grows the in-the-wild dataset of coding agent interactions to 230K prompts across 18K sessions from real users, now on HuggingFace, giving researchers grounded data on how developers work with coding agents.
Why it mattersA large dataset of real developer interactions with coding agents lets you study how people actually prompt and where agents fail, instead of relying on synthetic benchmarks.

ALoDLM 1.7B released
Amazon released ALoDLM-1.7B, a diffusion language model initialized from Qwen3 that adaptively loops shared transformer layers, keeps latent states for unresolved tokens.
Why it mattersShows a diffusion LM that spends variable compute per token via recurrent loops and a learned halting policy, with tunable thresholds trading refinement against parallel decoding speed.

ALoDLM 8B released
Amazon's ALoDLM-8B is a diffusion language model built on a Qwen3 backbone that refines unresolved tokens through adaptive recurrent passes and commits confident tokens in parallel.
Why it mattersOffers an 8B diffusion LM that allocates compute per token and decodes in parallel, with reported accuracy and throughput trade-offs on a B200.
Epoch AI estimates that AI infrastructure could soon run hundreds of millions to billions of agents, rivaling global human working hours at the high end. The thread walks through scenarios driven by model efficiency and whether AI demand keeps climbing.
Why it mattersEpoch AI sizes how many concurrent agents planned compute could run under different efficiency and demand scenarios, which helps engineers reason about future agent availability and cost.

GPT-6 Astra and Claude 5.5 Opus race to create the best StarCraft bot
StarSkirmish Hillclimb tests frontier LLMs on writing Protoss bots in C++ against BWAPI, climbing five opponent tiers up to top human-written bots.
Why it mattersRemoves the one-hour reasoning cap from an agentic coding benchmark, showing how frontier models perform on long-horizon implementation tasks against strong human-written bots.

GitHub September ship log: HydraFusion, new Copilot models, star history API
GitHub's September 2026 ship log covers Project HydraFusion, a multi-model orchestration preview now in the Copilot app and VS Code, new Copilot models.
Why it mattersLists what changed in Copilot this month: multi-model HydraFusion preview, three new models in the picker, and a privacy-safe star history API.

Factory launched revamped Analytics for its coding agent, breaking down consumption, model efficiency and adoption by model and user across all organization sessions, aimed at cost visibility for agentic engineering teams.
Why it mattersReports agent consumption by model and by user across every session instead of a single abstract compute number. Helps engineering leaders see where agent spend goes.
Ai2 releases AstaBrief 8B, a Qwen3-8B fine-tune (SFT plus DPO) that turns research questions and literature excerpts into cited reports. Weights and training data are open, and it powers Fast mode in Asta, running about 3.5x faster than the Claude-based mode.
Why it mattersAn open 8B model produces cited research reports in about 51 seconds versus 178 for a Claude-powered pipeline, and runs locally on your hardware. Weights and data are open, so you can reproduce or adapt the recipe.
An index of the vibe-coding frontier. Corrections welcome.