Intel
Page 06
OpenAI's Decisions API, powered by the Luna model, handles constrained decision making with visual inputs and is tuned to respond in a few hundred milliseconds end to end.
Why it mattersA dedicated low-latency endpoint for constrained decisions (choices, scores, yes/no) with image input suits classification, routing and moderation steps in agents without a full chat model call.

OpenAI releases GPT-6.1 Sol, an upgrade over GPT-6 Sol in coding, computer use and professional work. It approaches GPT-6 Astra on several benchmarks at lower cost, with $0.10 per million cached input tokens. It is live in ChatGPT Work and Codex.
Why it mattersNear-Astra coding and computer-use performance at roughly a fifth of the price, with cheap cached input, lowers the cost of running agents at scale. It is available in Codex today.

OpenAI introduced GPT-6.1 Sol, positioned near Astra in intelligence at one fifth the price with a 95% cache read discount. An ultrafast 8X speed option exists for Astra now and is coming to 6.1 Sol.
Why it mattersGPT-6.1 Sol claims near-Astra capability at about one fifth the price with a 95% cache read discount, which changes the cost math for agent workloads that reuse long contexts.

OpenAI launches dots, always-on agents powered by GPT-6 Astra. Each runs on its own cloud computer, works across 4,000+ apps through plugins, and has user-set boundaries for what it does alone. It is available to Pro, Business Premium and Enterprise users.
Why it mattersA persistent background agent with its own cloud computer and user-set rules for acting alone, asking first or never acting is a reference design for agent autonomy. Availability is limited to Pro, Business Premium and Enterprise.

OpenAI announced ChatGPT Space, collaborative rich notes with live visualizations that people and agents edit together and that agents keep updated in the background, pitched as AGENTS.md taken further.
Why it mattersShared rich notes that humans and agents edit and agents update in the background give a persistent context layer for delegated work, like AGENTS.md extended beyond a repo.
GPT 6.1 Sol: Near-Astra intelligence for a fifth of the price
OpenAI's release post for GPT-6.1 Sol, positioned as near-Astra intelligence at roughly one fifth of the price.
Why it mattersA new OpenAI model reportedly near GPT-6 Astra in capability at a much lower price, which affects model choice and cost for coding and agent workloads.

OpenAI's dots are 24/7 agents with their own computer and browser, powered by Astra and connected to over 4,000 apps. A primary dot is included in Pro without drawing usage, with multi-dot teams coming.
Why it mattersDots are persistent cloud agents with their own computer and browser, connected to 4k+ apps, included in Pro without consuming plan usage. Teams of dots are promised soon.

A study of 1,867 repositories finds agent instruction files like CLAUDE.md more than triple in size because rule rationale is lost. Commenting each rule with its failure and evidence kept prompts near ideal length with equal constraint adherence.
Why it mattersAgent instruction files like CLAUDE.md grow because nobody remembers why each rule exists. Attaching the failure, hypothesis and evidence behind each rule cut excess prompt length from 211% to 1.4% without hurting compliance.

Why has Shopify dropped React Native?
Gergely Orosz examines why Shopify reversed its React Native bet for native mobile development, arguing that stronger AI coding agents changed the tradeoff.
Why it mattersShopify concluded that AI coding agents now make native iOS and Android development viable at scale, undercutting the cross-platform case that led it to React Native.

NVIDIA Kumo Tabular Sets a New Accuracy-Efficiency Frontier for Tabular Prediction
NVIDIA releases Kumo Tabular, an open tabular foundation model in three sizes.
Why it mattersKumo Tabular predicts labels for new table rows in one forward pass, with no training, tuning or feature engineering. Open weights under a commercial license let you test in-context tabular prediction against gradient-boosted trees.

We Tested the New Wave of Personal AI Agents
The creator of Assistant Benchmark discusses testing dozens of personal AI agents on email, travel, and finance tasks, why proactivity may be the moat, agent-to-agent interaction.
Why it mattersReports what a real-task benchmark of personal agents finds: cost-saving tasks beat time-saving ones, proactivity is the differentiator, and merchants are splitting on letting agents in. Useful for anyone designing consumer agents.

Getting the Source Right, Not Just the Fact: Source-Aware Verification for MCP Agents
Multiverse Computing presents ProvenanceGuard, a verification layer over black-box MCP agents that tracks which tool output supports each claim.
Why it mattersPooled-evidence faithfulness checks pass claims that are true but attributed to the wrong MCP source. Source-aware verification catches those errors in data-sensitive agents like support and clinical assistants.

OpenRouter released an openrouter-decisions skill for coding agents. In its tests across three models and four tasks, decision-model use rose from 4 of 27 runs to 27 of 27, and a grader preferred the skill output every time.
Why it mattersCoding agents often fake decision models with keyword rules or yes/no chat prompts that return no probability to threshold. This skill steers them to ask in the right shape and pin a model version.

Using AI to chart a course for our post-quantum migration
Cloudflare explains how it uses AI-driven cryptography discovery to map where public-key crypto is used across its products.
Why it mattersShows an infrastructure team applying AI to inventory cryptographic usage across a large platform. It is a practical pattern for AI-assisted codebase auditing ahead of a migration deadline.

Search trace spans from the Vercel CLI
Vercel's CLI gained a traces search command that filters spans by status, service, deployment or request ID, supports a KQL subset for duration queries.
Why it mattersYour coding agent can now query production trace spans from the terminal with structured JSON output, so it can investigate errors and latency without the dashboard.

We tested our own WAF with frontier AI models. Here’s what we found
Cloudflare built an LLM-driven tester that starts from known exploits and iterates on encoding and delivery using only HTTP responses.
Why it mattersShows a black-box, LLM-driven payload mutation loop for testing whether defenses hold against AI attackers. You can reuse the approach to red-team your own WAF rules.

Introducing Threat Signals: agentic skills for open-source threat intelligence, free for every Cloudflare account
Cloudflare launches Threat Signals, agentic skills that summarize threat reports, extract indicators of compromise and tag them into a private dataset.
Why it mattersApplies AI skills to convert unstructured threat reports into structured, WAF-ready indicators. It is free on every Cloudflare account, so you can test it without extra cost.

Evolving our calendar assistant Reclaim to be AI-native without starting over
Dropbox engineers describe evolving the Reclaim calendar assistant to accept natural-language requests.
Why it mattersShows one path for adding conversational agents to a mature product: ground requests in existing calendar context and keep the current scheduler, rather than rebuilding from scratch.

Language Models for Text Classification: From Bag-of-Words to Jev
Sebastian Raschka traces language-model classification from bag-of-words to the Jev model, guesses at its methodology.
Why it mattersExplains where a general-purpose classifier model beats both frontier LLMs (on speed and cost) and task-specific classifiers (on generality). This helps engineers decide which to use for high-volume classification steps in agent pipelines.
DevDay 2026 Recap
OpenAI's own recap of DevDay 2026, spanning over 20 launches.
Why it mattersOpenAI shipped a new flagship model line (GPT-6 Astra) alongside Codex and API updates at DevDay 2026, changing what capabilities and tools are available to build agentic applications on.

OpenAI's Codex lead explains that Pro $200 reopens with usage calculation netting out at half the prior API dollar value, no 5-hour limit, and GPT-6 Sol and Luna API prices cut 50%.
Why it mattersThe reopened Pro $200 plan nets about half the API spend of the old plan, keeps no 5-hour limit, and GPT-6 Sol and Luna API prices dropped 50%. It resets cost expectations for heavy Codex use.

How distributed Claude Code is built on a tunnelling service
NFLTR describes extending Claude Code's hub-and-subagent model across machines.
Why it mattersShows how to run Claude Code subagents on multiple machines with outbound-only connections and exactly-once result delivery, useful when data or services cannot leave a particular host.

LLM Judge Validation Under Sparse Overlap: From Inference to Design
A NeurIPS 2026 paper proving that sparse human-annotation overlap drives wrong LLM-judge deployment decisions.
Why it mattersIf you validate an LLM judge against sparse human labels, low annotator overlap can pick the wrong judge 65% of the time. The paper gives an overlap threshold (0.25) and a stratified sampling scheme to reduce false rejections.

Methodological Harness in Agentic Software Engineering: An Empirical Study on Mining Software Repositories
An empirical mining study of repositories with agentic activity, measuring prevalence and co-occurrence of rule files, specifications and architectural decision records.
Why it mattersShows how teams actually configure coding agents in real repositories, including rule and context files, specs and decision records. It helps you compare your setup against observed adoption of context engineering and graduated autonomy.

Verification as an Architectural Layer for LLM Agents: A V-Model Design, and a Pilot Study of Its Deterministic Core
A V-model architecture for LLM agents that pairs each specification level with a verifier, splits deterministic gates from LLM judges.
Why it mattersReAct agents judge their own output and run until a budget stops them. This design separates deterministic gates from optional LLM judges per level, so a rejection localizes the faulty stage and the agent can halt cleanly.

What Drives Dialectal Jailbreaks? An Ablation of Surface Form, Cultural Framing, and Strategy Banks
A 36-cell ablation across Shanghainese, Cantonese, Mandarin and English shows that dialectal surface form does not drive jailbreaks.
Why it mattersJailbreak success comes from the expressiveness of the prompt-strategy bank, not from obscure dialects. A culture-neutral strategy bank reaches the same ceiling at near-single-query cost, so defenses should target optimized strategies.

Beyond the Model: Demystifying Harness Effects in Software Engineering Agents
An empirical study comparing mini-SWE-agent and OpenCode across ten Qwen and DeepSeek models on SWE-bench Pro, ProgramBench and GitTaskBench, then ablating five harness components in a lightweight NanoHarness to quantify how harness design shapes coding-agent results.
Why it mattersAgent performance depends on the harness as well as the model.

AsynCodeBench: Benchmarking Collaboration of Asynchronous Multi-Agent Systems in Software Engineering
AsynCodeBench represents each multi-agent software task as a dependency graph with executable checkers.
Why it mattersMulti-agent coding is usually scored by task pass rate, which conflates individual coding skill with coordination. This benchmark measures cross-agent dependency satisfaction directly, letting you evaluate orchestration separately.

[AINews] Opus 5.5 is good at explainer videos
A daily digest of the frontier model wave.
Why it mattersOpus 5.5 scores 62% on Terminal-Bench-Science at xhigh effort and falls to 59% at max, so choosing the effort level matters for agent runs. The digest also compares GPT-6 variants and other new models.

Claude Code’s Next Era — Thariq Shihipar, Anthropic
Latent Space talks with Anthropic's Thariq Shihipar about Claude Code's next phase, alongside a recap of Anthropic's recent releases including Sonnet 5.5, Opus 5.5, Claude Mods.
Why it mattersAn interview with an Anthropic Claude Code engineer on where the tool is heading, including the recent Claude Mods and plugin releases, useful for planning around coding agent workflows.

The Future of Claude Code: Mods, Mutable Software, & Multiplayer Agents — Thariq Shihipar, Anthropic
Anthropic's Thariq Shihipar covers Claude Code power-user workflows, Mods for rewriting the harness, multiplayer agents, mutable software.
Why it mattersDescribes where the Claude Code harness is heading: customizable Mods, multiplayer agent workflows, and the possible end of CLAUDE.md. It also walks through incidents where agents found unexpected exploits, informing sandboxing choices.

Helping personal agents shop more intelligently and reliably with Link
Stripe extends Link's wallet for agents with incremental authorization for price changes at checkout, guidance on checkout obstacles, spending analysis for better recommendations.
Why it mattersAgents that buy things can now raise an approved amount mid-purchase instead of restarting and placing a second card hold. Link also adds purchase protection if an agent errs, which lowers the trust barrier for agent-driven checkout.

What Is Web Bot Auth
Browserbase explains Web Bot Auth, a standard where agents sign requests to prove identity, why header and IP checks fail as bot traffic overtakes human traffic.
Why it mattersAgents that browse get blocked as bots. Web Bot Auth lets them prove identity cryptographically while sites keep control of admission, which shapes how browser agents will be allowed in.

Why You Can T Just Run Chromium In A Sandbox
Browserbase explains why launching Chromium in an agent's existing code sandbox works locally but struggles in production.
Why it mattersIf your agent uses a browser, running Chromium inside its code sandbox can break at scale, since the sandbox was built for bursty code execution. The post explains the resource and isolation gaps to weigh before doing it.

Image To Video Models Compared
OpenRouter compares Veo 3.1, Seedance, Kling and Grok Imagine Video for image-to-video work.
Why it mattersShows which video models accept first and last frames versus loose references, and their duration and resolution limits, so you can pick a model and set frame_images or input_references correctly.

GPT-6.1 Sol now available on AI Gateway
GPT-6.1 Sol is available on Vercel AI Gateway, improving on GPT-6 Sol for coding, computer use and document work.
Why it mattersA new OpenAI model with improved coding and computer use, and cheaper cached input, is callable through one gateway ID. That matters for agents that reuse long shared context.
llm-anthropic 0.30
The llm-anthropic plugin 0.30 adds Claude Sonnet 5.5, a refresh command that pulls the current model list from Anthropic's API.

Claude Sonnet 5.5
Simon Willison covers Anthropic's Sonnet 5.5: same price as Sonnet 5, faster, near Opus 5.5 on some coding tasks, and now the free-tier model.
Why it mattersSonnet 5.5 is 30%+ faster at the same price as Sonnet 5 and powers the free tier. Max thinking effort can burn 128,000 tokens ($1.28) and return nothing, so cap effort settings.

Building with Claude Sonnet 5.5: choosing, migrating, and tuning
Anthropic's playbook for building with Sonnet 5.5.
Why it mattersExplains when to pick Sonnet 5.5 over Opus 5.5, how it is priced at $2/$10 per million tokens, and that thinking is on by default, so code reading content[0].text can break. Helps with migration.

Automating eval design and hillclimbing with Claude Code
Anthropic lays out principles for trustworthy evals and hillclimbing, then shows how the claude-api skill's build-eval and hillclimb commands apply them in a codebase, including held-out sets to catch overfitting.
Why it mattersShows how to design evals that avoid fooling yourself and hillclimb one change at a time with a held-out set. The build-eval and hillclimb commands automate this inside Claude Code.

Vals AI ran GPT-6 Astra and Claude Opus 5.5 at every reasoning level on its new MysteryMechanism benchmark. Accuracy improves with longer thinking at up to 20x the cost, and trace inspection shows the two models spend effort in very different ways.
Why it mattersShows that raising reasoning effort on the same model can cost up to 20x more for its accuracy gain, and that GPT-6 Astra and Opus 5.5 use effort differently. Helps you choose reasoning levels by cost, not by default.

Cognition adds Claude Sonnet 5.5 to Devin Desktop and CLI, reporting 64.4% on FrontierCode 1.1 Main versus 56.2% for Sonnet 5 and ahead of Fable 5.1 at extra high effort. Anthropic's token prices are unchanged.
Why it mattersSonnet 5.5 is available in Devin Desktop and CLI. Cognition reports 64.4% on its FrontierCode benchmark against 56.2% for Sonnet 5, which helps when choosing a model for coding agents.

Cursor added Anthropic's Sonnet 5.5 as a selectable model, saying it matches Opus on many tasks. Coding-agent users get a new option for the model tier used in day-to-day editing and agent runs.
Why it mattersSonnet 5.5 can now be selected in Cursor, and Cursor reports it performs on par with Opus on many tasks. That may let you use a cheaper model for coding work you previously gave to Opus.

Stanford CME295 Transformers & LLMs | Autumn 2026 | Lecture 1 - Transformers
Opening lecture of Stanford's CME295 covering tokenization, word2vec, RNNs and LSTMs, self-attention with query, key and value, the encoder-decoder transformer, computational tricks and a worked end-to-end example.
Why it mattersA structured, chaptered walkthrough from tokenization to the full transformer, useful for engineers who want a solid grounding in how the models they build on actually work.

Andrew Ng argues agent breaches stem from weak sandboxing and describes OpenWorker's approach on NVIDIA OpenShell: pass in only task-relevant files, block secrets and arbitrary web access by default, enforce limits in deterministic code, and log every action.
Why it mattersEnforcing agent permissions in deterministic sandbox code, not prompts, keeps API keys, browser credentials and arbitrary web access away from an agent even under prompt injection.

How GLM5.3 Sparse Attention Affects HBM Memory Usage
SemiAnalysis examines why sparse attention does not shrink memory capacity needs, since top-k selection wants the whole context in HBM.
Why it mattersSparse attention cuts compute per token but top-k selection still needs the full KV cache in HBM. Tiered designs like HiSparse trade cache-miss I/O for capacity, which affects long-context serving cost.

How we found 24 Android vulnerabilities using our open source AI security agent
A GitHub Security Lab researcher shows how packaged taskflow prompts for the open-source Taskflow Agent guided an LLM through Android app audits, yielding more than 20 reported vulnerabilities.
Why it mattersSplitting an audit into incremental taskflow steps helps an LLM find complex vulnerabilities it would miss in one pass. The open-source taskflows can be run on your own repo, though they consume many premium requests.

Google's Project CC gives a family agent its own identity and sandbox
Google Labs expands Project CC into a family group agent.
Why it mattersShows one way to scope agent access: a separate identity, opt-in email forwarding and isolated compute rather than shared passwords. The design is useful to compare against your own multi-user agents.

Anthropic ships Sonnet 5.5, 30% faster than Sonnet 5 at unchanged pricing, with always-on thinking, effort guidance, and org-bound preserved thinking that resets when accounts switch mid-session.
Why it mattersSonnet 5.5 keeps $2/$10 pricing while claiming up to 30% lower cost for most work. Preserved thinking now stays within the originating org, so switching accounts mid-session forces new thinking.

Artificial Analysis benchmarks Claude Sonnet 5.5: 56 on its Intelligence Index, parity with Opus 5.5 on several agentic evals, but the heaviest output token use it has measured, raising cost per task to $7.60.
Why it mattersSonnet 5.5 nearly matches Opus 5.5 on agentic tasks but uses roughly 60% more output tokens, so cost per task is about 50% higher than Sonnet 5 at the same per-token price. Budget for tokens, not list price.
An index of the vibe-coding frontier. Corrections welcome.