Vibeleaderboard
← All Intel
Intel / post

Eleven Sakana AI papers at ICML 2026, from sparse memory to agent benchmarks

Source
Sakana AI
Date
Sakana AI@SakanaAILabs
Thread · 5 parts

Sakana AI is heading to #ICML2026 in Seoul (July 6–11)! 🐟🇰🇷 Our team will present 11 papers spanning multi-agent coordination, sparse and efficient LLMs, test-time scaling, long-term memory, and agent benchmarks. A thread of everything we're presenting:

"How Small Can a Tandem Speech Front-End Be? Diagnosing Front-End Capacity with Layer Removal" will be presented at the Workshop on Machine Learning for Audio on July 10 at #ICML2026 Link: https://t.co/5KC0s6vln7 Recent tandem speech-to-speech (S2S) dialogue systems delegate knowledge and reasoning to a back-end LLM, leaving a front-end S2S transformer to handle low-latency spoken interaction. This raises a practical capacity question: how small can the front-end become, what performance is preserved, and how does its behavior change as layers are removed? To diagnose this, we randomly remove different numbers of layers from a KAME-style front-end transformer, fine-tune each variant on the same data, and evaluate on speech MT-Bench.

"UnMaskFork: Test-Time Scaling for Masked Diffusion via Deterministic Action Branching" will be presented at #ICML2026 Paper: https://t.co/4JC9SYdTyX We introduce UnMaskFork, a test-time scaling framework for Masked Diffusion Language Models (MDLMs). Using Monte Carlo Tree Search, it explores diverse generation paths by dynamically switching between multiple pre-trained MDLMs to collaboratively generate the text. Evaluations on coding and mathematical reasoning tasks show that UnMaskFork consistently outperforms standard Best-of-N and other tree search baselines. The results demonstrate that deriving search space diversity from multiple distinct models is a highly effective test-time scaling strategy for MDLMs.

"Feedback-to-Rubrics: Can We Extract Expert Criteria from Inline Comments?" will be presented at the Workshop on Human-AI Co-Creativity: Advances, Opportunities, and Challenges on July 11 at #ICML2026 Paper: https://t.co/MPfNX6l2z1 (Extended full-paper version) LLMs are increasingly used for writing and review support, but their usefulness depends on context-dependent criteria, such as expert preferences or organization-specific conventions, that are often tacit, undocumented, and difficult to elicit directly. We propose a problem setting for learning reusable natural-language rubrics from accumulated inline comments on artifacts such as human-written or LLM-generated drafts.

Read the full thread on X
Terms in this piece · Glossary
  • inference — Running a trained model to get answers — the phase where AI is actually used, as opposed to trained.
  • token — The chunk of text a model reads and writes in — roughly three-quarters of a word — and the unit AI usage is billed in.
  • context window — The maximum amount of text a model can consider at once — its working memory for the current conversation or task.
Why it matters

A single first-party index of the lab's ICML results, including sparse fast weights written at test time that hold about 75% five-needle accuracy at 128K , and a method for learning reusable rubrics from accumulated inline review comments.

More from Sakana AI
Recommended reads
Comments

Checking sign-in…

Loading comments…