Sakana AI is heading to #ICML2026 in Seoul (July 6–11)! 🐟🇰🇷 Our team will present 11 papers spanning multi-agent coordination, sparse and efficient LLMs, test-time scaling, long-term memory, and agent benchmarks. A thread of everything we're presenting:

"How Small Can a Tandem Speech Front-End Be? Diagnosing Front-End Capacity with Layer Removal" will be presented at the Workshop on Machine Learning for Audio on July 10 at #ICML2026 Link: https://t.co/5KC0s6vln7 Recent tandem speech-to-speech (S2S) dialogue systems delegate knowledge and reasoning to a back-end LLM, leaving a front-end S2S transformer to handle low-latency spoken interaction. This raises a practical capacity question: how small can the front-end become, what performance is preserved, and how does its behavior change as layers are removed? To diagnose this, we randomly remove different numbers of layers from a KAME-style front-end transformer, fine-tune each variant on the same data, and evaluate on speech MT-Bench.

"UnMaskFork: Test-Time Scaling for Masked Diffusion via Deterministic Action Branching" will be presented at #ICML2026 Paper: https://t.co/4JC9SYdTyX We introduce UnMaskFork, a test-time scaling framework for Masked Diffusion Language Models (MDLMs). Using Monte Carlo Tree Search, it explores diverse generation paths by dynamically switching between multiple pre-trained MDLMs to collaboratively generate the text. Evaluations on coding and mathematical reasoning tasks show that UnMaskFork consistently outperforms standard Best-of-N and other tree search baselines. The results demonstrate that deriving search space diversity from multiple distinct models is a highly effective test-time scaling strategy for MDLMs.
"Feedback-to-Rubrics: Can We Extract Expert Criteria from Inline Comments?" will be presented at the Workshop on Human-AI Co-Creativity: Advances, Opportunities, and Challenges on July 11 at #ICML2026 Paper: https://t.co/MPfNX6l2z1 (Extended full-paper version) LLMs are increasingly used for writing and review support, but their usefulness depends on context-dependent criteria, such as expert preferences or organization-specific conventions, that are often tacit, undocumented, and difficult to elicit directly. We propose a problem setting for learning reusable natural-language rubrics from accumulated inline comments on artifacts such as human-written or LLM-generated drafts.

A single first-party index of the lab's ICML results, including sparse fast weights written at test time that hold about 75% five-needle accuracy at 128K , and a method for learning reusable rubrics from accumulated inline review comments.
Checking sign-in…
Loading comments…