One update equation unifies evolution strategies and consensus optimizers
- Source
- Sakana AI
- Date
We are pleased to present our latest research at #ICML2026, “Bridging Spherical Black-Box Optimizers” https://t.co/3FT6vn0dSn When optimizing through simulators, external APIs, or in reinforcement learning, gradients are often unavailable. Black-Box Optimization (BBO) fills this gap, but the field has been historically split into two categories: 1. Parametric Methods: Algorithms like Evolution Strategies (ES) scale to high dimensions but only find a single solution. 2. Nonparametric Methods: Algorithms like Consensus-Based Optimization (CBO) find multiple solutions but fail in high dimensions. Our team asked a simple question: what if they are all doing the same thing? In our paper, we showed that these distinct families are actually variations of a single update equation. By bridging this theoretical gap, we can now engineer custom hybrid optimizers for specific tasks. A key application of this is merging foundation models. Building on our previous work in Evolutionary Model Merging, we faced a computational challenge. Evaluating large language models at every step is resource-intensive, but using a smaller evaluation dataset causes standard unimodal optimizers to overfit. By treating LLM merging as a multimodal problem and deploying our newly developed hybrid optimizers, AdaPol and SchedPol, we successfully navigated this issue. The algorithms identified multiple distinct optima on the smaller dataset, allowing us to find generalized, high-quality merges at a fraction of the compute cost.
"Bridging Spherical Black-Box Optimizers" will be presented at #ICML2026 Paper: https://t.co/3FT6vn0dSn Black-box optimizers can be used when gradients are unavailable. We show that multiple different such optimizers from different families can be written using a single unified update equation, and can use this update to design new optimizers. We then apply them to LLM model merging, reducing the training cost to a fraction. https://t.co/NzdWpxrmb0
- eval — A repeatable test for AI quality — a set of tasks plus scoring — used the way software teams use test suites, because model output is too variable to judge by eyeballing.
- multimodal — A model that works with more than text — reading images, audio, or video, and sometimes generating them too.
Gradient-free optimizers from separate families share one update equation, which lets you build hybrids that treat model merging as search and find generalising merges on small sets at a fraction of the compute.
articleBelief-Calibrated Optimization: An Explicit World Model for Agentic OptimizationYuhan Chen, Zhihua Tian, Mahavir Dabas, Charith Peris, Rahul Gupta, Ming Jin, Feiyang Kang, Siyuan Zhang, Nan Wang, Ruoxi Jia- postSheaf-ADMM: agents with partial views negotiate to a correct global answerSakana AI
articleEfficient MoE Training for Biological Foundation ModelsMichelle Horton
Checking sign-in…
Loading comments…


