Sakana AI released two new systems, Fugu Max and Fugu Ultra v2, that do not compete as single models but as orchestration layers sitting on top of a pool of open weight and specialized models, including NVIDIA's Nemotron. Instead of picking one closed frontier model and sending every request to it, Fugu routes each task to whichever model in its pool handles that kind of work best, mixing cheap models for simple steps with stronger ones for complex, multi step reasoning. Sakana claims the combined system matches or beats Opus 5 on its benchmarks at a fraction of the cost, and because the pool is swappable, the approach avoids locking a product to any single vendor's pricing or availability. This extends a wider pattern this year of orchestration layers competing directly with single frontier models rather than just wrapping them, betting that smart routing across many specialized models can outperform one large generalist on both cost and capability.

Introducing Fugu Max and Fugu Ultra v2: the next evolution of Sakana Fugu’s multi-agent orchestration system. Try: https://t.co/hhO6qT9YqD Blog: https://t.co/vjyNB5lWGp The frontier that actually matters is the Pareto frontier: capability on one axis, cost on the other. But the industry still treats it as a static menu of isolated models. Today we are resolving that with a dynamic architecture: Fugu Max expands the Pareto efficiency frontier. By orchestrating our largest pool of open-weights and specialized models to date, including NVIDIA Nemotron family, it dynamically routes tasks to the leanest capable model. Fugu Max delivers performance within striking distance of elite models at two to six times lower cost. Fugu Ultra v2 pushes the peak capability of orchestration higher than ever before. On Chartography, it outperforms Opus 5 and Fable 5. On DeepSWE, it outperforms models that cost three to five times more per token. Crucially, it does all of this without Fable 5, Fable 5.1, or GPT-6-Astra in its agent pool. The Fugu orchestration system evolved with model resiliency in mind. It does not rely on individual frontier models to deliver frontier output. By orchestrating a…
Checking sign-in…
Loading comments…