Gives builders a defensible rule for when routing to a cheaper model actually saves money and when it silently costs more. The KV-cache and arguments are directly applicable to any long-running coding .
“we're reducing the cost of Fable level intelligence by 40%. The way we do that is we allow Fable to still do like the planning and the the hard decision making”
Walden Yan
“Like if you run terminal bench on Opus and Haiku, like Opus will do about three times better at 1/10 the cost of Haiku, even though Haiku's significantly cheaper per token.”
Alex Atallah
“you look at like the top model being used by dollar spent on classification tasks, well, guess what it is. It's Opus.”
Alex Atallah
“I would like never recommend using like these models past like 200K tokens, under 100K if you can.”
Walden Yan
“you kind of just always have this like main frontier agent that's watching, even if it's not the one doing the work.”
Walden Yan
videoOperating Distributed Inference Systems at Scale — Nishant Gupta & Naman Ahuja, Meta
videoVertical Mobility: Inference from MVP to Trillion-Parameter Workloads — Sitanshu Gupta, CoreWeave
videoAre LLM Performance Benchmarks Reliable? — Ashok Chandrasekar & Jason Kramberger, Google
videoRouting LLM Inference in Production: From Engine Signals to Policy — Qianru Lao & Lu Zhang, OpenAIChecking sign-in…
Loading comments…