
"Beyond autoregressive: why diffusion is the future of language models." @StefanoErmon's keynote at @StartupGrind last week at the Fox Theatre. Your AI product doesn't make one model call per session. It makes thousands. Most run quietly under the hood. The work that keeps the agent moving. Using a frontier model for all of it means paying frontier pricing for work that doesn't need frontier intelligence. Mercury 2 is built for that work: >1,000 tokens/sec on standard GPUs, comparable quality to frontier speed-optimized models for a fraction of the cost. The question isn't which model is smartest. It's which model is most efficient, on the highest-volume tasks. @adityagrover_ @volokuleshov
Cost in an agentic product is driven by the thousands of routine calls, not the few hard ones. Tiering those calls to a cheap fast model is the lever, and this puts a concrete argument behind it.
Checking sign-in…
Loading comments…