
Offers hard-won, specific engineering lessons — like the collapse pitfall and RL-against-real-environments approach — that can save practitioners from repeating the same mistakes when building or agentic/multimodal models.
“it aligns with our mission that we want to have intelligence with everyone”
Olive (MiniMax)
“it was trained multimodal from scratch. So, from step zero, we trained not only text data, we also trained image data.”
Olive (MiniMax)
“If someone tells me you can't do the thousand first thing, yeah, like I don't know, try harder.”
Dan (Together AI)
“It's like the type of thing that you should have done in your third year of undergrad or something like that, but most of us actually skipped that class, so now we're rediscovering it live in in industry.”
Dan (Together AI)
“we're we're seeing with models like M3 and GLM and Kimmy and and all those models that um the open-source frontier really can catch up”
Dan (Together AI)
video"My name is... my name is...": A Linguistic Map for Voice Agents — Midam Kim, ServiceNow
videoAct, Confirm, or Stop? Smarter behavior for AI assistants, wearables & robots — Amit Desai, Roku
videoSpeech-to-Speech Model Research at Google DeepMind — Valeria Wu Fon & Tom Ouyang, Google DeepMind
videoVoice Agents Can Just Do Things — Charlie Guo, OpenAI
videoWhy AI Agents Need Million-Token Context — Thomas Wolf & Olive Song, MiniMaxAI EngineerChecking sign-in…
Loading comments…