
If you rely on public coding leaderboards to pick or trust a model, this piece exposes specific ways those benchmarks can be gamed or mis-graded — useful for calibrating skepticism before you build pipelines of your own.
“I have not met an engineer in the last 6 months that would choose a model or choose um an LLM based on the leaderboards.”
Ali
“Leaderboards are what we see in benchmarks today. They tell you who wins, but they don't to you why.”
Ali
“So, instead of actually trying to fix the to to apply a patch to a task, they try to go and find dot git folders”
Ali
“Um it is one thing to have a test a task that is failing the LLM proven that the LLM is not there yet.”
Ali
“Um benchmarks are not hard. We need to look under the hood.”
Ali
video"My name is... my name is...": A Linguistic Map for Voice Agents — Midam Kim, ServiceNow
videoAct, Confirm, or Stop? Smarter behavior for AI assistants, wearables & robots — Amit Desai, Roku
videoSpeech-to-Speech Model Research at Google DeepMind — Valeria Wu Fon & Tom Ouyang, Google DeepMind
videoVoice Agents Can Just Do Things — Charlie Guo, OpenAIChecking sign-in…
Loading comments…