
Finding domains where correctness is machine-checkable is the bottleneck in building RL environments.
“the output of a single experiment can exceed what a scientist can safely store on a consumer laptop in many cases”
Kenny Workman
“just like code provided a verifiable substrate for complex software tasks that are not inherently verifiable, uh data analysis might do the same thing in bio”
Kenny Workman
“frontier models cannot be trusted to do real work. They're missing some capability between knowing biology and writing code.”
Kenny Workman
“cool thing about evaluation like coding is it forces you to reason about things more rigorously than you would when you're doing the thing yourself”
Kenny Workman
“We try to get the labs to compete um on the benchmarks cuz then it makes the models better at our products.”
Kenny Workman
videoData and Environment Curation for Post-Training LLMs — Mahesh Sathiamoorthy, Bespoke LabsAI Engineer
videoTrading Desks to Clinical Trials: Parallels in Applied Vertical AI — Ayush Bhardwaj, Allos AIAI Engineer
videoTraining Frontier Models to Out-Think Hackers — Uri Rolls, Arithmetic & Thom Wolf, Hugging FaceAI EngineerChecking sign-in…
Loading comments…