← All IntelClip / AI ToolsData (including RL environments) is the real bottleneck for post-training
From Data and Environment Curation for Post-Training LLMs — Mahesh Sathiamoorthy, Bespoke Labs · ≈4:48
“Most of the places where people struggle especially enterprises is that they don't have access to good quality data and RLNs and this obviously also applies to frontier labs where they have all this infra setup and they are you know needing good quality RLMs right”
“for post training be it SFT or or uh reinforcement learning data is the bottleneck right”
What’s in it
- Argues data quality, not compute, is the real post-training bottleneck
- Reframes RL environments as just a new shape of training data
- Notes even frontier labs struggle to source good RL environments
Clip transcript
longer durations and for post training one of the popular techniques as you know is reinforcement learning and that's kind of um something you know a lot of you are excited about is the you know notion of RL environments but ultimately for post training be it SFT or or uh reinforcement learning data is the bottleneck right so when when I talk about data RLNs are also something I'm calling it as data it's just the data is now in a very different shape u again here you know compute is kind of well definfined models you know uh good sort of models exist and the uh infrastructure to post train for example u the There are various providers like fireworks, tinker or uh slime world and whatnot. So all all of those are somewhat well defined. Most of the places where people struggle especially enterprises is that they don't have access to good quality data and RLNs and this obviously also applies to frontier labs where they have all this infra setup and they are you know needing good quality RLMs right uh beyond so that
Comments
Sign in to comment.
Loading comments…