
Reliability over long autonomous runs is an environment-and-data problem before it is an algorithm problem.
“what is it that's blocking the uh autonomy of agents. It's basically reliability, right?”
Mahesh Sathiamoorthy
“for post training be it SFT or or uh reinforcement learning data is the bottleneck”
Mahesh Sathiamoorthy
“RLNs are also something I'm calling it as data it's just the data is now in a very different shape”
Mahesh Sathiamoorthy
“the stronger teachers are not always the best uh uh stronger models are not always the better teachers”
Mahesh Sathiamoorthy
“in in this whole process of building this open thoughts agent SFT still contributed a lot to the gains”
Mahesh Sathiamoorthy
Checking sign-in…
Loading comments…