← All IntelClip / OtherConcrete recipe: build an RL environment for your specific use case
From Local Models: Trust, Control, Optimization — Carter Abdallah, NVIDIA · ≈24:08
“there's a reason why like uh chatbot isn't good at self-driving because like it's not trained on that.”
“it's sort of the the white pill for like the AI application builders and and the AI startups to actually have a a huge opportunity to build kind of their modes”
What’s in it
- Argues RL environments built for one use case beat generic chatbots
- Explains how production deployment creates a data flywheel for agents
- Frames domain specialization as the real moat for AI startups
Clip transcript
the the frontier models? >> Yeah, I think I like like I can start on this. Like I think the the power what that we've seen with a lot of different customers is is really kind of this idea that like if you want to make a specific use case work, like we can take the example of like if you want to figure out like a way that agents can actually automate your text. Like the the the most like concrete way you can do it today really is like build an RL environment for that use case. Like train on it and then deploy it into production with those users, right? Like let's say with a a million accountants that then now use this agent to ultimately get it towards full autonomy. It's a bit like almost like Tesla's levels towards full autonomy where like you kind of need to deploy it into like do the last mile of actually like training for that specific use case, but then also um deploying it to those specific users, right? So, like there's a reason why like uh chatbot isn't good at self-driving because like it's not trained on that. It's not deployed into that context, right? And like I think it's the same even for these specific like knowledge work use cases, where it's like if you want to have the perfect like financial agent, it's much more likely that you'll be able to get there if you have like our environment for that use case, if you deploy it into production, for example, as a bank, right? Like to millions of customers, than if you're there's like one got model chatbot. Like And I think this is sort of like what we've seen now with a lot of verticals and customers that like um there's like a huge unlock there to um really go into this like specialized domains, post train on them, deploy into them, and then continuously learn from production traces. So, we work with like some um also big AI natives on things like computer use, where ultimately having like millions of of traces from production data really can help you to to um continuously improve uh those those agents. And I think this kind of applies to almost every single domain, and I think it's sort of the the white pill for like the AI application builders and and the AI startups to actually have a a huge opportunity to build kind of their modes and and to get to this data flywheel um of like specialized models even in a broader sense. Like just going after like let's say computer use agents, right? And I think um yeah, this is something where I think um we're just seeing a lot of like movement especially now with like all models catching up to the frontier.
Comments
Sign in to comment.
Loading comments…