← All IntelClip / EducationClosing the loop: production issues become new training tasks
From RL Without Verifiable Rewards (Will Brown, Prime Intellect) · ≈18:05
“all of this is in spirit of making post-training easier, making continual learning easier, giving people the ability to create agents and not worry too much about having to to fuss with the research pieces and fine-tune all the the small details.”
“we're kind of seeing paths forward of how we start automating this more and more by like having environments as the anchor, which we can then spend compute on refining.”
“all of this is well allows us to ultimately close the loop where models are then able to stay within the guardrails we give them, they go find the issues in production, and then they turn these back into new tasks that can then be trained on for getting better in the real world.”
What’s in it
- Explains how coding agents could self-improve by mining production issues into new training tasks
- Argues compute can automate reward and environment design, cutting manual fine-tuning work
- Frames 'environments as the anchor' as the key lever for scaling continual learning
Clip transcript
information into its weights over time as well. Um and so all of this is in spirit of making post-training easier, making continual learning easier, giving people the ability to create agents and not worry too much about having to to fuss with the research pieces and fine-tune all the the small details. Today we still do, but we're kind of seeing paths forward of how we start automating this more and more by like having environments as the anchor, which we can then spend compute on refining. We can kind of use compute to mine the data we have from the real world to refine the the signals where the humans are kind of just in the same way that with coding agents we're kind of going to higher levels of abstraction, we can do this with environment and reward design as well. Um and all of this is well allows us to ultimately close the loop where models are then able to stay within the guardrails we give them, they go find the issues in production, and then they turn these back into new tasks that can then be trained on for getting better in the real world. Um we
Comments
Checking sign-in…
Loading comments…