← All IntelIntel / video
Hugging Face Journal Club: Direct On-Policy Distillation
- Source
- youtube.com
- Author
- Hugging Face
- Date
Why it matters
Describes a cheaper route to RL-grade improvements on a large model by reusing the policy shift from a small RL-trained model as a dense reward.
Terms in this piece · Glossary
- distillation — Training a small, cheap model to imitate a big one's outputs, keeping much of the capability at a fraction of the cost.
Read the source www.youtube.com
More from Hugging Face
Recommended reads
Comments
Checking sign-in…
Loading comments…






