Vibeleaderboard
← All Intel
Intel / video

Hugging Face Journal Club: Direct On-Policy Distillation

Source
youtube.com
Author
Hugging Face
Date
Why it matters

Describes a cheaper route to RL-grade improvements on a large model by reusing the policy shift from a small RL-trained model as a dense reward.

Terms in this piece · Glossary
  • distillation — Training a small, cheap model to imitate a big one's outputs, keeping much of the capability at a fraction of the cost.
Read the source www.youtube.com
More from Hugging Face
Recommended reads
Comments

Checking sign-in…

Loading comments…