
Our latest post explores on-policy distillation, a training approach that unites the error-correcting relevance of RL with the reward density of SFT. When training it for math reasoning and as an internal chat assistant, we find that on-policy distillation can outperform other approaches for a fraction of the cost.

On-policy combines RL's error correction with the dense reward signal of supervised , and Thinking Machines reports it beating other approaches on math reasoning and an internal assistant for a fraction of the training cost.
Checking sign-in…
Loading comments…