Explains the R1 pipeline of pure RL with verifiable rewards, then SFT and preference tuning to recover general usability, which is directly useful for anyone building reasoning models.
Terms in this piece · Glossary
RLHF — Reinforcement learning from human feedback — training a model to prefer answers humans rate as better, which turns a raw text predictor into a usable assistant.