GSPO stabilizes long-horizon RL training for language models by optimizing at the sequence level rather than per-, directly addressing the model-collapse and instability failures teams hit when scaling GRPO. If you're doing RL or training reasoning models, it's a concrete algorithmic fix worth understanding before your next long run.
“However, we observe that existing RL algorithms (such as GRPO) exhibit severe instability issues during long training and lead to irreversible model collapse, hindering further performance improvements with increased compute.”
articleGRPO Beyond English: A Large-Scale Study of GRPO in Non-English and Multilingual SettingsKonstantin Dobler, Federico Scozzafava, Jonathan Janke, Mohamed Ali, Simon LehnererChecking sign-in…
Loading comments…