
Our Single-rollout Asynchronous Optimization (SAO), is able to train stably for one thousand steps and consistently outperform GRPO and its variants on agentic coding and reasoning benchmarks, such as SWE-Bench Verified, BeyondAIME, and IMOAnswerBench. https://t.co/oca1DvDvdB
Zhipu's SAO training method trains stably for a thousand steps and beats GRPO variants on , BeyondAIME, and IMOAnswerBench, a concrete RL technique for agentic and builders.
articleSERPO: Self-Evolving Rubric Policy Optimization for Open-Ended Test-Time Reinforcement LearningJianze Wang, Kunwang Zheng, Ying Liu, Yu Cao, Qilong Zhang, Jinlong Chen, Hua Yang, Qianglong Chen
articleRL Systems: Mind the Gap — Matching Trainer and Generator ThroughputKimbo ChenChecking sign-in…
Loading comments…