
From IcePop to KPop — our team keeps pushing on RL training stability for large MoE models. 👇 KPop replaces the fixed-ratio mask with an adaptive binary-KL region that matches each token's inherent noise. More robust updates, stable long-horizon agentic RL. Ring-2.6-1T → 76+ on SWE-bench Verified, pure RL. Congrats to @Jia__Guo & team! Blog:

-level mismatch in reinforcement learning is not uniform, and sizing the trust region per token instead of by a fixed ratio is what let a 1T model train stably to 76+ on with pure RL.
Checking sign-in…
Loading comments…