Vibeleaderboard
← All Intel
Intel / post

KPop replaces fixed-ratio masking to stabilize agentic RL

Source
Ant Ling
Date
Ant Ling@AntLingAGI

From IcePop to KPop — our team keeps pushing on RL training stability for large MoE models. 👇 KPop replaces the fixed-ratio mask with an adaptive binary-KL region that matches each token's inherent noise. More robust updates, stable long-horizon agentic RL. Ring-2.6-1T → 76+ on SWE-bench Verified, pure RL. Congrats to @Jia__Guo & team! Blog:

Context

Reinforcement learning on a large model has to keep each training update close enough to the policy that generated the data, or training destabilizes; the usual way to enforce that is a KL-divergence constraint applied the same way to every . Ant Ling describes KPop as a successor to the team's earlier IcePop method that instead uses an adaptive binary-KL region sized to each token's own noise level, rather than one fixed ratio applied uniformly. The company reports this produced more robust updates and stable long-horizon agentic RL training, and that its Ring-2.6-1T model reached a score above 76 using pure reinforcement learning, without additional supervised credited for that result.

Terms in this piece · Glossary
  • mixture-of-experts — A model built from many specialist sub-networks where only a few activate per token, giving big-model capability at small-model running cost.
  • token — The chunk of text a model reads and writes in — roughly three-quarters of a word — and the unit AI usage is billed in.
  • SWE-bench — The standard benchmark for AI coding agents: real GitHub issues from real repositories, scored by whether the agent's patch passes the project's own tests.
  • fine-tuning — Taking a trained model and training it a bit more on your own examples so it gets better at one specific job.
More from Ant Ling
Recommended reads
Comments

Checking sign-in…

Loading comments…