Zero RL at trillion scale, with no format rewards required
- Source
- Ant Ling
- Date
Introducing Ring-2.5-1T-Zero: Scaling Zero RL to a 1-Trillion parameter model. We train this giant directly from base with NO extra human annotations, achieving competitive reasoning performance.

💡 The Bitter Lesson! On smaller models, we need complex format rewards to enforce reasoning. But at 1T scale, hand-crafted heuristics become completely redundant! 🚫🛠️ With a minimalist setup, the model spontaneously discover optimal strategies. Scale wins!
🌟 Spontaneous Cognitive Emergence! Without any templates, the 1T model developed advanced behaviors: 🧠 Anthropomorphism: Complaining, I might have a brain fart here ↔️ Parallel Reasoning: exploring alternate paths 😰 Context Anxiety: A strategic guess before limits

⚙️ 4-Stage Zero RL Pipeline: 1️⃣ Elicit: Token-level RL to grow CoT lengths 2️⃣ Distill: SFT to trim redundancy & reset train-infer gap 3️⃣ Refine: Sample-level RL for sustained improvement 4️⃣ Adaptive: Tiered training for cognitive routing

Context
'Zero RL' refers to training a directly from a pretrained base using reinforcement learning, without a supervised fine-tuning stage on human-written examples first. Ant Ling introduces Ring-2.5-1T-Zero, which the company describes as scaling this Zero RL approach to a trillion-parameter model for the first time. Ant Ling reports that on smaller models, complex reward shaping is needed to push a model toward structured reasoning, but that at 1 trillion parameters the model developed reasoning behaviors on its own with a simpler reward setup, including what the company characterizes as spontaneous parallel exploration of alternate solution paths.
The company describes a four-stage training pipeline: growing length with token-level reinforcement learning, distilling that into a more concise model with supervised fine-tuning, refining further with sample-level reinforcement learning, and finally routing between reasoning depths adaptively. Ant Ling reports the resulting model solves AIME problems using less than half the of baseline models it compares against, and that distilling 100,000 of its own reasoning traces into a much smaller Qwen-32B model outperformed using 800,000 traces from DeepSeek-R1. These are the company's own reported comparisons; the arXiv paper's full experimental setup was not independently reviewed for this explanation.
- reasoning model — A model trained to think — generating extended internal reasoning before answering — trading time and tokens for accuracy on hard problems.
- chain-of-thought — Having a model write out intermediate reasoning steps before its answer, which markedly improves performance on hard problems.
- token budget — A cap on how many tokens a task, session, or agent run may consume — the practical control on both cost and how long an agent will grind.
- distillation — Training a small, cheap model to imitate a big one's outputs, keeping much of the capability at a fraction of the cost.
Checking sign-in…
Loading comments…






