Introducing Ring-2.5-1T-Zero: Scaling Zero RL to a 1-Trillion parameter model. We train this giant directly from base with NO extra human annotations, achieving competitive reasoning performance.

💡 The Bitter Lesson! On smaller models, we need complex format rewards to enforce reasoning. But at 1T scale, hand-crafted heuristics become completely redundant! 🚫🛠️ With a minimalist setup, the model spontaneously discover optimal strategies. Scale wins!
🌟 Spontaneous Cognitive Emergence! Without any templates, the 1T model developed advanced behaviors: 🧠 Anthropomorphism: Complaining, I might have a brain fart here ↔️ Parallel Reasoning: exploring alternate paths 😰 Context Anxiety: A strategic guess before limits

⚙️ 4-Stage Zero RL Pipeline: 1️⃣ Elicit: Token-level RL to grow CoT lengths 2️⃣ Distill: SFT to trim redundancy & reset train-infer gap 3️⃣ Refine: Sample-level RL for sustained improvement 4️⃣ Adaptive: Tiered training for cognitive routing

Format rewards that smaller models need fall away at 1T scale, and 100K of this model's reasoning traces into Qwen-32B beat 800K DeepSeek-R1 traces, a concrete signal that trace quality outweighs volume.
Checking sign-in…
Loading comments…