Ling and Ring 2.6 technical report lands with two open base checkpoints
- Source
- Ant Ling
- Date
Ling & Ring 2.6 technical report is out, with two open-weight base models. We co-design model + system across architecture, training, and agentic capability: • 7:1 hybrid linear attention • KPop for stable agentic RL: SWE-bench Verified 76.28% • ~4× token efficiency

Architecture · Hybrid Linear Attention GQA attention becomes costly and fragile at 256K context. Ling 2.6 makes 256K practical through a hybrid design: 7 Lightning Attention layers + 1 MLA layer.

RL Training · From IcePop to KPop IcePop (Ring-1T) used a uniform KL constraint for stable RL on MoE. But token-level mismatch is heterogeneous, which means rare tokens need wider tolerance. KPop replaces the uniform bound with adaptive Binary KL. As a result, our SWE-bench Verified improved from 70.8% → 76.28%, with pure RL.

Efficiency · Intelligence Density Over Token Count Ling-2.6-1T reaches AAII 34 with ~16M output tokens — comparable to GPT-5.4 in non-reasoning mode, with ~4× higher token efficiency than Ling-2.0-1T. Methods: Evo-CoT, LPO, Bidirectional RLHF, Shortest-correct Distillation.

Context
Ant Ling published a technical report for the Ling and Ring 2.6 model family, alongside two open base checkpoints. On architecture, the company says standard attention becomes costly and fragile at 256,000 tokens of context, and that Ling 2.6 instead stacks seven Lightning Attention layers to one MLA layer, which it says makes that context length practical. On training, the report credits KPop, the adaptive per-token KL method the team described the prior week, with lifting from 70.8% to 76.28% using pure reinforcement learning.
The company also reports Ling-2.6-1T reaches an Artificial Analysis Intelligence Index score of 34 using about 16 million output tokens, which it describes as roughly four times more token-efficient than its earlier Ling-2.0-1T model at comparable capability, achieved through methods it names but does not detail in the thread: Evo-, LPO, bidirectional and shortest-correct distillation. Ant Ling released base checkpoints for Ling-2.6-flash and Ling-2.6-1T, saying they are meant for research use rather than as production-ready instruct models; the report's full methodology was not reviewed for this explanation.
- SWE-bench — The standard benchmark for AI coding agents: real GitHub issues from real repositories, scored by whether the agent's patch passes the project's own tests.
- chain-of-thought — Having a model write out intermediate reasoning steps before its answer, which markedly improves performance on hard problems.
- RLHF — Reinforcement learning from human feedback — training a model to prefer answers humans rate as better, which turns a raw text predictor into a usable assistant.
- mixture-of-experts — A model built from many specialist sub-networks where only a few activate per token, giving big-model capability at small-model running cost.
Checking sign-in…
Loading comments…






