Vibeleaderboard
← All Intel
Intel / post

Ring-2.6-1T Ships With a Reasoning-Effort Dial for Agent Workloads

Source
Ant Ling
Date
Ant Ling@AntLingAGI
Thread · 5 parts

We are launching Ring-2.6-1T, a trillion-parameter flagship thinking model engineered for real-world complex tasks and production env: 🚀 - Adjustable Thinking Effort: dynamic compute mechanism to flexibly balance cognitive depth, token cost, and execution speed; - Agent-Optimized: Built for high-frequency workflows, delivering rapid multi-step execution and tool orchestration with SOTA stability; - Deep Thinking: Unlocks the model's maximum capability ceiling for rigorous mathematical logic and scientific research;

1/4 Not all tasks need equal compute Format conversion differs vastly from math olympiads.🧑‍🔬 Ring 1T delivers lower token overhead and rapid multi-step execution, making it the ideal production default for tool orchestration, coding, and multi-turn interactions. Benchmarks🫶

2/4 Agentic "high-ness" In real-world task execution benchmarks, Ring-2.6-1T 'high' demonstrates top-tier stability and API routing capabilities, suitable for general and coding agent use cases: - PinchBench: 87.60 (outperforming GPT-5.4 xHigh & Gemini-3.1-Pro high) @kilocode - ClawEval: 63.82 - Tau2-Bench Telecom: 95.32 📊

3/4 Intelligence "overload-xhigh" Built to provide thought space needed for rigorous logical analysis,🧠'xhigh' unlocks our highest capability ceiling for math, research, and multi-path exploration, suitable for planning and reasoning heavy use cases: - AIME 26: 95.83 - GPQA Diamond: 88.27 - ARC-AGI-V2: 77.78

Read the full thread on X

Context

Ant Ling launches Ring-2.6-1T, a trillion-parameter from InclusionAI built with an adjustable compute setting rather than one fixed reasoning depth. The company reports that in its faster 'high' mode, meant for everyday and coding workflows, the model scored 87.60 on PinchBench and 95.32 on the Tau2-Bench Telecom agent ; in its deeper 'xhigh' mode, meant for math and research tasks needing more thinking space, it reports 95.83 on AIME 26 and 88.27 on GPQA Diamond. The company frames the split as a way to avoid paying the token cost of deep reasoning on tasks, like format conversion, that do not need it.

These figures are the company's own reported benchmark runs, not independently reproduced here, and the post does not specify how each mode's compute budget is set beyond the general high and xhigh labels. Ant Ling made the model free on OpenRouter for one week and said open-source weights were coming soon; it followed through six days later, releasing the weights with the same two-mode design.

Terms in this piece · Glossary
  • reasoning model — A model trained to think — generating extended internal reasoning before answering — trading time and tokens for accuracy on hard problems.
  • AI agent — An AI system that doesn't just answer once but works toward a goal in a loop — taking actions, reading the results, and deciding what to do next.
  • benchmark — A standard public test set for comparing AI models — the shared scoreboards behind every "model X beats model Y" claim.
  • open weights — A model whose trained parameters are published for anyone to download and run — unlike API-only models you can access but never possess.
More from Ant Ling
Recommended reads
Comments

Checking sign-in…

Loading comments…