What if an LLM could EDIT its own tokens in real-time, not just generate them? 🤯 Introducing LLaDA2.1 — a diffusion model that breaks from autoregressive dominance. It drafts fast, then fixes its own mistakes on the fly with Token-to-Token editing. The result? 892 tokens/sec on a 100B model. 🔥 ⚡ 892 TPS on HumanEval+ (coding) ⚡ 801 TPS on BigCodeBench 🧠 Real-time self-correction via T2T editing ✅ @lmsysorg SGLang Day 0 support — production-ready now A "non-consensus" architecture now challenging the mainstream. Open-sourced TODAY. 👇 #LLaDA #TokenEditing #OpenSource #LLM #dLLM

What's the secret? Our revolutionary Error-Correcting Editable (ECE) engine. 🧠 LLaDA2.1 doesn't just write; it drafts, then intelligently reviews and edits its own work, much like a human expert. This two-phase approach fixes a major weakness of early diffusion models — exposure bias and error accumulation.
Flexibility is key. LLaDA2.1 introduces a dual-mode design: ⚡ Speedy Mode for rapid iteration and prototyping 💎 Quality Mode for precision-critical production tasks You choose between speed and perfection based on your needs. From research to real-world applications.
We've also pioneered the first-ever large-scale Reinforcement Learning (RL) framework for a 100B-parameter diffusion model. Through EBPO (Block-level Policy Optimization), LLaDA2.1 learns to understand and follow complex human instructions with incredible fidelity. It doesn't just generate, it understands. 🤝

Diffusion LLMs now have a self-correction pass and day-zero SGLang support, so a 100B model reporting 892 per second becomes a real option to test for latency-bound coding workloads.
Checking sign-in…
Loading comments…