
🚀 Meet LLaDA2.2-flash — the first agent-oriented MoE Diffusion LLM. Levenshtein editing—deletion, insertion, and substitution—enables self-correction during decoding. A more elegant solution to model collapse than forcing diffusion back into left-to-order generation. Performance: ⚡ 703.82 TPS on BFCL-V4 function calling 519.0 TPS on SWE-bench Verified coding ⚡ 1.64x BF16 throughput vs autoregressive models Built for high-throughput, low-latency AI agents. 🔗 GitHub: https://t.co/3TJUYNSLdT 🤗 Hugging Face: https://t.co/8OrHXBiGp2 📄 Technical Report:https://t.co/liV6nOk8YU #LLaDA #OpenSource #DiffusionLLM #AIAgents

Diffusion decoding with insert, delete and substitute edits offers a non autoregressive route to high throughput , with a reported 1.64x BF16 throughput advantage and to verify it.
Checking sign-in…
Loading comments…