
⚡️ 892 tokens/s — our 100B diffusion LLM, LLaDA2.1-flash, is now live on @ZenMuxAI! With Token Editing, LLaDA 2.1 goes from research breakthrough to production-ready speed. Diffusion models just got real. Try it via API or Chat 👇 https://t.co/8ObarWTPio #LLaDA #ZenMux #AI #dLLM https://t.co/zUUROipUAs
A 100B diffusion is now reachable through a hosted API, so its speed claims can be benchmarked against an autoregressive stack without standing up serving infrastructure.
Checking sign-in…
Loading comments…