👋 Introducing LLaDA2.0-Uni — the 💥first unified MoE multimodal model in the LLaDA2.0 series of @TheInclusionAI, designed for native multimodal understanding and generation. #dLLM #inclusionAI #LLaDA #Multimodal #Opensource #DiffusionModels 1/2 Key features: 🧠Chain-of-Thought Generation LLaDA2.0-Uni doesn't just generate images — it thinks first. Through reasoning-augmented training, the model performs step-by-step reasoning before visual generation. Achieves 0.78 on WISE-Bench with thinking mode.
2/2 Key features: 🔄Interleaved Generation & Reasoning Unified discrete representations enable seamless interleaved text-image sequences. From story telling to science problems with step-by-step visuals — all in one coherent generation flow.
🔬 Built on three core model designs: 1️⃣ LLaDA2.0 Backbone for Unified Discrete Modelling: Leveraging the MoE dLLM architecture from LLaDA2.0, it formulates both multimodal understanding and generation as a unified block-wise mask prediction paradigm. 2️⃣ Fully Semantic Visual Tokens: Departing from traditional VQ methods that focus on reconstruction, LLaDA2.0-Uni transforms image inputs into purely semantic discrete tokens. 3️⃣ Diffusion Decoder for High-Quality Generation: Leveraging a custom Diffusion Decoder optimized for semantic tokens and few-step distillation, LLaDA2.0-Uni delivers high-fidelity image generation in just 8 steps of ultra-fast inference.

🔧Try out and explore how #LLaDA understanding and generate the world. 📄 Technical Report: https://t.co/DxlLxJ3zMy 🐙 Code: https://t.co/sX0lALW2AA 🤗 Model: https://t.co/n15P5dO7jS
One model covering interleaved text and image generation, with a reasoning pass before drawing and an 8-step decoder, collapsing the usual understand-then-generate pipeline into a single system you can run.
Checking sign-in…
Loading comments…