👋 LLaDA-Image combines the LLaDA2.0-mini dLLM understanding backbone with a 6B DiT for high-quality image generation and instruction-guided editing. Both backone and Image-Gen are diffusion models, trained in a unified framework. 🔑 Image-first training at scale: Of ~220M cumulative generation-training samples, >90% use image-only supervision. Image-text pairs are introduced later for language alignment. The result is a model family that covers high-quality generation, instruction-guided editing, and fast 2-4-step inference with LLaDA-Image-Turbo. #OpenSource #inclusionAI #dLLM


An open-source stack where both understanding and generation are diffusion, not autoregressive, with a Turbo variant that does instruction-guided editing in 2 to 4 sampling steps.
Checking sign-in…
Loading comments…