🚀10B parameters, 200B+ performance! Introducing STEP3-VL-10B : our open-source SOTA vision language model. 🥊 At just 10B, it redefines efficiency by matching or exceeding the capabilities of 100B/200B-scale models. SOTA Performance : ✅STEM/Multimodal: Outperforms GLM-4.6V (106B-A12B) and Qwen3-VL (235B-A22B) on MMMU, MathVision, MathVerse,etc. ✅Math: Near-perfect scores on AIME 24/25, achieving elite-level reasoning. ✅2D/3D Spatial Understanding: Beating same-scale models on BLINK/CVBench/OmniSpatial. ✅Coding:Dominates LiveCodeBench in real-world dynamic programming. Key Breakthroughs: 🔹 1.2T token full-parameter pre-training. 🔹 1,400+ RL iterations for superior reasoning. 🔹 Innovative PaCoRe tech for dynamic compute allocation. 💡We prove that scale isn't everything. With high-quality targeted data and systematic post-training, a 10B model can go toe-to-toe with the industry's largest giants." 📱Complex AI reasoning is now accessible for every device. Check out the Base & Thinking versions on HuggingFace now! Homepage:https://t.co/xn82bwwTzD Paper:https://t.co/flxKXRbCBW HuggingFace:https://t.co/6YV5wF6iai ModelScope:

ModelScope playground: https://t.co/MCI2UBWEAT
A 10B vision reporting results at or above GLM-4.6V and Qwen3-VL on MMMU and MathVision puts strong reasoning within reach of single-device deployment, with weights and a paper available to verify.
article👏🏻Congratulations!Step3-VL-10B was selected for HuggingFace Daily Papers…
postStep-Audio-R1.1 opens weights for a speech model that reasons in real timeChecking sign-in…
Loading comments…