⚡️ Step 3.5 Flash is coming: Fast Enough to Think. Reliable Enough to Act! We’re dropping our most capable open-source foundation model yet. Frontier reasoning meets extreme efficiency. It leverages a sparse Mixture of Experts (MoE) architecture, 196B total → 11B active. Key Capabilities: ✅Reasoning at Speed: MTP-3 powered throughput at 100–300 tok/s (350 tok/s peak for single-stream coding tasks). ✅Agentic Power: ⚡️ 74.4% SWE-bench Verified ⚡️ 51.0% Terminal-Bench 2.0. Proven stability for complex, long-horizon tasks. ✅256K Efficient Context: 3:1 SWA ratio + Full Attention. Massive datasets or long codebases support with minimal overhead. Consistent performance, hybrid efficiency. ✅Local-First Deployment: Optimized for Mac Studio M4 Max, NVIDIA DGX Spark. Secure, private, and frontier-capable. Your data, your hardware, your agent. You can try Step 3.5 Flash right now: 👉 OpenRouter: https://t.co/MRDG3R4F64 👉 GitHub: https://t.co/qqmpslxK9t 👉 HuggingFace:https://t.co/y70Qzsc8DE 👉 Blog:https://t.co/rMH67kYmvA 👉 ModelScope: https://t.co/T4W7we5GEr 🌌 The Next:Step 4 training is officially LIVE! We're calling on the world's boldest builders to co-creat the Step 4 right now. Let's define the Agentic Era together! Join our Discord:

huggingface🤗 mtp3_bf16: https://t.co/0QIM32AlWz mtp3_fp8: https://t.co/5sIaLQ9Ekd int4: https://t.co/PfwfxQLGtK
A 196B activating 11B parameters posts 74.4% and 51.0% Terminal-Bench 2.0 at 100 to 300 tokens per second, with and local deployment on Mac Studio M4 Max and DGX Spark.
article👏🏻Congratulations!Step3-VL-10B was selected for HuggingFace Daily Papers…
postStep-Audio-R1.1 opens weights for a speech model that reasons in real timeChecking sign-in…
Loading comments…