
Great to see Step 3.7 Flash live on @FireworksAI_HQ. Designed for inference from day one, Step 3.7 Flash combines a hardware-friendly architecture with MTP-assisted decoding to reach up to 400 tokens/s. Fast, multimodal, and ready to power capable agents in real-world workflows.
Serving a 198B at roughly 400 output per second changes the latency budget for loops, and the stated architecture explains where the speed comes from instead of just asserting it.
article👏🏻Congratulations!Step3-VL-10B was selected for HuggingFace Daily Papers…
postStep-Audio-R1.1 opens weights for a speech model that reasons in real timeChecking sign-in…
Loading comments…