⚡Meet Qwen3.8-Flash, a multimodal MoE and an early preview of the Qwen4 architecture, now open-weight! The production version Qwen3.8-Flash will be available soon via QwenCloud API at just $ 0.16/1M input tokens and $ 0.47/1M output tokens. 125B parameters + 51B N-gram embeddings, with just 6B activated per token. Unmatched cost-efficiency. What's new: 🥳 - Next architecture: GDN + QSA hybrid attention, Gated Residual, N-gram Embedding & Muon optimizer, serving as a precursor to the architecture used in Qwen4. - Dramatically lower training and inference costs: trained at just 1/9 the cost of Qwen3.7-Plus, while outperforming it across the board with especially strong gains in coding and office tasks. - Strong performance: scoring 58.7 on DeepSWE 1.1, 62.5 on SWE-bench Pro, 73.9 on CoWorkBench, 84.5 on AndroidWorld, and 95.7 on MathVision (with CI). - 262K native context, extensible to 1M with YaRN. We’re also releasing the weights for Qwen3.8-Flash-Next, giving the community an early look at the new architecture we’re exploring for Qwen4.🚀 We can't wait to see what you build with Qwen3.8-Flash!👀👇 - Blog: https://t.co/M5hYypFLgJ - Technical Report: https://t.co/IF0gObIkQO - Hugging Face: https://t.co/6ow8QVAABt - ModelScope:

Model Architecture Four core upgrades for maximum capability, efficiency, capacity, and stability: - Attention: GDN + QSA Hybrid. Gated DeltaNet (GDN) compresses history. Qwen Sparse Attention (QSA) uses a lightweight indexer for micro-block context selection. Lower the cost of attention on long sequences. - Residual: Gated Residual (GR) widens the residual stream to 4 branches with a dynamic read and write gating, strengthening cross-layer information flow and significantly improving training stability. - Embedding: N-gram Embedding uses local context lookups to expand model capacity at minimal compute cost, while keeping the embedding table in host memory with asynchronous prefetching. - Optimization: Muon optimizer. Refines Muon through improved orthogonalization, smarter parameter assignment between Muon and AdamW, and fused-parameter splitting, with scaling laws refitted for the new architecture.

At a 1M-token context length, QSA’s attention kernel is up to 7.6× faster in prefill and 4.9× faster in decode. With a 90% prefix-cache hit rate, Qwen3.8-Flash-Next delivers 8.6× the prefill throughput of Qwen3.7-Plus.

for a 125B with only 6B active per token, priced around $0.16/$0.47 per 1M tokens and claiming 62.5 on Pro — plus the first concrete look at the architecture Qwen4 will be built on.
Checking sign-in…
Loading comments…