Excited to see Qwen next-generation architecture, Qwen3.8-Flash-Next, out in the wild! Huge thanks to the Qwen team for building it, and to NVIDIA AI for the GPU resources, engineering support, and for featuring TokenSpeed in the blog. More to come. Keep building! 🚀 https://lnkd.in/gVcEE6CQ
⚡Meet Qwen3.8-Flash, a multimodal MoE and an early preview of the Qwen4 architecture, now open-weight! The production version Qwen3.8-Flash will be available soon via QwenCloud API at just $ 0.16/1M input tokens and $ 0.47/1M output tokens. 125B parameters + 51B N-gram embeddings, with just 6B activated per token. Unmatched cost-efficiency. What's new: 🔵 Next architecture: GDN + QSA hybrid attention, Gated Residual, N-gram Embedding & Muon optimizer, serving as a precursor to the architecture used in Qwen4. ⚪️ Dramatically lower training and inference costs: trained at just 1/9 the cost of Qwen3.7-Plus, while outperforming it across the board with especially strong gains in coding and office tasks. 🔵 Strong performance: scoring 58.7 on DeepSWE 1.1, 62.5 on SWE-bench Pro, 73.9 on CoWorkBench, 84.5 on AndroidWorld, and 95.7 on MathVision (with CI). ⚪️ 262K native context, extensible to 1M with YaRN. We’re also releasing the weights for Qwen3.8-Flash-Next, giving the community an early look at the new architecture we’re exploring for Qwen4. ✨ We can't wait to see what you build with Qwen3.8-Flash! - Blog: https://lnkd.in/g8CTfwsx - Tech Report:https://lnkd.in/gyU23Rp2 - Hugging Face:https://lnkd.in/guZ9UEje - ModelScope:https://lnkd.in/gGzKchMy