Qwen3.8-Flash multimodal MoE previews Qwen4 architecture

OpenSource APIs

TL;DR: Alibaba released open-weights Qwen3.8-Flash, a multimodal MoE with 6B activated parameters that outperforms Qwen3.7-Plus at 1/9 training cost.

Summary: Qwen3.8-Flash is a 125B-parameter multimodal MoE with 51B N-gram embeddings and just 6B activated per token, offering a 262K native context extensible to 1M. It introduces the GDN + QSA hybrid attention, Gated Residual, N-gram Embedding, and Muon optimizer as a precursor to Qwen4. Weights for Qwen3.8-Flash-Next were also released for early architecture exploration.

Why it matters: This gives AI builders a preview of Qwen's upcoming architecture with dramatically lower training and inference costs, plus strong coding, office, and multimodal benchmarks. Try it via open weights or the QwenCloud API at $0.16/1M input and $0.47/1M output tokens.

Source: x_com