Qwen3.8 Flash Next
Qwen/Qwen3.8-Flash-Next
Open Source · chat · open-weights
Open
Alert me on changes
Intelligence
Context
262.1K
Max output
131.1K
Weights
Open
API $/1M
$0.16 / $0.47
Modalities
text · image
Released
24 Aug 2026
Intelligence Index via Artificial Analysis · 0–100, higher is better
License: other · Qwen/Qwen3.8-Flash-Next
AI summary
● machine-written
Alibaba releases Qwen3.8-Flash-Next, a 125B mixture-of-experts model
Qwen3.8-Flash-Next is an open-weight multimodal model serving as an architecture preview for Qwen4, featuring 125 billion total parameters with only 6 billion activated per token. The model natively supports a 262,144-token context window and performs competitively with much larger models on coding and reasoning tasks. The production version, Qwen3.8-Flash, is available on QwenCloud at significantly lower pricing than flagship alternatives.
What's new
- 125B parameters with only 6B active per token using mixture-of-experts
- 51B n-gram embedding layer stored in system RAM rather than GPU memory
- Native 262,144-token context window, scalable to 1 million tokens
- Outperforms larger models on coding benchmarks (SWE-bench Pro: 62.5)
- Production API priced at $0.16 per million input tokens, $0.47 per million output tokens
Best for
Agentic coding tasks and bug fixingOffice and productivity workflowsCost-optimized deployments requiring long contextResearch on mixture-of-experts architectures