Alibaba Cloud logo Open weights · Qwen Community 1.0

Qwen3.8-Flash-Next

Alibaba's open-weight preview of the architecture it says will underpin Qwen4: a 125B mixture-of-experts model with 6B active per token, under the Qwen Community 1.0 licence.

Key Specifications

Context window
262K tokens
Max output
131K tokens
Released
Parameters
125B, 6B active
Inputs
Text, image, video
Capabilities Show details
Reasoning Thinks before it answers, always on or as a switchable mode. Function calling Connect to external tools, APIs, and systems. Structured output Return responses in structured formats like JSON.

Hosted API pricing

Provider Input / 1M tokens Output / 1M tokens Cost at 10M in + 2M out
Alibaba Cloud logo Alibaba Cloud Creator Open weights, no hosted price View

Heads up: Base-tier, on-demand rates per 1M tokens; cached, batch and long-context tiers excluded. A provider may serve a shorter context or a quantized build than the creator's release. Verify before provisioning. More on how we price.

Estimated cost to self-host

Qwen3.8-Flash-Next needs about 72 GB of GPU memory at 4-bit with 32K context. The cheapest rental that fits is 8× RTX 3060 at about $288 a month.

Precision Cheapest, 32K context Cheapest, full 262K context
4-bitINT4 / FP4
Nvidia logo 8× RTX 3060 72 GB VRAM · $288/mo
Nvidia logo 8× RTX 3060 77 GB VRAM · $288/mo
8-bitFP8 / INT8
Nvidia logo 8× RTX 3090 128 GB VRAM · $806/mo
Nvidia logo 8× RTX 3090 133 GB VRAM · $806/mo
16-bitFP16 / BF16
Nvidia logo 8× RTX A6000 253 GB VRAM · $3,168/mo
Nvidia logo 8× RTX A6000 258 GB VRAM · $3,168/mo

Estimates based on median on-demand rates for Nvidia GPUs. Memory is weights plus KV cache for one request, using FP8 KV cache where supported and FP16/BF16 otherwise. No guarantee of runtime support, usable performance, or that a matching quantized build exists. How we estimate costs.

Frequently Asked Questions

What is Qwen3.8-Flash-Next good for?

Text, image and video agent and coding work you run yourself, with thinking on by default and a reasoning effort setting. Qwen Community 1.0 weights, on a multi-GPU node.

When is Qwen3.8-Flash-Next not a good fit?

Audio input or image output, and there is no built-in web search, though a host or your own tool loop can add one. It takes a multi-GPU node, and Alibaba calls it an experimental preview, with the hosted Qwen3.8-Flash carrying the production features instead.

What is the cheapest way to run Qwen3.8-Flash-Next?

No provider we track hosts it; self-hosting on the cheapest rental that fits (8x RTX 3060, $0.40 per hour) costs about $288 a month.

Can I self-host Qwen3.8-Flash-Next?

Yes. The weights are Qwen Community 1.0 licensed. At 4-bit it needs about 72 GB of GPU memory, which starts at roughly $288 a month on the cheapest rental that fits.

More from Alibaba Cloud

Model Context Input / 1M Output / 1M
Alibaba Cloud Qwen3-Coder-Plus 1M $1.00 $5.00
Alibaba Cloud Qwen3.5-Flash 1M $0.10 $0.40
Alibaba Cloud Qwen3.5-Plus 1M $0.40 $2.40
Alibaba Cloud Qwen3.6-Flash 1M $0.25 $1.50
Alibaba Cloud Qwen3.6-Plus 1M $0.50 $3.00
Alibaba Cloud Qwen3.7-Flash 1M $0.03 $0.13
Alibaba Cloud Qwen3.7-Max 1M $1.25 $3.75
Alibaba Cloud Qwen3.7-Plus 1M $0.28 $1.10