Alibaba Cloud logo

Qwen3.6-35B-A3B

Open weights · Apache 2.0

The first and smaller of the two open Qwen3.6 checkpoints, a 35B MoE with 3B active across 256 experts pairing Gated DeltaNet linear attention with sparse experts, under Apache 2.0.

Cheapest of 2 providers, via Novita
Input
$0.25 / 1M tokens
Output
$1.49 / 1M tokens

About $5.48 for 10M input and 2M output tokens. Estimate yours

Novita Alibaba Cloud 2 providers

Key Specifications

Context window
262K tokens
Max output
66K tokens
Inputs
Text, Image, Video Outputs: Text

Qwen3.6-35B-A3B pricing by provider

Provider Input / 1M tokens Output / 1M tokens Cost at 10M in + 2M out
Novita logo Novita Cheapest $0.25 $1.49 $5.48 View
Alibaba Cloud logo Alibaba Cloud Creator $0.38 $2.25 $8.30 View

Heads up: We do our best to keep these specs & prices accurate. However, cloud costs may fluctuate based on region, usage, and other factors not listed here. These are estimates based on common setups and are for informational purposes only. Always verify current rates & exact specs with the provider before provisioning. LLM rates are base-tier, on-demand prices per 1M tokens; cached-input, batch and long-context tiers are not included.

Compare every model at this volume in the LLM cost calculator.

Capabilities

Function calling

Function calling

Connect to external tools, APIs, and systems.

Structured output

Structured output

Return responses in structured formats like JSON.

Estimated cost to self-host Qwen3.6-35B-A3B

Qwen3.6-35B-A3B has 35B parameters, about 3B active per token. At 4-bit it needs about 21 GB of GPU memory, at 8-bit 42 GB and at BF16 84 GB, counting 20% on top of the weights for KV cache and runtime overhead.

Precision Memory needed Cheapest rentals that fit Per month, 24/7
4-bit 21 GB Nvidia logo RTX 3090 ($0.13/hr, tight) Nvidia logo RTX 3090 Ti ($0.19/hr, tight) $94
8-bit 42 GB Nvidia logo RTX A6000 ($0.56/hr, tight) Nvidia logo A40 ($0.74/hr, tight) $403
BF16 84 GB Nvidia logo RTX PRO 6000 ($2.19/hr, tight) Nvidia logo GH200 ($2.52/hr, tight) $1,577

$0.13 an hour (RTX 3090) buys about 285K tokens an hour at the cheapest hosted rate we track ($0.25 in / $1.49 out per 1M tokens on Novita, 5:1 input to output); below that volume the API is cheaper, before idle time. Compare hosted costs.

Memory is parameters × bytes per weight at each precision, plus 20% for KV cache and runtime overhead. Rentals are the cheapest cards that hold it, at provider-weighted median on-demand prices for the week of August 17, 2026, consumer cards included, in nodes of up to 8 GPUs; a fit with under 15% headroom is marked tight. A month is 720 hours. See the best-value GPUs guide for the same table across models, and cloud GPU pricing for every card.

More from Alibaba Cloud

Model Context Input / 1M Output / 1M
Alibaba Cloud logo Qwen3.6-Flash 1M $0.25 $1.50
Alibaba Cloud logo Qwen3.5-35B-A3B Open weights 262K $0.25 $2.00
Alibaba Cloud logo Qwen-MT-Plus 16K $0.25 $0.75
Alibaba Cloud logo Qwen3-235B-A22B Thinking (2507) Open weights 262K $0.30 $3.00
Alibaba Cloud logo Qwen3-VL-235B-A22B Instruct Open weights 262K $0.30 $1.50
Alibaba Cloud logo Qwen3.5-27B Open weights 262K $0.30 $2.40
Alibaba Cloud logo Qwen3-VL-30B-A3B Instruct Open weights 262K $0.20 $0.70
Alibaba Cloud logo Qwen3-VL-Plus 262K $0.20 $1.60

Models near this price

The models nearest this one by input rate, five each way, at each one's cheapest listed provider.

Frequently Asked Questions

How much does Qwen3.6-35B-A3B cost?

Qwen3.6-35B-A3B costs $0.25 per 1M input tokens and $1.49 per 1M output tokens via Novita, the cheapest of 2 providers listing it. The highest rate is Alibaba Cloud at $0.38 in / $2.25 out. Alibaba Cloud publishes the weights, so it can also be run on your own hardware.

What does Qwen3.6-35B-A3B cost for 10M input and 2M output tokens?

At Novita's rates, 10M input tokens and 2M output tokens cost about $5.48: $2.50 for input and $2.98 for output. Input is prompts and context, output is what the model writes back; a workload that generates more than it reads shifts the cost toward the output rate of $1.49 per 1M tokens.

Which providers offer Qwen3.6-35B-A3B?

2 providers list Qwen3.6-35B-A3B: Novita ($0.25 in / $1.49 out) and Alibaba Cloud ($0.38 in / $2.25 out). Rates are per 1M tokens in USD, cheapest input rate first.

What is Qwen3.6-35B-A3B's context window?

Qwen3.6-35B-A3B accepts up to 262K tokens of input per request and returns up to 66K tokens per response. The context window is the prompt plus any documents, conversation history and tool results sent with it; every token in it is billed at the input rate.

What inputs and outputs does Qwen3.6-35B-A3B support?

Qwen3.6-35B-A3B accepts text, images and video as input and produces text. Its listed capabilities are function calling and structured output.

Can I self-host Qwen3.6-35B-A3B?

Yes. Alibaba Cloud publishes Qwen3.6-35B-A3B's weights under the Apache 2.0 license. At 4-bit it needs about 21 GB of GPU memory; the cheapest rental that holds it is RTX 3090 at $0.13 per hour, about $94 a month. $0.13 an hour (RTX 3090) buys about 285K tokens an hour at the cheapest hosted rate we track ($0.25 in / $1.49 out per 1M tokens on Novita, 5:1 input to output); below that volume the API is cheaper, before idle time. The estimate above prices 8-bit and BF16 too.

How does Qwen3.6-35B-A3B compare with Qwen3.6-Flash?

Qwen3.6-35B-A3B and Qwen3.6-Flash both cost $0.25 per 1M input tokens ($1.49 vs $1.50 per 1M output tokens). The context window is 262K tokens against 1M.

What are cheaper alternatives to Qwen3.6-35B-A3B?

Models from other creators with a lower input rate and at least Qwen3.6-35B-A3B's 262K-token context window: GPT-5.6 Luna at $0.20 per 1M input tokens (1M context), GPT-5.4 Nano at $0.20 per 1M input tokens (400K context) and Llama 4 Scout at $0.17 per 1M input tokens (10M context). Rates are the cheapest listed provider for each; whether the quality holds for a given task is a separate question.

Cheaper alternatives to Qwen3.6-35B-A3B