Alibaba Cloud logo

Qwen3-VL-235B-A22B Thinking

Reasoning variant of Qwen3-VL-235B-A22B with extended chain-of-thought for complex visual and multimodal reasoning.

Params
235B, 22B active
Context
262K tokens
Max output
66K
License
Apache 2.0

Hosted API price

List price via Alibaba Cloud · per 1M tokens

USD
Input
$0.40 /1M
Output
$4.00 /1M
Hosted by
Alibaba Cloud Novita 2 providers

Hosted API pricing

How we price

Order

Providers are listed cheapest input rate first, then by output rate, a tie going to the creator's own listing. A provider that lists the model without a published rate sits at the end.

Prices

Base-tier, on-demand rates in USD per 1M tokens, converted at ECB reference rates where a provider publishes in another currency. Cached-input and batch rates show only where a provider publishes them. "Cheapest" marks the one provider with the lowest input and output pair; a tie wears no badge.

Cost

Input rate times the input tokens plus output rate times the output tokens, at the monthly volume set above the table. Cached-input, batch, long-context and reasoning-token billing are not modelled.

Transparency and funding

  • Affiliates: Affiliate links are marked. We may earn a commission if you click them, but commissions never affect the order.

Every provider serving Qwen3-VL-235B-A22B Thinking, cheapest input first. Set your monthly volume to see what each would bill.

Every provider serving Qwen3-VL-235B-A22B Thinking, with its price per 1M input and output tokens and what the volume above would cost
Provider Input /1M Output /1M Cost at 10M in + 2M out Link
Alibaba Cloud logo Alibaba Cloud Creator Cheapest $0.40 $4.00 $12.00 Visit website
Novita logo Novita $0.98 $3.95 $17.70 Visit website

Heads up: Base-tier, on-demand rates per 1M tokens; cached, batch and long-context tiers excluded. A provider may serve a shorter context or a quantized build than the creator's release. Verify before provisioning. More on how we price.

Estimated cost to self-host

Qwen3-VL-235B-A22B Thinking needs about 137 GB of GPU memory at 4-bit with 32K context. The cheapest rental that fits is 8× RTX 3090 at about $1,440 a month.

That costs the same as roughly 1B tokens a month on Alibaba Cloud's API. Self-hosting is more expensive below that volume.

Single node ≤ 8
What it costs to self-host Qwen3-VL-235B-A22B Thinking at each precision: the memory it needs, the cheapest rental that holds it, that rental's month and the API volume it pays off at
Precision Memory Cheapest fit (32K context) Cost /mo Break-even vs API
4-bit INT4 / FP4 137 GB Nvidia logo 8× RTX 3090 $1,440 1B tokens /mo
8-bit FP8 / INT8 243 GB Nvidia logo 8× RTX A6000 $3,283 3B tokens /mo
16-bit FP16 / BF16 478 GB Nvidia logo 8× A100 $9,850 10B tokens /mo

Estimates based on median on-demand rates for Nvidia GPUs. Memory is weights plus KV cache for one request, using FP8 KV cache where supported and FP16/BF16 otherwise. Break-even assumes a 5:1 input-to-output ratio. No guarantee of runtime support, usable performance, or that a matching quantized build exists. Pricing methodology.

Similarly priced models

The models nearest Qwen3-VL-235B-A22B Thinking by blended rate, each at its own cheapest provider.

Models priced nearest Qwen3-VL-235B-A22B Thinking per 1M tokens, each at its cheapest provider
Model Blended /1M Input /1M Output /1M Context Cutoff vs Qwen3-VL-235B-A22B Thinking
Z.AI GLM-5.2 Z.AI $0.95 $0.70 $2.20 1M −5%
Nvidia Nemotron 3 Ultra 550B A55B Nvidia $0.99 $0.66 $2.64 1M Sep 2025 −1%
DeepSeek DeepSeek R1 DeepSeek $1.00 $0.70 $2.50 128K 0%
DeepSeek DeepSeek R1 Turbo DeepSeek $1.00 $0.70 $2.50 64K 0%
Meta Llama 3 70B Meta $1.00 $0.65 $2.75 8K Dec 2023 0%
Alibaba Cloud Qwen3-VL-235B-A22B Thinking This model Alibaba Cloud $1.00 $0.40 $4.00 262K
Moonshot AI Kimi K2.6 Moonshot AI $1.0333 $0.60 $3.20 262K +3%
xAI Grok Build 0.1 xAI $1.1667 $1.00 $2.00 256K +17%
Cohere Command A+ Cohere $1.20 $0.80 $3.20 128K +20%
Z.AI GLM-5.1 Z.AI $1.2377 $0.825 $3.301 205K +24%
Google Cloud Gemini 3.6 Flash Google Cloud $1.25 $0.75 $3.75 1M Mar 2026 +25%

USD per 1M tokens at each model's cheapest listed provider. Blended is the cost of 10M input plus 2M output tokens, spread over the 12M.

Capabilities

What Qwen3-VL-235B-A22B Thinking accepts and can do, as published by Alibaba Cloud.

  • Text Accepts and generates natural-language text.
  • Image Accepts images as input alongside text.
  • Video Accepts video as input.
  • Reasoning Thinks before it answers, always on or as a switchable mode.
  • Function calling Connect to external tools, APIs, and systems.
  • Structured output Return responses in structured formats like JSON.

Common questions

What is Qwen3-VL-235B-A22B Thinking good for?

Visual reasoning with a long chain of thought over images and video: operating GUIs, spatial reasoning and STEM problems. Apache 2.0 weights, on a multi-GPU node.

When is Qwen3-VL-235B-A22B Thinking not a good fit?

Cost-sensitive volume work: it sits near the top of Alibaba's price range. No audio input, no built-in web search, though a host or your own tool loop can add one, and self-hosting takes a multi-GPU node.

What is the cheapest way to run Qwen3-VL-235B-A22B Thinking?

Hosted, unless you push serious volume. Alibaba Cloud charges $0.40 in / $4.00 out per 1M tokens. The cheapest rental that fits is 8x RTX 3090 at $1,440 a month, which costs the same as about 1B tokens a month on that API.

Can I self-host Qwen3-VL-235B-A22B Thinking?

Yes. The weights are Apache 2.0 licensed. At 4-bit it needs about 137 GB of GPU memory, which starts at roughly $1,440 a month on the cheapest rental that fits.

More from Alibaba Cloud

  • Qwen3.6-Plus 1M context
    From $0.50 / $3.00 /1M in / out at the cheapest provider
  • Qwen3-Max 262K context
    From $0.85 / $3.38 /1M in / out at the cheapest provider
  • Qwen3.5-Plus 1M context
    From $0.40 / $2.40 /1M in / out at the cheapest provider
  • Qwen3-Coder-Plus 1M context
    From $1.00 / $5.00 /1M in / out at the cheapest provider
  • Qwen3.7-Max 1M context
    From $1.25 / $3.75 /1M in / out at the cheapest provider
  • Qwen3-235B-A22B Thinking (2507) Open weights 262K context
    From $0.23 / $2.30 /1M in / out at the cheapest provider
  • Qwen3-Coder-480B-A35B Open weights 262K context
    From $0.38 / $1.55 /1M in / out at the cheapest provider
  • Qwen3-VL-30B-A3B Thinking Open weights 262K context
    From $0.20 / $2.40 /1M in / out at the cheapest provider