Tencent logo Open weights · Apache 2.0

Hy3

Tencent's flagship open-weight Hunyuan model, released July 2026: a 295B-parameter MoE (21B active) under Apache 2.0 with three reasoning modes and a design emphasis on low hallucination rates.

Cheapest via GPUhub
Input
$0.14 / 1M tokens
Output
$0.56 / 1M tokens
GPUhub Novita Tencent 3 providers

Key Specifications

Context window
262K tokens
Released
Parameters
295B, 21B active
Inputs
Text
Capabilities Show details
Reasoning Thinks before it answers, always on or as a switchable mode. Function calling Connect to external tools, APIs, and systems. Structured output Return responses in structured formats like JSON.

Hosted API pricing

Provider Input / 1M tokens Output / 1M tokens Cost at 10M in + 2M out
GPUhub logo GPUhub Cheapest $0.14 $0.56 $2.52 View
Novita logo Novita $0.14 $0.58 $2.56 View
Tencent logo Tencent Creator Open weights, no hosted price View

Heads up: Prices are estimates using base-tier, on-demand rates per 1M tokens from published pages; cached, batch and long-context tiers are shown only where listed. Providers may serve a shorter context or a quantized build than the creator's release. Verify with the provider before provisioning. How we estimate costs.

Estimated cost to self-host

Hy3 needs about 172 GB of GPU memory at 4-bit. The cheapest rental that fits is 8× RTX 3090 at about $749 a month.

That costs the same as roughly 4B tokens a month on GPUhub's API. Self-hosting is more expensive below that volume.

Precision Cheapest, 1x concurrency Cheapest, 8x concurrency
4-bit
Nvidia logo 8× RTX 3090 172 GB · $749/mo
Nvidia logo 8× RTX 5090 228 GB · $2,650/mo
8-bit
Nvidia logo 8× RTX A6000 305 GB · $3,226/mo
Nvidia logo 8× A100 361 GB · $10,138/mo
16-bit
Nvidia logo 8× RTX PRO 6000 600 GB · $12,614/mo
Nvidia logo 8× RTX PRO 6000 656 GB · $12,614/mo

Estimates, not quotes. Nvidia cards at median on-demand rates, 32K context per request. Break-even assumes 5:1 input to output. How we estimate costs.

Similarly priced models

The models nearest Hy3 by blended rate, five each way, each at its own cheapest provider.

Model Blended / 1M vs Hy3
Google Cloud Gemma 4 31B Google Cloud $0.17 −19%
Alibaba Cloud Qwen3-235B-A22B Instruct (2507) Alibaba Cloud $0.17 −18%
Google Cloud Gemma 4 26B A4B Google Cloud $0.18 −17%
Meta Llama 3.3 70B Meta $0.18 −13%
Mistral Ministral 3 14B Mistral $0.20 −5%
Tencent Hy3 This model Tencent $0.21
Alibaba Cloud Qwen-MT-Turbo Alibaba Cloud $0.22 +2%
Cohere Command R 08-2024 Cohere $0.23 +7%
OpenAI GPT-4o Mini OpenAI $0.23 +7%
Mistral Mistral Small 4 Mistral $0.23 +7%
Alibaba Cloud Qwen3-Coder-Next Alibaba Cloud $0.23 +10%

Prices are USD per 1M tokens at each model's cheapest listed provider. Blended is the cost of 10M input plus 2M output tokens, spread over the 12M.

Frequently Asked Questions

What is Hy3 good for?

General text tasks where you can set reasoning effort from no-think up to high. It supports function calling and structured output, and Tencent says it has low hallucination rates. Apache 2.0 weights allow commercial self-hosting at about 177 GB at 4-bit.

When is Hy3 not a good fit?

No image or video input, and no published max output for long single responses. Tencent's own API has no published input and output price split, so hosted pricing comes from third parties.

What is the cheapest way to run Hy3?

Hosted, unless you push serious volume. GPUhub charges $0.14 in / $0.56 out per 1M tokens. The cheapest rental that fits is 8x RTX 3090 at $749 a month, which costs the same as about 4B tokens a month on that API.

Can I self-host Hy3?

Yes. The weights are Apache 2.0 licensed. At 4-bit it needs about 172 GB of GPU memory, which starts at roughly $749 a month on the cheapest rental that fits. See the table above for 8-bit and 16-bit.