Hy3
Tencent's flagship open-weight Hunyuan model, released July 2026: a 295B-parameter MoE (21B active) under Apache 2.0 with three reasoning modes and a design emphasis on low hallucination rates.
Key Specifications
- Context window
- 262K tokens
- Released
- Parameters
- 295B, 21B active
- Inputs
- Text
- Capabilities Show details
- Reasoning Thinks before it answers, always on or as a switchable mode. Function calling Connect to external tools, APIs, and systems. Structured output Return responses in structured formats like JSON.
Hosted API pricing
| Provider | Input / 1M tokens | Output / 1M tokens | Cost at 10M in + 2M out | |
|---|---|---|---|---|
|
|
$0.14 | $0.56 | $2.52 | View |
|
|
$0.14 | $0.58 | $2.56 | View |
|
|
Open weights, no hosted price | View | ||
Heads up: Prices are estimates using base-tier, on-demand rates per 1M tokens from published pages; cached, batch and long-context tiers are shown only where listed. Providers may serve a shorter context or a quantized build than the creator's release. Verify with the provider before provisioning. How we estimate costs.
Estimated cost to self-host
Hy3 needs about 172 GB of GPU memory at 4-bit. The cheapest rental that fits is 8× RTX 3090 at about $749 a month.
That costs the same as roughly 4B tokens a month on GPUhub's API. Self-hosting is more expensive below that volume.
| Precision | Cheapest, 1x concurrency | Cheapest, 8x concurrency |
|---|---|---|
| 4-bit |
|
|
| 8-bit |
|
|
| 16-bit |
|
|
Estimates, not quotes. Nvidia cards at median on-demand rates, 32K context per request. Break-even assumes 5:1 input to output. How we estimate costs.
Similarly priced models
The models nearest Hy3 by blended rate, five each way, each at its own cheapest provider.
| Model | Blended / 1M | Input / 1M | Output / 1M | Context | Cutoff | vs Hy3 |
|---|---|---|---|---|---|---|
|
|
$0.17 | $0.13 | $0.37 | 256K | Jan 2025 | −19% |
|
|
$0.17 | $0.09 | $0.58 | 262K | −18% | |
|
|
$0.18 | $0.13 | $0.40 | 256K | Jan 2025 | −17% |
|
|
$0.18 | $0.14 | $0.40 | 128K | Dec 2023 | −13% |
|
|
$0.20 | $0.20 | $0.20 | 256K | −5% | |
|
|
$0.21 | $0.14 | $0.56 | 262K | ||
|
|
$0.22 | $0.16 | $0.49 | 16K | +2% | |
|
|
$0.23 | $0.15 | $0.60 | 128K | Jun 2024 | +7% |
|
|
$0.23 | $0.15 | $0.60 | 128K | Oct 2023 | +7% |
|
|
$0.23 | $0.15 | $0.60 | 256K | +7% | |
|
|
$0.23 | $0.08 | $0.98 | 262K | +10% |
Prices are USD per 1M tokens at each model's cheapest listed provider. Blended is the cost of 10M input plus 2M output tokens, spread over the 12M.
Frequently Asked Questions
What is Hy3 good for?
General text tasks where you can set reasoning effort from no-think up to high. It supports function calling and structured output, and Tencent says it has low hallucination rates. Apache 2.0 weights allow commercial self-hosting at about 177 GB at 4-bit.
When is Hy3 not a good fit?
No image or video input, and no published max output for long single responses. Tencent's own API has no published input and output price split, so hosted pricing comes from third parties.
What is the cheapest way to run Hy3?
Hosted, unless you push serious volume. GPUhub charges $0.14 in / $0.56 out per 1M tokens. The cheapest rental that fits is 8x RTX 3090 at $749 a month, which costs the same as about 4B tokens a month on that API.
Can I self-host Hy3?
Yes. The weights are Apache 2.0 licensed. At 4-bit it needs about 172 GB of GPU memory, which starts at roughly $749 a month on the cheapest rental that fits. See the table above for 8-bit and 16-bit.