LLM API pricing
Compare 185 large language models across 26 providers.
- Median input
- $0.42 /1M
- across 168 priced models
- Models
- 185
- 102 with open weights
- Providers
- 26
- 510 listings
LLM models
Cheapest listed price per 1M tokens, input and output from the same provider. Newest cutoff first.
How this list works
Order
Models are sorted by three factors, in order of priority:
- Knowledge cutoff: Models with the most recent training data appear first. Models without a known cutoff date appear last.
- Context window: Among models with the same knowledge cutoff, those with larger context windows rank higher.
- Name: Models that tie on both factors are sorted alphabetically.
Prices
Each row shows the cheapest provider for the model, with input and output rates from that same provider. Rates are base-tier, on-demand prices in USD per 1M tokens. "Open weights" marks a model whose creator publishes the weights and lists no hosted price.
Sorting and filtering
The list opens on the newest 20 models; "Show all" reveals the rest, and so does any search, filter or sort. Click a column header to sort by it, again to reverse it, and a third time to restore the default order. The search, the creator chips and "Open weights" (models whose creator publishes the weights) combine: a model has to match all three. A retired model no provider lists any more is not in the list; its own page stays up.
Search
We match the model name, its creator and the providers that serve it. Matching is partial and case-insensitive, and every word has to match, so "llama novita" shows only Novita's servings of Llama.
Transparency and funding
- Affiliates: Affiliate links are marked. We may earn a commission if you click them, but commissions never affect the order.
| Model | Inputs | Context | Cutoff | Input /1M | Output /1M |
|---|---|---|---|---|---|
|
|
1M | Jun 2026 | $10.00 /1M | $50.00 /1M | |
|
|
1M | May 2026 | $5.00 /1M | $25.00 /1M | |
|
|
1M | Apr 2026 | $10.00 /1M | $50.00 /1M | |
|
|
1M | Mar 2026 | $0.30 /1M | $2.50 /1M | |
|
|
1M | Mar 2026 | $0.75 /1M | $3.75 /1M | |
|
|
1M | Mar 2026 | $0.75 /1M | $3.75 /1M | |
|
|
1M | Mar 2026 | $0.75 /1M | $3.75 /1M | |
|
|
1M | Feb 2026 | $0.20 /1M | $1.20 /1M | |
|
|
1M | Feb 2026 | $2.00 /1M | $10.00 /1M | |
|
|
1M | Feb 2026 | $2.00 /1M | $12.00 /1M | |
|
|
500K | Feb 2026 | $2.00 /1M | $6.00 /1M | |
|
|
1M | Jan 2026 | $10.00 /1M | $50.00 /1M | |
|
|
1M | Jan 2026 | $5.00 /1M | $25.00 /1M | |
|
|
1M | Jan 2026 | $5.00 /1M | $25.00 /1M | |
|
|
1M | Jan 2026 | $2.00 /1M | $10.00 /1M | |
|
|
500K | Jan 2026 | $2.00 /1M | $6.00 /1M | |
|
|
131K | Jan 2026 | $0.20 /1M | $0.80 /1M | |
|
|
1M | Dec 2025 | $5.00 /1M | $30.00 /1M | |
|
|
1M | Dec 2025 | $30.00 /1M | $180.00 /1M | |
|
|
1M | Sep 2025 | $0.66 /1M | $2.64 /1M |
No models matching your filters.
Showing 20 of 185 models
Heads up: Base-tier, on-demand rates per 1M tokens; cached, batch and long-context tiers excluded. A provider may serve a shorter context or a quantized build than the creator's release. Verify before provisioning. More on how we price.
Common questions
What is the cheapest LLM API?
The cheapest priced model we track is Llama 3.1 8B at $0.02 per 1M input tokens and $0.05 per 1M output tokens via Novita. The median across all 168 priced models is $0.42 per 1M input tokens.
What does a price per 1M tokens mean?
US dollars per one million tokens (about 750,000 English words), with input and output billed at separate rates. A 2,000-token prompt with a 500-token reply costs 2,000 x the input rate plus 500 x the output rate, divided by one million; the LLM cost calculator does this for any workload.
Why do output tokens cost more than input tokens?
Output is generated one token at a time, re-reading everything before it, so it takes several times the compute of input and is priced 3x to 8x higher. Read-heavy work (classification, extraction) follows the input rate; drafting and code generation follow the output rate.
Why does the same model cost different amounts at different providers?
Anyone can serve an open-weight model, so hosts compete on hardware, batching, quantisation and margin, and the same weights can differ severalfold in price. Closed models are resold at rates close to the creator's own.
What is an open-weight model, and can I run one myself?
One whose creator publishes the trained weights, so it can run on your own hardware or a rented GPU; 102 models here are listed that way. Each open model's page estimates the GPU cost; see cloud GPU pricing and the LM Studio guide for running one locally.
Which prices are listed here?
Base-tier, on-demand rates in USD per 1M tokens, converted at ECB reference rates where a provider publishes in another currency. Cached-input, batch, long-context and prepaid tiers are not included.
How often are LLM prices updated?
By hand, against each provider's published pricing page; last updated September 11, 2026. Every row links to the page its figure came from.
Last updated .
Back to top ↑