LLM API pricing

Compare 185 large language models across 26 providers.

Cheapest input
$0.02 /1M
Llama 3.1 8B via Novita
Median input
$0.42 /1M
across 168 priced models
Models
185
102 with open weights
Providers
26
510 listings

LLM models

Cheapest listed price per 1M tokens, input and output from the same provider. Newest cutoff first.

How this list works

Order

Models are sorted by three factors, in order of priority:

  1. Knowledge cutoff: Models with the most recent training data appear first. Models without a known cutoff date appear last.
  2. Context window: Among models with the same knowledge cutoff, those with larger context windows rank higher.
  3. Name: Models that tie on both factors are sorted alphabetically.

Prices

Each row shows the cheapest provider for the model, with input and output rates from that same provider. Rates are base-tier, on-demand prices in USD per 1M tokens. "Open weights" marks a model whose creator publishes the weights and lists no hosted price.

Sorting and filtering

The list opens on the newest 20 models; "Show all" reveals the rest, and so does any search, filter or sort. Click a column header to sort by it, again to reverse it, and a third time to restore the default order. The search, the creator chips and "Open weights" (models whose creator publishes the weights) combine: a model has to match all three. A retired model no provider lists any more is not in the list; its own page stays up.

Search

We match the model name, its creator and the providers that serve it. Matching is partial and case-insensitive, and every word has to match, so "llama novita" shows only Novita's servings of Llama.

Transparency and funding

  • Affiliates: Affiliate links are marked. We may earn a commission if you click them, but commissions never affect the order.
Creator
Large language models with context window, knowledge cutoff and the cheapest listed price per 1M input and output tokens
Model Inputs Context Cutoff Input /1M Output /1M
Anthropic Claude Fable 5.1 Anthropic · 2 providers 1M Jun 2026 $10.00 /1M $50.00 /1M
Anthropic Claude Opus 5 Anthropic · 2 providers 1M May 2026 $5.00 /1M $25.00 /1M
OpenAI GPT-6 Astra OpenAI 1M Apr 2026 $10.00 /1M $50.00 /1M
Google Cloud Gemini 3.5 Flash-Lite Google Cloud 1M Mar 2026 $0.30 /1M $2.50 /1M
Google Cloud Gemini 3.6 Flash Google Cloud 1M Mar 2026 $0.75 /1M $3.75 /1M
Google Cloud Gemini 3.7 Flash Google Cloud 1M Mar 2026 $0.75 /1M $3.75 /1M
Google Cloud Gemini 3.8 Flash Google Cloud 1M Mar 2026 $0.75 /1M $3.75 /1M
OpenAI GPT-5.6 Luna OpenAI · 4 providers 1M Feb 2026 $0.20 /1M $1.20 /1M
OpenAI GPT-5.6 Sol OpenAI · 4 providers 1M Feb 2026 $2.00 /1M $10.00 /1M
OpenAI GPT-5.6 Terra OpenAI · 4 providers 1M Feb 2026 $2.00 /1M $12.00 /1M
xAI Grok 4.5 xAI 500K Feb 2026 $2.00 /1M $6.00 /1M
Anthropic Claude Fable 5 Anthropic · 3 providers 1M Jan 2026 $10.00 /1M $50.00 /1M
Anthropic Claude Opus 4.7 Anthropic · 3 providers 1M Jan 2026 $5.00 /1M $25.00 /1M
Anthropic Claude Opus 4.8 Anthropic · 2 providers 1M Jan 2026 $5.00 /1M $25.00 /1M
Anthropic Claude Sonnet 5 Anthropic · 3 providers 1M Jan 2026 $2.00 /1M $10.00 /1M
xAI Grok 4.6 xAI 500K Jan 2026 $2.00 /1M $6.00 /1M
Meta Muse Glimmer 30B Meta · Open weights · 3 providers 131K Jan 2026 $0.20 /1M $0.80 /1M
OpenAI GPT-5.5 OpenAI · 3 providers 1M Dec 2025 $5.00 /1M $30.00 /1M
OpenAI GPT-5.5 Pro OpenAI 1M Dec 2025 $30.00 /1M $180.00 /1M
Nvidia Nemotron 3 Ultra 550B A55B Nvidia · Open weights · 3 providers 1M Sep 2025 $0.66 /1M $2.64 /1M

No models matching your filters.

Showing 20 of 185 models

Heads up: Base-tier, on-demand rates per 1M tokens; cached, batch and long-context tiers excluded. A provider may serve a shorter context or a quantized build than the creator's release. Verify before provisioning. More on how we price.

No models matching your filters.

Cheapest listed provider per model. Models without a hosted price are not on the chart.

Common questions

What is the cheapest LLM API?

The cheapest priced model we track is Llama 3.1 8B at $0.02 per 1M input tokens and $0.05 per 1M output tokens via Novita. The median across all 168 priced models is $0.42 per 1M input tokens.

What does a price per 1M tokens mean?

US dollars per one million tokens (about 750,000 English words), with input and output billed at separate rates. A 2,000-token prompt with a 500-token reply costs 2,000 x the input rate plus 500 x the output rate, divided by one million; the LLM cost calculator does this for any workload.

Why do output tokens cost more than input tokens?

Output is generated one token at a time, re-reading everything before it, so it takes several times the compute of input and is priced 3x to 8x higher. Read-heavy work (classification, extraction) follows the input rate; drafting and code generation follow the output rate.

Why does the same model cost different amounts at different providers?

Anyone can serve an open-weight model, so hosts compete on hardware, batching, quantisation and margin, and the same weights can differ severalfold in price. Closed models are resold at rates close to the creator's own.

What is an open-weight model, and can I run one myself?

One whose creator publishes the trained weights, so it can run on your own hardware or a rented GPU; 102 models here are listed that way. Each open model's page estimates the GPU cost; see cloud GPU pricing and the LM Studio guide for running one locally.

Which prices are listed here?

Base-tier, on-demand rates in USD per 1M tokens, converted at ECB reference rates where a provider publishes in another currency. Cached-input, batch, long-context and prepaid tiers are not included.

How often are LLM prices updated?

By hand, against each provider's published pricing page; last updated September 11, 2026. Every row links to the page its figure came from.

Last updated .

Back to top ↑