GPT-3.5 Turbo
OpenAI's lightweight model of the GPT-3.5 generation.
Key Specifications
- Context window
- 16K tokens
- Max output
- 4K tokens
- Knowledge cutoff
- Released
- Inputs
- Text
Hosted API pricing
| Provider | Input / 1M tokens | Output / 1M tokens | Cost at 10M in + 2M out | |
|---|---|---|---|---|
|
|
$0.50 | $1.50 | $8.00 | View |
|
|
$0.50 | $1.50 | $8.00 | View |
Heads up: Base-tier, on-demand rates per 1M tokens; cached, batch and long-context tiers excluded. A provider may serve a shorter context or a quantized build than the creator's release. Verify before provisioning. More on how we price.
Similarly priced models
The models nearest GPT-3.5 Turbo by blended rate, each at its own cheapest provider.
| Model | Blended / 1M | Input / 1M | Output / 1M | Context | Cutoff | vs GPT-3.5 Turbo |
|---|---|---|---|---|---|---|
|
|
$0.575 | $0.23 | $2.30 | 262K | −14% | |
|
|
$0.575 | $0.38 | $1.55 | 262K | −14% | |
|
|
$0.60 | $0.40 | $1.60 | 1M | Jun 2024 | −10% |
|
|
$0.6667 | $0.30 | $2.50 | 1M | Jan 2025 | 0% |
|
|
$0.6667 | $0.30 | $2.50 | 1M | Mar 2026 | 0% |
|
|
$0.6667 | $0.50 | $1.50 | 16K | Sep 2021 | |
|
|
$0.6667 | $0.50 | $1.50 | 256K | 0% | |
|
|
$0.7333 | $0.40 | $2.40 | 1M | +10% | |
|
|
$0.7667 | $0.60 | $1.60 | 200K | +15% | |
|
|
$0.80 | $0.80 | $0.80 | 131K | +20% | |
|
|
$0.8667 | $0.60 | $2.20 | 205K | +30% |
Prices are USD per 1M tokens at each model's cheapest listed provider. Blended is the cost of 10M input plus 2M output tokens, spread over the 12M.
Frequently Asked Questions
What is GPT-3.5 Turbo good for?
Cheap text-only chat and classification. OpenAI calls it a legacy model and has recommended GPT-4o Mini instead since July 2024.
When is GPT-3.5 Turbo not a good fit?
Images, tool calling and structured output, since it has none of them. Its knowledge is years out of date, and OpenAI plans to shut down the API model in October 2026.
What is the cheapest way to run GPT-3.5 Turbo?
The cheapest hosted rate is $0.50 / $1.50 per 1M tokens via OpenAI.
How does GPT-3.5 Turbo compare with GPT-4o Mini?
OpenAI recommends GPT-4o Mini in place of GPT-3.5 Turbo and says it is smarter, cheaper and just as fast, with better long-context performance, and can be fine-tuned or distilled from GPT-4o. GPT-3.5 Turbo lists at 233% more per 1M input tokens and 150% more per 1M output tokens than GPT-4o Mini. It came out 16 months earlier, has an older knowledge cutoff (September 2021 against October 2023), a smaller context (16K against 128K tokens) and a smaller max output (4K against 16K tokens) and takes no image input, which GPT-4o Mini does.
Can I self-host GPT-3.5 Turbo?
No. OpenAI does not publish the weights; GPT-3.5 Turbo is available only through hosted APIs.
More from OpenAI
| Model | Context | Input / 1M | Output / 1M |
|---|---|---|---|
|
|
1M | $0.40 | $1.60 |
|
|
400K | $0.25 | $2.00 |
|
|
400K | $0.20 | $1.25 |
|
|
1M | $0.20 | $1.20 |
|
|
400K | $0.75 | $4.50 |
|
|
200K | $1.00 | $4.00 |
|
|
200K | $1.10 | $4.40 |
|
|
128K | $0.15 | $0.60 |