Meta logo

Llama 3 8B

Meta's 8B-parameter open-weight Llama 3 model, released April 2024.

Params
8B
Context
8K tokens
Cutoff
Released
License
Meta Llama 3 Community License

Hosted API price

Cheapest via Replicate · per 1M tokens

USD
Input
$0.05 /1M
Output
$0.25 /1M
Hosted by
Replicate Together AI Meta 3 providers

Hosted API pricing

How we price

Order

Providers are listed cheapest input rate first, then by output rate, a tie going to the creator's own listing. A provider that lists the model without a published rate sits at the end.

Prices

Base-tier, on-demand rates in USD per 1M tokens, converted at ECB reference rates where a provider publishes in another currency. Cached-input and batch rates show only where a provider publishes them. "Cheapest" marks the one provider with the lowest input and output pair; a tie wears no badge.

Cost

Input rate times the input tokens plus output rate times the output tokens, at the monthly volume set above the table. Cached-input, batch, long-context and reasoning-token billing are not modelled.

Transparency and funding

  • Affiliates: Affiliate links are marked. We may earn a commission if you click them, but commissions never affect the order.

Every provider serving Llama 3 8B, cheapest input first. Set your monthly volume to see what each would bill.

Every provider serving Llama 3 8B, with its price per 1M input and output tokens and what the volume above would cost
Provider Input /1M Output /1M Cost at 10M in + 2M out Link
Replicate logo Replicate Cheapest $0.05 $0.25 $1.00 Visit website
Together AI logo Together AI $0.14 $0.14 $1.68 Visit website
Meta logo Meta Creator Open weights, no hosted price Visit website

Heads up: Base-tier, on-demand rates per 1M tokens; cached, batch and long-context tiers excluded. A provider may serve a shorter context or a quantized build than the creator's release. Verify before provisioning. More on how we price.

Estimated cost to self-host

Llama 3 8B needs about 7 GB of GPU memory at 4-bit with its full 8K context. The cheapest rental that fits is 1× RTX 3060 at about $58 a month.

That costs the same as roughly 691M tokens a month on Replicate's API. Self-hosting is more expensive below that volume.

Fits on 1 Single node ≤ 8
What it costs to self-host Llama 3 8B at each precision: the memory it needs, the cheapest rental that holds it, that rental's month and the API volume it pays off at
Precision Memory Cheapest fit (Full 8K context) Cost /mo Break-even vs API
4-bit INT4 / FP4 7 GB Nvidia logo 1× RTX 3060 $58 691M tokens /mo
8-bit FP8 / INT8 11 GB Nvidia logo 1× RTX 4060 Ti $94 1B tokens /mo
16-bit FP16 / BF16 19 GB Nvidia logo 2× RTX 3060 $115 1B tokens /mo

Estimates based on median on-demand rates for Nvidia GPUs. Memory is weights plus KV cache for one request, using FP8 KV cache where supported and FP16/BF16 otherwise. Break-even assumes a 5:1 input-to-output ratio. No guarantee of runtime support, usable performance, or that a matching quantized build exists. Pricing methodology.

Similarly priced models

The models nearest Llama 3 8B by blended rate, each at its own cheapest provider.

Models priced nearest Llama 3 8B per 1M tokens, each at its cheapest provider
Model Blended /1M Input /1M Output /1M Context Cutoff vs Llama 3 8B
Meta Llama 3.2 1B Meta $0.056 $0.027 $0.201 128K Dec 2023 −33%
Cohere Command R7B Cohere $0.0583 $0.04 $0.15 128K −30%
Google Cloud Gemma 3 12B Instruct Google Cloud $0.0583 $0.05 $0.10 128K Aug 2024 −30%
OpenAI GPT-OSS-120B OpenAI $0.0628 $0.039 $0.182 131K Jun 2024 −25%
Nvidia Nemotron 3 Nano 30B A3B Nvidia $0.075 $0.05 $0.20 262K Jun 2025 −10%
Meta Llama 3 8B This model Meta $0.0833 $0.05 $0.25 8K Mar 2023
Alibaba Cloud Qwen3-Coder-30B-A3B Alibaba Cloud $0.0917 $0.06 $0.25 262K +10%
Meta Llama 3.2 3B Meta $0.0983 $0.051 $0.335 128K Dec 2023 +18%
Alibaba Cloud Qwen3-30B-A3B FP8 Alibaba Cloud $0.0983 $0.051 $0.335 33K +18%
Mistral Ministral 3 3B Mistral $0.10 $0.10 $0.10 256K +20%
Z.AI GLM-5.3-Flash Z.AI $0.1042 $0.075 $0.25 1M +25%

USD per 1M tokens at each model's cheapest listed provider. Blended is the cost of 10M input plus 2M output tokens, spread over the 12M.

Capabilities

What Llama 3 8B accepts and can do, as published by Meta.

  • Text Accepts and generates natural-language text.

Common questions

What is Llama 3 8B good for?

Small text-only tasks such as classification, short chat and extraction. Open weights under the Meta Llama 3 Community License, small enough for a single consumer GPU.

When is Llama 3 8B not a good fit?

Image input and tool calling, which it has neither of, and anything needing recent knowledge. Llama 3.1 8B takes far longer inputs and adds function calling under a similar licence.

What is the cheapest way to run Llama 3 8B?

Hosted, unless you push serious volume. Replicate charges $0.05 in / $0.25 out per 1M tokens. The cheapest rental that fits is 1x RTX 3060 at $58 a month, which costs the same as about 691M tokens a month on that API.

How does Llama 3 8B compare with Llama 3.1 8B?

Meta says Llama 3.1 8B upgrades Llama 3 8B with more and better training data and a new multi-round post-training, adding a much longer context, multilingual support, tool use and what Meta calls stronger reasoning. Same text input. Llama 3 8B costs 150% more per 1M input tokens and 400% more per 1M output tokens than Llama 3.1 8B, at each model's cheapest listed provider. It came out 3 months earlier, has an older knowledge cutoff (March 2023 against December 2023) and a smaller context (8K against 131K tokens) and is licensed under Meta Llama 3 Community License against Llama 3.1 Community License.

Can I self-host Llama 3 8B?

Yes. The weights are released under the Meta Llama 3 Community License. At 4-bit it needs about 7 GB of GPU memory, which starts at roughly $58 a month on the cheapest rental that fits.

More from Meta

  • Llama 3.2 3B Open weights 128K context · cutoff Dec 2023
    From $0.051 / $0.335 /1M in / out at the cheapest provider
  • Llama 3.2 1B Open weights 128K context · cutoff Dec 2023
    From $0.027 / $0.201 /1M in / out at the cheapest provider
  • Llama 3.2 11B Vision Open weights 128K context · cutoff Dec 2023
    From $0.049 / $0.676 /1M in / out at the cheapest provider
  • Llama 3.3 70B Open weights 128K context · cutoff Dec 2023
    From $0.13 / $0.40 /1M in / out at the cheapest provider
  • Llama 4 Scout Open weights 10M context · cutoff Aug 2024
    From $0.16 / $0.64 /1M in / out at the cheapest provider
  • Llama 3.1 8B Open weights 131K context · cutoff Dec 2023
    From $0.02 / $0.05 /1M in / out at the cheapest provider
  • Llama 4 Maverick Open weights 1M context · cutoff Aug 2024
    From $0.20 / $0.696 /1M in / out at the cheapest provider
  • Muse Glimmer 30B Open weights 131K context · cutoff Jan 2026
    From $0.20 / $0.80 /1M in / out at the cheapest provider