Meta logo Open weights · Llama 3.2 Community License

Llama 3.2 11B Vision

Meta's smallest multimodal Llama, a 10.6B model that adds image understanding through cross-attention adapters and suits single-GPU deployment.

Cheapest via Cloudflare
Input
$0.05 /1M tokens
Output
$0.68 /1M tokens
Cloudflare Meta 2 providers

Key Specifications

Context window
128K tokens
Knowledge cutoff
Parameters
10.6B
Inputs
Text, image
Capabilities Show details
Function calling Connect to external tools, APIs, and systems.

Hosted API pricing

Provider Input /1M tokens Output /1M tokens Cost at 10M in + 2M out
Cloudflare logo Cloudflare $0.05 $0.68 $1.86 View
Meta logo Meta Creator Open weights, no hosted price View

Heads up: Base-tier, on-demand rates per 1M tokens; cached, batch and long-context tiers excluded. A provider may serve a shorter context or a quantized build than the creator's release. Verify before provisioning. More on how we price.

Estimated cost to self-host

Llama 3.2 11B Vision needs about 10 GB of GPU memory at 4-bit with 32K context. The cheapest rental that fits is 1× RTX 4070 at about $72 a month.

That costs the same as roughly 465M tokens a month on Cloudflare's API. Self-hosting is more expensive below that volume.

Fits on 1 Single node ≤ 8
Precision Memory Cheapest fit (32K context) Cost /mo Break-even vs API
4-bitINT4 / FP4
10 GB
$72
465M tokens /mo
8-bitFP8 / INT8
17 GB
$86
557M tokens /mo
16-bitFP16 / BF16
27 GB
$158
1B tokens /mo

Estimates based on median on-demand rates for Nvidia GPUs. Memory is weights plus KV cache for one request, using FP8 KV cache where supported and FP16/BF16 otherwise. Break-even assumes a 5:1 input-to-output ratio. No guarantee of runtime support, usable performance, or that a matching quantized build exists. How we estimate costs.

Similarly priced models

The models nearest Llama 3.2 11B Vision by blended rate, each at its own cheapest provider.

Model Blended /1M vs Llama 3.2 11B Vision
Alibaba Cloud Qwen3-32B FP8 Alibaba Cloud $0.1333 −14%
Google Cloud Gemini 2.5 Flash-Lite Google Cloud $0.15 −3%
OpenAI GPT-4.1 Nano OpenAI $0.15 −3%
Mistral Ministral 3 8B Mistral $0.15 −3%
Alibaba Cloud Qwen3.5-Flash Alibaba Cloud $0.15 −3%
Meta Llama 3.2 11B Vision This model Meta $0.155
Nvidia Nemotron 3 Super 120B A12B Nvidia $0.1583 +2%
Google Cloud Gemma 4 31B Google Cloud $0.17 +10%
Alibaba Cloud Qwen3-235B-A22B Instruct (2507) Alibaba Cloud $0.1717 +11%
Meta Llama 3.3 70B Meta $0.175 +13%
Alibaba Cloud Qwen3.5-27B Alibaba Cloud $0.19 +23%

Prices are USD per 1M tokens at each model's cheapest listed provider. Blended is the cost of 10M input plus 2M output tokens, spread over the 12M.

Frequently Asked Questions

What is Llama 3.2 11B Vision good for?

Image understanding with text: captions, document and chart questions. Open weights under the Llama 3.2 Community License, small enough for a single consumer GPU.

When is Llama 3.2 11B Vision not a good fit?

Tool calling on image inputs: it works only with text-only prompts. No structured output, and its knowledge is older than the current open models'.

What is the cheapest way to run Llama 3.2 11B Vision?

Hosted, unless you push serious volume. Cloudflare charges $0.05 in / $0.68 out per 1M tokens. The cheapest rental that fits is 1x RTX 4070 at $72 a month, which costs the same as about 465M tokens a month on that API.

Can I self-host Llama 3.2 11B Vision?

Yes. The weights are released under the Llama 3.2 Community License. At 4-bit it needs about 10 GB of GPU memory, which starts at roughly $72 a month on the cheapest rental that fits.

More from Meta

Model Context Input /1M Output /1M
Meta Llama 3.3 70B Open weights 128K $0.13 $0.40
Meta Llama 4 Scout Open weights 10M $0.16 $0.64
Meta Llama 3.2 3B Open weights 128K $0.05 $0.34
Meta Llama 4 Maverick Open weights 1M $0.20 $0.70
Meta Llama 3 8B Open weights 8K $0.05 $0.25
Meta Muse Glimmer 30B Open weights 131K $0.20 $0.80
Meta Llama 3.2 1B Open weights 128K $0.03 $0.20
Meta Llama 3.1 8B Open weights 131K $0.02 $0.05