Google Cloud logo Open weights · Gemma Terms of Use

Gemma 3 4B

A compact 4B model in Google's open-weight Gemma 3 family, built for local and edge deployments.

Key Specifications

Context window
128K tokens
Max output
128K tokens
Knowledge cutoff
Released
Parameters
4.3B
Inputs
Text, image
Capabilities Show details
Function calling Connect to external tools, APIs, and systems.

Hosted API pricing

Provider Input /1M tokens Output /1M tokens Cost at 10M in + 2M out
Google Cloud logo Google Cloud Creator Open weights, no hosted price View

Heads up: Base-tier, on-demand rates per 1M tokens; cached, batch and long-context tiers excluded. A provider may serve a shorter context or a quantized build than the creator's release. Verify before provisioning. More on how we price.

Estimated cost to self-host

Gemma 3 4B needs about 5 GB of GPU memory at 4-bit with 32K context. The cheapest rental that fits is 1× GTX 1660 Super at about $43 a month.

Fits on 1
Precision Memory Cheapest fit (32K context) Cost /mo
4-bitINT4 / FP4
5 GB
$43
8-bitFP8 / INT8
7 GB
$43
16-bitFP16 / BF16
11 GB
$79

Estimates based on median on-demand rates for Nvidia GPUs. Memory is weights plus KV cache for one request, using FP8 KV cache where supported and FP16/BF16 otherwise. No guarantee of runtime support, usable performance, or that a matching quantized build exists. How we estimate costs.

Frequently Asked Questions

What is Gemma 3 4B good for?

Local and edge deployments taking text and images. Open weights under the Gemma Terms of Use, small enough for a single consumer GPU.

When is Gemma 3 4B not a good fit?

Video or audio input, and structured output. The Gemma Terms of Use are a custom license rather than Apache 2.0, and the Gemma 4 family knows more recent events.

What is the cheapest way to run Gemma 3 4B?

No provider we track hosts it; self-hosting on the cheapest rental that fits (1x GTX 1660 Super, $0.06 per hour) costs about $43 a month.

Can I self-host Gemma 3 4B?

Yes. The weights are released under the Gemma Terms of Use. At 4-bit it needs about 5 GB of GPU memory, which starts at roughly $43 a month on the cheapest rental that fits.

More from Google Cloud

Model Context Input /1M Output /1M
Google Cloud Gemini 3.5 Flash-Lite 1M $0.30 $2.50
Google Cloud Gemini 3.6 Flash 1M $0.75 $3.75
Google Cloud Gemini 3.7 Flash 1M $0.75 $3.75
Google Cloud Gemini 3.8 Flash 1M $0.75 $3.75
Google Cloud Gemini 2.5 Flash 1M $0.30 $2.50
Google Cloud Gemini 2.5 Flash-Lite 1M $0.10 $0.40
Google Cloud Gemini 2.5 Pro 1M $1.25 $10.00
Google Cloud Gemini 3 Flash Preview 1M $0.50 $3.00