Google Cloud logo Open weights · Apache 2.0

Gemma 4 E2B

The smallest Gemma 4 model at 2.3B effective parameters, built for mobile and IoT edge devices. Apache 2.0 license.

Key Specifications

Context window
128K tokens
Knowledge cutoff
Released
Parameters
5.1B
Inputs
Text, image, audio
Capabilities Show details
Reasoning Thinks before it answers, always on or as a switchable mode. Function calling Connect to external tools, APIs, and systems. Structured output Return responses in structured formats like JSON.

Hosted API pricing

Provider Input / 1M tokens Output / 1M tokens Cost at 10M in + 2M out
Google Cloud logo Google Cloud Creator Open weights, no hosted price View

Heads up: Base-tier, on-demand rates per 1M tokens; cached, batch and long-context tiers excluded. A provider may serve a shorter context or a quantized build than the creator's release. Verify before provisioning. More on how we price.

Estimated cost to self-host

Gemma 4 E2B needs about 5 GB of GPU memory at 4-bit with 32K context. The cheapest rental that fits is 1× GTX 1660 Super at about $43 a month.

Fits on 1
Precision Memory Cheapest fit (32K context) Cost / mo
4-bitINT4 / FP4
5 GB
$43
8-bitFP8 / INT8
7 GB
$43
16-bitFP16 / BF16
12 GB
$79

Estimates based on median on-demand rates for Nvidia GPUs. Memory is weights plus KV cache for one request, using FP8 KV cache where supported and FP16/BF16 otherwise. No guarantee of runtime support, usable performance, or that a matching quantized build exists. How we estimate costs.

Frequently Asked Questions

What is Gemma 4 E2B good for?

Mobile and IoT edge deployments taking text, images and audio. Apache 2.0 weights, at the light end of the Gemma 4 family.

When is Gemma 4 E2B not a good fit?

Video input, and hosted tools like web search or code execution. No maximum output is published.

What is the cheapest way to run Gemma 4 E2B?

No provider we track hosts it; self-hosting on the cheapest rental that fits (1x GTX 1660 Super, $0.06 per hour) costs about $43 a month.

Can I self-host Gemma 4 E2B?

Yes. The weights are Apache 2.0 licensed. At 4-bit it needs about 5 GB of GPU memory, which starts at roughly $43 a month on the cheapest rental that fits.

More from Google Cloud

Model Context Input / 1M Output / 1M
Google Cloud Gemini 3.5 Flash-Lite 1M $0.30 $2.50
Google Cloud Gemini 3.6 Flash 1M $0.75 $3.75
Google Cloud Gemini 3.7 Flash 1M $0.75 $3.75
Google Cloud Gemini 3.8 Flash 1M $0.75 $3.75
Google Cloud Gemini 2.5 Flash 1M $0.30 $2.50
Google Cloud Gemini 2.5 Flash-Lite 1M $0.10 $0.40
Google Cloud Gemini 2.5 Pro 1M $1.25 $10.00
Google Cloud Gemini 3 Flash Preview 1M $0.50 $3.00