Google Cloud logo

Gemma 4 12B

Open weights · Apache 2.0

Released June 2026, an 11.95B dense Gemma 4 model and the first mid-sized Gemma with an encoder-free design that feeds image and audio inputs directly into the transformer. Apache 2.0 license.

Pricing Open weights

No provider we track hosts it. Self-hosting starts at $0.05 an hour: see the estimate.

Google Cloud 1 provider

Key Specifications

Context window
256K tokens
Knowledge cutoff
Inputs
Text, Image, Audio Outputs: Text

Gemma 4 12B pricing by provider

Provider Input / 1M tokens Output / 1M tokens Cost at 10M in + 2M out
Google Cloud logo Google Cloud Creator Open weights, no hosted price View

Heads up: We do our best to keep these specs & prices accurate. However, cloud costs may fluctuate based on region, usage, and other factors not listed here. These are estimates based on common setups and are for informational purposes only. Always verify current rates & exact specs with the provider before provisioning. LLM rates are base-tier, on-demand prices per 1M tokens; cached-input, batch and long-context tiers are not included.

Compare every model at this volume in the LLM cost calculator.

Capabilities

Function calling

Function calling

Connect to external tools, APIs, and systems.

Structured output

Structured output

Return responses in structured formats like JSON.

Estimated cost to self-host Gemma 4 12B

Gemma 4 12B has 11.95B parameters. At 4-bit it needs about 7 GB of GPU memory, at 8-bit 14 GB and at BF16 29 GB, counting 20% on top of the weights for KV cache and runtime overhead.

Precision Memory needed Cheapest rentals that fit Per month, 24/7
4-bit 7 GB Nvidia logo GTX 1070 ($0.05/hr, tight) Nvidia logo RTX 3060 Ti ($0.07/hr, tight) $36
8-bit 14 GB Nvidia logo RTX 5060 Ti ($0.09/hr, tight) Nvidia logo RTX 4060 Ti ($0.10/hr, tight) $65
BF16 29 GB Nvidia logo RTX 5090 ($0.46/hr, tight) Nvidia logo RTX A6000 ($0.56/hr) $331

Memory is parameters × bytes per weight at each precision, plus 20% for KV cache and runtime overhead. Rentals are the cheapest cards that hold it, at provider-weighted median on-demand prices for the week of August 17, 2026, consumer cards included, in nodes of up to 8 GPUs; a fit with under 15% headroom is marked tight. A month is 720 hours. See the best-value GPUs guide for the same table across models, and cloud GPU pricing for every card.

More from Google Cloud

Model Context Input / 1M Output / 1M
Google Cloud logo Gemini 3.5 Flash-Lite 1M $0.30 $2.50
Google Cloud logo Gemini 3.6 Flash 1M $1.50 $7.50
Google Cloud logo Gemini 3.7 Flash 1M $1.50 $7.50
Google Cloud logo Gemini 2.5 Flash 1M $0.30 $2.50
Google Cloud logo Gemini 2.5 Flash-Lite 1M $0.10 $0.40
Google Cloud logo Gemini 2.5 Pro 1M $1.25 $10.00
Google Cloud logo Gemini 3 Flash Preview 1M $0.50 $3.00
Google Cloud logo Gemini 3 Pro Preview 1M

Models near this price

The models nearest this one by input rate, five each way, at each one's cheapest listed provider.

Frequently Asked Questions

How much does Gemma 4 12B cost?

Google Cloud publishes Gemma 4 12B's weights and lists no hosted price. Running it yourself costs whatever the hardware does; see cloud GPU pricing.

Which providers offer Gemma 4 12B?

One provider lists Gemma 4 12B: Google Cloud (open weights, no hosted price).

What is Gemma 4 12B's context window?

Gemma 4 12B accepts up to 256K tokens of input per request. The context window is the prompt plus any documents, conversation history and tool results sent with it; every token in it is billed at the input rate.

What is Gemma 4 12B's knowledge cutoff?

Gemma 4 12B's knowledge cutoff is January 2025: its training data runs up to that month and it has no built-in knowledge of later events.

What inputs and outputs does Gemma 4 12B support?

Gemma 4 12B accepts text, images and audio as input and produces text. Its listed capabilities are function calling and structured output.

Can I self-host Gemma 4 12B?

Yes. Google Cloud publishes Gemma 4 12B's weights under the Apache 2.0 license. At 4-bit it needs about 7 GB of GPU memory; the cheapest rental that holds it is GTX 1070 at $0.05 per hour, about $36 a month. The estimate above prices 8-bit and BF16 too.