Google Cloud logo Retired Jun 2026

Gemini 2.0 Flash-Lite

The lightest Gemini 2.0 model, built for high-throughput text generation.

Key Specifications

Context window
1M tokens
Max output
8K tokens
Knowledge cutoff
Released
Inputs
Text, image, video, audio
Capabilities Show details
Function calling Connect to external tools, APIs, and systems. Structured output Return responses in structured formats like JSON.

Frequently Asked Questions

What is Gemini 2.0 Flash-Lite good for?

High-throughput text generation from multimodal input, at the light end of the Gemini 2.0 lineup.

When is Gemini 2.0 Flash-Lite not a good fit?

Any new work: Google retired it in June 2026. It had no web search or code execution, and its short maximum output limited long answers.

How does Gemini 2.0 Flash-Lite compare with Gemini 2.5 Flash-Lite?

Google says Gemini 2.5 Flash-Lite has higher quality than Gemini 2.0 Flash-Lite across coding, math, science, reasoning and multimodal tasks, with lower latency, and adds optional thinking plus web search and code execution. Same 1M-token context and text, image, audio and video input. Gemini 2.0 Flash-Lite came out 5 months earlier and has an older knowledge cutoff (August 2024 against January 2025) and a smaller max output (8K against 66K tokens).

Can I self-host Gemini 2.0 Flash-Lite?

No. Google Cloud does not publish the weights; Gemini 2.0 Flash-Lite is available only through hosted APIs.

More from Google Cloud

Model Context Input / 1M Output / 1M
Google Cloud Gemini 3.5 Flash-Lite 1M $0.30 $2.50
Google Cloud Gemini 3.6 Flash 1M $0.75 $3.75
Google Cloud Gemini 3.7 Flash 1M $0.75 $3.75
Google Cloud Gemini 3.8 Flash 1M $0.75 $3.75
Google Cloud Gemini 2.5 Flash 1M $0.30 $2.50
Google Cloud Gemini 2.5 Flash-Lite 1M $0.10 $0.40
Google Cloud Gemini 2.5 Pro 1M $1.25 $10.00
Google Cloud Gemini 3 Flash Preview 1M $0.50 $3.00