Gemini 2.0 Flash-Lite
The lightest Gemini 2.0 model, built for high-throughput text generation.
Key Specifications
- Context window
- 1M tokens
- Max output
- 8K tokens
- Knowledge cutoff
- Released
- Inputs
- Text, image, video, audio
- Capabilities Show details
- Function calling Connect to external tools, APIs, and systems. Structured output Return responses in structured formats like JSON.
Frequently Asked Questions
What is Gemini 2.0 Flash-Lite good for?
High-throughput text generation from multimodal input, at the light end of the Gemini 2.0 lineup.
When is Gemini 2.0 Flash-Lite not a good fit?
Any new work: Google retired it in June 2026. It had no web search or code execution, and its short maximum output limited long answers.
How does Gemini 2.0 Flash-Lite compare with Gemini 2.5 Flash-Lite?
Google says Gemini 2.5 Flash-Lite has higher quality than Gemini 2.0 Flash-Lite across coding, math, science, reasoning and multimodal tasks, with lower latency, and adds optional thinking plus web search and code execution. Same 1M-token context and text, image, audio and video input. Gemini 2.0 Flash-Lite came out 5 months earlier and has an older knowledge cutoff (August 2024 against January 2025) and a smaller max output (8K against 66K tokens).
Can I self-host Gemini 2.0 Flash-Lite?
No. Google Cloud does not publish the weights; Gemini 2.0 Flash-Lite is available only through hosted APIs.
More from Google Cloud
| Model | Context | Input / 1M | Output / 1M |
|---|---|---|---|
|
|
1M | $0.30 | $2.50 |
|
|
1M | $0.75 | $3.75 |
|
|
1M | $0.75 | $3.75 |
|
|
1M | $0.75 | $3.75 |
|
|
1M | $0.30 | $2.50 |
|
|
1M | $0.10 | $0.40 |
|
|
1M | $1.25 | $10.00 |
|
|
1M | $0.50 | $3.00 |