Gemini 3.5 Flash-Lite: Google's Fastest 3.5-Class Model Reaches General Availability
Alongside Gemini 3.6 Flash, Google brought Gemini 3.5 Flash-Lite to general availability as the fastest and cheapest option in the 3.5 model family, built for high-volume automation and subagent workloads. The model runs at 350 output tokens per second, dramatically outperforms the prior Flash-Lite generation on coding tasks (54% vs. 31% on Terminal-Bench), and is priced at $0.30/$2.50 per million input/output tokens. It's available today through the Gemini API, Google AI Studio, Android Studio, and Enterprise platforms, and is rolling out into Google Search.
Key Takeaways
- 350 output tokens per second, making it the fastest model in the current Gemini 3.5 lineup.
- Coding accuracy jumped from 31% to 54% on Terminal-Bench versus the prior-generation Flash-Lite model.
- Priced at $0.30 input / $2.50 output per million tokens — the cheapest tier in the current Flash family.
- Despite its "lite" branding, it retains configurable thinking levels and built-in computer use.
- Available immediately via Gemini API, AI Studio, Android Studio, and Enterprise, and is rolling out into Google Search.
- Developer commentary grouped it with the 3.6 Flash launch, with builders welcoming the price/throughput tradeoff more warmly than the lack of a new Pro-tier model.
Sources & Mentions
5 external resources covering this update
Google Releases Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber: A Cheaper, More Token-Efficient Flash Tier Built for Agentic Workloads
MarkTechPost
Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber
Hacker News
Google's Gemini 3.6 Flash model cuts AI agent token costs by up to 65%... and 3.5 Pro is on the way
VentureBeat
Google reveals dev-focused Gemini 3.1 Flash-Lite, promises best-in-class intelligence for your highest-volume workloads
TechRadar
Google has released Gemini 3.6 Flash & Gemini 3.5 Flash-Lite
Overview
Google released Gemini 3.5 Flash-Lite as the newest entry in its Flash-Lite line, positioned as the fastest and most cost-effective model in the 3.5 family. Where Gemini 3.6 Flash targets general-purpose agentic work, Flash-Lite is built specifically for high-throughput, low-latency use cases: subagent orchestration, agentic search, and high-volume document processing where per-call cost and speed matter more than raw capability.
Performance and Speed
Google reported Gemini 3.5 Flash-Lite running at 350 output tokens per second, and the model shows a substantial jump over the prior generation on coding tasks — scoring 54% versus 31% for Gemini 3.1 Flash-Lite on Terminal-Bench. It also exceeds the performance of older, larger Gemini 3 Flash models on several benchmarks, making the cost/performance tradeoff notably better than the model's tier name might suggest.
Capabilities
Despite its "lite" positioning, the model supports configurable thinking levels, built-in computer use functionality, and performs well at agentic search and document processing — capabilities previously reserved for larger Flash-tier models.
Pricing and Availability
Gemini 3.5 Flash-Lite is priced at $0.30 per million input tokens and $2.50 per million output tokens, undercutting both the standard Flash tier and prior Flash-Lite generations. It's available immediately through the Gemini API, Google AI Studio, Android Studio, and Enterprise platforms, and Google says it is currently rolling out into Google Search — an indicator of how the company is using its own products to validate the model's real-world throughput.
Developer Reception
Coverage of the launch largely grouped Flash-Lite together with the Gemini 3.6 Flash announcement, with commentary noting that "builders welcomed the price and efficiency" of the new Flash tier as a whole. As with 3.6 Flash, some Hacker News discussion questioned why no accompanying Pro-tier model shipped alongside the Flash refresh, though the Flash-Lite pricing and throughput numbers specifically were well received as a strong fit for high-volume, cost-sensitive automation workloads.