Gemini 3.5 Flash-Lite: Google's Fastest 3.5-Class Model Reaches General Availability

Gemini CLIView original changelog

Alongside Gemini 3.6 Flash, Google brought Gemini 3.5 Flash-Lite to general availability as the fastest and cheapest option in the 3.5 model family, built for high-volume automation and subagent workloads. The model runs at 350 output tokens per second, dramatically outperforms the prior Flash-Lite generation on coding tasks (54% vs. 31% on Terminal-Bench), and is priced at $0.30/$2.50 per million input/output tokens. It's available today through the Gemini API, Google AI Studio, Android Studio, and Enterprise platforms, and is rolling out into Google Search.

Key Takeaways

  • 350 output tokens per second, making it the fastest model in the current Gemini 3.5 lineup.
  • Coding accuracy jumped from 31% to 54% on Terminal-Bench versus the prior-generation Flash-Lite model.
  • Priced at $0.30 input / $2.50 output per million tokens — the cheapest tier in the current Flash family.
  • Despite its "lite" branding, it retains configurable thinking levels and built-in computer use.
  • Available immediately via Gemini API, AI Studio, Android Studio, and Enterprise, and is rolling out into Google Search.
  • Developer commentary grouped it with the 3.6 Flash launch, with builders welcoming the price/throughput tradeoff more warmly than the lack of a new Pro-tier model.

Overview

Google released Gemini 3.5 Flash-Lite as the newest entry in its Flash-Lite line, positioned as the fastest and most cost-effective model in the 3.5 family. Where Gemini 3.6 Flash targets general-purpose agentic work, Flash-Lite is built specifically for high-throughput, low-latency use cases: subagent orchestration, agentic search, and high-volume document processing where per-call cost and speed matter more than raw capability.

Performance and Speed

Google reported Gemini 3.5 Flash-Lite running at 350 output tokens per second, and the model shows a substantial jump over the prior generation on coding tasks — scoring 54% versus 31% for Gemini 3.1 Flash-Lite on Terminal-Bench. It also exceeds the performance of older, larger Gemini 3 Flash models on several benchmarks, making the cost/performance tradeoff notably better than the model's tier name might suggest.

Capabilities

Despite its "lite" positioning, the model supports configurable thinking levels, built-in computer use functionality, and performs well at agentic search and document processing — capabilities previously reserved for larger Flash-tier models.

Pricing and Availability

Gemini 3.5 Flash-Lite is priced at $0.30 per million input tokens and $2.50 per million output tokens, undercutting both the standard Flash tier and prior Flash-Lite generations. It's available immediately through the Gemini API, Google AI Studio, Android Studio, and Enterprise platforms, and Google says it is currently rolling out into Google Search — an indicator of how the company is using its own products to validate the model's real-world throughput.

Developer Reception

Coverage of the launch largely grouped Flash-Lite together with the Gemini 3.6 Flash announcement, with commentary noting that "builders welcomed the price and efficiency" of the new Flash tier as a whole. As with 3.6 Flash, some Hacker News discussion questioned why no accompanying Pro-tier model shipped alongside the Flash refresh, though the Flash-Lite pricing and throughput numbers specifically were well received as a strong fit for high-volume, cost-sensitive automation workloads.