How Google’s New Gemini Flash Models Compare to its Rivals

Share this article
Share this article
Prioritise Us on Google
Google steps up the competition against AI rivals with a faster, more cost-effective Gemini Flash model lineup. Credit: Getty Images
Google's newly released Gemini 3.6 Flash excels at document drafting and review while completing tasks 12% faster on average, says Harvey’s Niko Grupen

Trailing Anthropic and OpenAI in model rollout timing this year, Google is upping the competition with a faster, leaner Gemini lineup designed to undercut rivals, including its Chinese competitors on both price and performance.

The tech giant aims to deliver higher token efficiency, lower latency and more reliable execution for developers with the release that aims to build high-volume production AI agents.

Headlining the launch is Gemini 3.5 Flash Cyber, arriving roughly a month after Anthropic’s Mythos and Fable security models faced an abrupt freeze and ease in restriction following its launch. 

Positioned to compete directly with these offerings, Google introduces three distinct, cost-effective models engineered to address specific throughput, security and cost requirements.

The new Gemini Flash models offer lower token costs and enhanced speed for production AI agents. Credit: Google

The new competitors

Building on Gemini 3.5 Flash and meeting the sweet spot of efficiency and quality, the Flash series of models include: 

  • Gemini 3.6 Flash: Delivering improved coding, knowledge work and multimodal performance while reducing output token usage by 17% on the Artificial Analysis Index at a lower cost per output token
  • Gemini 3.5 Flash-Lite: Operating as the fastest, most cost-effective 3.5-class model, processing 350 output tokens per second on the Artificial Analysis Index to outperform prior generations in agentic workflows
  • Gemini 3.5 Flash Cyber in CodeMender: Combining a highly efficient, specialised cyber-focused model with the CodeMender security agent to deliver competitive frontier performance.

Beyond these releases, Gemini 3.5 Pro is undergoing partner testing ahead of its broader availability.

The team has also teased the next generation of models as it has started its ambitious pre-training run for Gemini 4.

These developments mark Alphabet’s clearest strategy yet to counter Anthropic in automated code defence and accelerate high-volume AI agent deployment.

Gemini 3.6 Flash and 3.5 Flash-Lite have been made available via Google AI Studio, Android Studio, Google Antigravity, Gemini Enterprise and the Gemini App.

Youtube Placeholder

Strong coding and computer use gains

Artificial Analysis data shows Gemini Flash undercuts comparable models from Anthropic, OpenAI and Chinese rivals on cost. It is cheaper per task than GPT-5.6 Terra Max, Kimi K3 and Qwen 3.7 Max, which compete for high-volume production workloads.

Gemini 3.6 Flash delivers higher precision with fewer unwanted code edits on DeepSWE (49% vs 37%). It also shows significant improvement in machine learning research on MLE Bench (63.9% vs 49.7%).

Computer use capabilities improve on OSWorld-Verified (83% vs 78.4%). The feature operates as a built-in client-side tool via the Gemini API and Gemini Enterprise.

The model outperforms 3.5 Flash in knowledge work on GDPval-AA v2 (1,421 vs 1,349). Customers like Hebbia and Harvey find it capable at document parsing, chart analysis and report drafting.

Niko Grupen, Head of Applied Research at Harvey, says: “Gemini 3.6 Flash excels at document drafting and review in practice areas like capital markets and corporate M&A. 

“Compared to its predecessor, Gemini 3.6 Flash showed strong gains in performance on our benchmarks and was notably more efficient completing tasks 12% faster on average.”

Niko Grupen, Head of Applied Research at Harvey. Credit: Niko Grupen/LinkedIn

Scaling high-volume agentic tasks

Gemini 3.5 Flash-Lite is designed for low-latency tasks and high-throughput workflows like agentic search and document processing. Running at 350 output tokens per second, it is the fastest model in the 3.5 series.

Priced at US$0.30 per 1m input tokens and US$2.50 per 1m output tokens, it offers a strong price-to-performance ratio. Across thinking levels, it outperforms 3.1 Flash-Lite to enable efficient scaling.

Developers can configure minimal thinking levels for high-volume execution or engage higher thinking levels for multi-step subagent workloads. Computer use acts as a built-in tool to support agentic tasks across surfaces.

It improves coding and agentic tasks on Terminal-Bench 2.1 (54% vs 31%), long context on GDM-MRCR v2 (72.2% vs 60.1%) and real-world execution on GDPval-AA v2 (1,140 vs 642). 

On SWE-Bench Pro (54.2% vs 49.6%) and OSWorld-Verified (74% vs 65.1%), it outperforms 3 Flash.

The model extracts product features from massive ecommerce datasets and synthesises data. Working alongside 3.6 Flash as master agent, it generates 25 web design concepts instantly.

It scales receipt translation with multimodal understanding and builds games by iterating through multiple options. Early customers have highlighted its speed, intelligence and cost efficiency.

Gemini 3.5 Flash Lite compared to 3 Flash. Credit: Google

Strengthening automated code defence

Conceding to the reality of modern cyber threats, Google highlights how AI models find security vulnerabilities faster than systems can fix them. 

Tackling this threat, the company takes a unique approach that is both capable and efficient. Gemini 3.5 Flash Cyber is built on 3.5 Flash and fine-tuned for finding and fixing cybersecurity vulnerabilities at a lower price per token. 

Within CodeMender, multiple 3.5 Flash Cyber agents work together to produce a single combined report.

The combination reaches competitive performance at the frontier on the benchmark CyberGym. The model helps Google narrow its cybersecurity gap with Anthropic, which built an early lead with its Mythos model.

The model will be exclusively available to governments and trusted partners via CodeMender as part of a limited pilot program. This gives frontline defenders a head start in fixing critical vulnerabilities before exploitation.

On the other side, Chinese rivals are also gaining momentum as Moonshot AI limits Kimi K3 subscriptions due to capacity constraints, while Alibaba teases Qwen 3.8 Max. Building a competitive model requires enough computing capacity to serve it at scale.

Google has an advantage through custom chips, cloud infrastructure and co-designing models and hardware together. 

With this balance of speed, cost and performance, the new Flash lineup solidifies the company’s position in competing for production AI agent dominance as developers seek faster and more cost-effective alternatives to existing market rivals.

Executives