Google Launches Gemini 3.6 Flash as Pro Version Stalls in Testing
In brief
- Gemini 3.6 Flash outperforms 3.5 Flash on coding benchmarks at faster, cheaper pricing.
- Gemini 3.5 Pro remains missing after Google missed its own deadline due to coding task failures.
- Flash-Lite model runs at 350 tokens per second, costs $0.30 per million input tokens.
- Gemini 3.5 Flash Cyber restricted to governments and vetted security partners only.
- Alphabet stock fell 4.4% on Pro delay reports, erasing roughly $200 billion in market cap.
The Flash Strategy
Google launched three new AI models today: Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber. The Flash series is Google's line of speed-optimized models—fast, cost-effective, and built for AI agents, which are programs that operate semi-autonomously to handle tasks like managing documents, processing data pipelines, or browsing the web without a human clicking through each step.
The 3.6 Flash iteration shows real gains. It uses 17% fewer output tokens than 3.5 Flash and costs $1.50 per million input tokens and $7.50 per million output tokens, down from $9 on the output side for 3.5 Flash. On benchmarks, it scored 49% on DeepSWE v1.1 (a test for long-horizon software engineering) versus 37% for 3.5 Flash. On MLE-Bench, a machine learning engineering test, it scored 63.9% versus 49.7%.
It topped the table on OSWorld-Verified—a test where the AI takes control of a computer screen to complete real tasks—at 83.0%, ahead of Claude Sonnet 5 at 81.2% and GPT-5.6 Luna at 72.6%. But real-world coding work tells a different story. In testing, the model produced an unusable HTML file with improper formatting and rendering errors on a simple coding task. Deepseek identified 11 bugs in the output and implemented 8 key fixes to make the code functional.
The Pro Problem
Here's where the story gets uncomfortable for Google. After unveiling Gemini 3.5 Flash at Google I/O 2026 in May and promising a Pro version within a month, Google quietly missed its own deadline. Gemini 3.5 Pro was held back because it fell short of internal targets, particularly on coding tasks. A late-June attempt to fix it by updating the training data produced disappointing results.
The last Pro-tier model Google shipped was Gemini 3.1 Pro back in February. That's a six-month gap at a moment when OpenAI is shipping GPT-4o variants and Anthropic is rolling out Claude Sonnet 5. The delay carries real cost: Alphabet stock fell roughly 4.4% on the report, erasing an estimated $200 billion in market cap in a single session.
The Smaller Models
The Flash-Lite variant is built purely for volume: 350 output tokens per second at $0.30 per million input tokens and $2.50 per million output tokens. It outperforms the older 3 Flash on key coding tasks, including Terminal-Bench 2.1 (54% vs. 31%).
The third model, Gemini 3.5 Flash Cyber, won't be publicly available. Google is restricting it to governments and vetted partners for finding and fixing software vulnerabilities. Pro models are the heavy lifters: slower, pricier, and built for complex reasoning where raw power matters more than speed. Until Google ships a working Pro version, that gap stays open.


