跪拜 Guibai
← All articles
Gemini · AI Coding · AIGC

Gemini 3.6 Flash Cuts Token Costs 17% but Flatlines on Intelligence Benchmarks

By ServBay ·
Read original on juejin.cn ↗ Google Translate ↗ Alt translation

A model that gets cheaper without getting smarter changes the unit economics of AI deployment but not the capability ceiling. Teams running high-volume agent workloads will see lower bills, but anyone choosing a model for reasoning quality has no reason to switch.

Summary

Google's Gemini 3.6 Flash ships as a cost-optimized workhorse model with a 17% reduction in output tokens and a corresponding price cut to $7.50 per million output tokens. Official benchmarks claim double-digit gains on DeepSWE and MLE-Bench, and the knowledge cut-off advances to March 2026. A companion Flash-Lite model undercuts it further at $2.50 per million output tokens while beating the older Gemini 3 Flash on several agent tasks.

Third-party evaluations tell a different story. Artificial Analysis scores 3.6 Flash's Intelligence Index at 50 — flat against 3.5 Flash and trailing Meta Spark 1.1, GLM-5.2, Sonnet 5, and Grok 4.5. Developers also flag that its cost per unit of intelligence now exceeds GPT-5.6 Sol Medium, eroding the Flash line's traditional value proposition.

The release arrives as Google's flagship Gemini 3.5 Pro remains delayed. Reports indicate code-generation targets were missed and training was restarted from scratch, while public messaging has already shifted toward Gemini 4.

Takeaways
Output token usage drops roughly 17% versus 3.5 Flash, with savings reaching 65% on the DeepSWE software-engineering benchmark.
Pricing falls to $1.50 per million input tokens and $7.50 per million output tokens, a 17% reduction on the output side.
Artificial Analysis rates the Intelligence Index at 50, unchanged from 3.5 Flash and below several competing models.
Flash-Lite, priced at $0.30 input / $2.50 output per million tokens, scores 54.2% on SWE-Bench Pro, surpassing the older Gemini 3 Flash at 49.6%.
Computer Use is now a built-in API tool requiring no extra configuration.
Gemini 3.5 Pro has missed internal code-generation targets; training was reportedly restarted from scratch.
Gemini 3.5 Flash Cyber, a security-vulnerability detection model, is limited to trusted partners and not publicly available.
Conclusions

Google is optimizing a model whose raw intelligence hasn't moved, which turns the Flash line into a cost-reduction story rather than a capability story.

Flash-Lite outperforming the older full Flash on agent benchmarks suggests Google's distillation and training recipes improved even when flagship scaling stalled.

Keeping the vulnerability-detection model closed to all but trusted partners is a tacit admission that better offensive security tooling carries dual-use risk.

Shifting public narrative to Gemini 4 while 3.5 Pro remains undelivered is a classic expectation-management tactic that signals internal trouble with the current generation.

Concepts & terms
DeepSWE
A software engineering benchmark by Datacurve that measures an AI model's ability to resolve real-world coding tasks, often used to evaluate agentic coding performance.
MLE-Bench
A benchmark focused on machine-learning engineering tasks, testing a model's ability to handle end-to-end ML workflows rather than isolated code snippets.
OSWorld-Verified
A benchmark for computer-use agents that evaluates how well a model can operate a desktop environment — clicking, typing, and navigating interfaces to complete tasks.
Intelligence Index
A composite metric from Artificial Analysis that aggregates performance across multiple reasoning and knowledge benchmarks into a single score for model comparison.
Source: juejin.cn ↗ Google Translate ↗ Backup ↗