Blog · 2026-07-22 · Vynaris Team
Gemini 3.6 Flash output falls to $7.50; Flash-Lite rises 67%
Google's Flash refresh moved two prices opposite ways: Gemini 3.6 Flash output fell 16.7% to $7.50 while 3.5 Flash-Lite output rose 67%. The per-workload re-route math, prices verified 2026-07-22.
Google's Flash-tier refresh (2026-07-21) moved two prices in opposite directions on the same page. Gemini 3.6 Flash keeps its $1.50 input rate but drops output from $9.00 to $7.50 per million tokens, a 16.7% cut. One row down, the new Gemini 3.5 Flash-Lite raises output from $1.50 to $2.50 (+66.7%) and input from $0.25 to $0.30 (+20%) against the 3.1 Flash-Lite it replaces. The mid-tier got cheaper; the budget tier got more expensive. Prices verified 2026-07-22.
That makes this a genuine re-route event, not a free upgrade. If you route output-heavy work to Flash, 3.6 pays you back. If you route cheap, high-volume traffic to Flash-Lite on the assumption a "refresh" holds the line, your next invoice is up to 67% higher on output. Below is the math on a real token shape, every price live-verified and every efficiency claim shown as a checkable derivation.
Verdict: who moves, who stays
Your workload What changed Move to
------------------------------------------ ---------------------------------------------------- -------------------------------------------------------------------------------------------------------------------------------------
Output-heavy agent (coding, tool loops) Output rate -16.7%, plus fewer output tokens claimed Test 3.6 Flash; the double win compounds
Input-heavy (summarize, classify, extract) Input unchanged at $1.50 Little reason to migrate; bill barely moves
Cheap high-volume on 3.1 Flash-Lite Output +66.7%, input +20% Re-price before migrating; 3.1 Flash-Lite is still listed
Sticker-shopping on price alone 3.6 Flash is not the cheapest per token Compare [Luna](https://vynaris.com/models#gpt-5-6-luna) / [Haiku 4.5](https://vynaris.com/models#claude-haiku-4-5) on your own tokensThe price cut is real; the token claim is Google's
Two separate numbers drive 3.6 Flash, and they carry different weights of evidence.
The price is a fact, read live twice today on Google's pricing page: $1.50 input, $7.50 output. That output rate is 16.7% below the $9.00 the 3.5 Flash predecessor charged, at the same input.
The token efficiency is a vendor-and-benchmark claim, not a price. Google reports 3.6 Flash emits fewer output tokens per task: about 17% fewer on the independent Artificial Analysis Intelligence Index (28,000 down to 23,000 tokens on their reported run), and up to 65% fewer on DeepSWE v1.1, a long-horizon coding benchmark (276,000 down to 97,000 tokens). We print the absolute counts so you can check the percentages: 23,000 / 28,000 = 0.821, a 17.9% cut; 97,000 / 276,000 = 0.351, a 64.9% cut. These are Google's and a benchmark's numbers, not ours, and they will not generalize to every workload.
Stack the verified price cut with the claimed token cut and the effective output cost falls further than the sticker, because you pay a lower rate on fewer tokens:
effective output cost vs 3.5 Flash output-cost cut
-----------------------------------------------------
sticker only (0.833x price) 16.7%
x AA-Index tokens (0.821x) 31.5%
x DeepSWE tokens (0.351x) 70.7%Read that as a ceiling, not a promise. It touches output tokens only, and only if your traffic behaves like the benchmark.
The total-task saving depends on your input:output ratio
"About 30% cheaper" is an average, and the average hides the routing decision. The 3.6 Flash win comes almost entirely from output, so the saving scales with how output-heavy your task is, not with total token count. Two representative shapes, priced per 1,000 tasks:
Workload Tokens (in / out) 3.5 Flash 3.6 Flash, same tokens + AA-Index -17% + DeepSWE -65%
-------------- ----------------- --------- ---------------------- --------------- ---------------
Summarization 12,000 / 600 $23.40 $22.50 (-3.8%) $21.70 (-7.3%) $19.58 (-16.3%)
Agentic coding 20,000 / 8,000 $102.00 $90.00 (-11.8%) $79.29 (-22.3%) $51.09 (-49.9%)
Same refresh, opposite outcomes. The summarization job is input-dominated, input did not change, and even a generous token cut barely dents the bill. The agentic loop is output-dominated, so it captures both the rate cut and the fewer-tokens claim. Before you migrate anything, run your own token shape through the calculator rather than trusting a benchmark average.
Flash-Lite went the other way
The budget tier moved against you. On the new Gemini 3.5 Flash-Lite, the hike depends on output share too, and it cuts the opposite direction from Flash:
Cheap-tier workload Tokens (in / out) 3.1 Flash-Lite 3.5 Flash-Lite Change
------------------- ----------------- -------------- -------------- ------
Classification 2,000 / 100 $0.65 $0.85 +30.8%
Short generation 1,000 / 1,000 $1.75 $2.80 +60.0%If your budget tier does output-light classification, the +20% input hike dominates and you pay about 31% more. If it generates, the +67% output hike stings for 60%. The one piece of good news: the old 3.1 Flash-Lite is still listed at $0.25 / $1.50, so the cheapest call did not disappear, it just is not the new default. A "refresh" is a pricing event; re-price your actual traffic on both rows before you let a client library pick the newer string for you.
The honest part: 3.6 Flash is not the cheapest sticker
Credibility first. On identical token counts, 3.6 Flash is not the cheapest per-token option for output-heavy work. On the agentic shape (20,000 in / 8,000 out), priced with no efficiency credited to anyone, per 1,000 tasks:
Gemini 3.5 Flash-Lite $ 26.00 (smaller model; quality tradeoff)
Claude Haiku 4.5 $ 60.00
GPT-5.6 Luna $ 68.00
Gemini 3.6 Flash $ 90.00
Claude Sonnet 5 (intro) $120.00GPT-5.6 Luna undercuts 3.6 Flash by $22.00 per 1,000 tasks (24.4%) on the same tokens. So the entire 3.6 Flash case rests on the fewer-tokens claim, which is the one number you cannot read off a pricing page. The break-even is concrete: 3.6 Flash must emit at most 5,067 output tokens against Luna's 8,000 on this shape, a 36.7% token cut, to draw level on price. That sits between the AA-Index 17% and the DeepSWE 65% figures, which means the answer is workload-specific and only your own eval settles it. Do not migrate on the vendor benchmark; measure your token count on your traffic first. This is the same lesson our GPT-5.6 migration teardown reached from the other side: token efficiency is real but small next to a deliberate model-routing decision.
What to actually do this week
- Output-heavy agents: benchmark 3.6 Flash on your traces and count output tokens per task. If you see even the modest 17% cut, the compounding rate-plus-token effect makes it a straightforward right-sizing win over the old 3.5 Flash.
- Input-heavy jobs: skip the migration noise. Input did not change, so your bill does not either.
- Anyone on 3.1 Flash-Lite: treat the new Flash-Lite as a price increase, not an upgrade. Re-price, and keep the 3.1 string if it is still the cheaper call for your shape.
- Router users: this is a routing table edit, not a rewrite. Point output-heavy inference at 3.6 Flash, keep the planning step on a frontier tier, keep cheap classification where it is cheapest, and read the invoice.
For the full cross-provider rate sheet behind every model named here, see the July 2026 LLM price list; for the per-tier task math on the OpenAI foils, the GPT-5.6 Sol vs Terra vs Luna breakdown. A cheaper right-sized model only helps your agent if it clears your quality bar, so route the new tier onto steps where a wrong answer is caught cheaply, never onto planning.
FAQ
Is Gemini 3.6 Flash cheaper than 3.5 Flash? On output, yes: $7.50 vs $9.00 per million tokens, a 16.7% cut, with input unchanged at $1.50. Total-task savings depend on your input:output ratio and range from about 4% on summarization to double digits on output-heavy agents.
Did Gemini 3.5 Flash-Lite get cheaper? No. It rose versus the 3.1 Flash-Lite it succeeds: input +20% ($0.25 to $0.30), output +66.7% ($1.50 to $2.50). Real-workload increases run about 31% to 60% depending on output share.
Where does the "30% cheaper" figure come from? From stacking the verified 16.7% output-rate cut with Google's claim of roughly 17% fewer output tokens per task. The 30% is an effective output-cost figure, not a total-bill figure, and it is contingent on the token claim holding on your traffic.
Is 3.6 Flash the cheapest option? No. On identical tokens, GPT-5.6 Luna, Claude Haiku 4.5, and even the new Flash-Lite are cheaper per token. The 3.6 Flash argument rests on emitting fewer tokens per task, which you must verify on your own workload.
Are the token-efficiency numbers reliable? They are vendor-and-benchmark-reported (Google, Artificial Analysis Index, DeepSWE v1.1), not independently measured by us. We show the absolute token counts so you can check the percentages, and we recommend counting tokens on your own traces before crediting any model with a saving.
Sources
- Google Gemini API pricing (3.6 Flash $1.50/$7.50, 3.5 Flash $1.50/$9.00, 3.5 Flash-Lite $0.30/$2.50, 3.1 Flash-Lite $0.25/$1.50), read live twice 2026-07-22: https://ai.google.dev/gemini-api/docs/pricing
- OpenRouter, Gemini 3.6 Flash ($1.50/$7.50, 1M context), 2026-07-22: https://openrouter.ai/google/gemini-3.6-flash
- Google release announcement (token-efficiency claims, benchmark scores), 2026-07-21: https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-6-flash-3-5-flash-lite-3-5-flash-cyber/
- VentureBeat, token-cost cut on long-horizon tasks (DeepSWE absolute counts), 2026-07-21: https://venturebeat.com/technology/googles-gemini-3-6-flash-model-cuts-ai-agent-token-costs-by-up-to-65-on-long-horizon-engineering-tasks-and-3-5-pro-is-on-the-way
- OpenAI API pricing (GPT-5.6 Luna $1/$6), verified 2026-07-22: https://developers.openai.com/api/docs/pricing
- Anthropic pricing (Sonnet 5 intro $2/$10, Haiku 4.5 $1/$5), verified 2026-07-22: https://platform.claude.com/docs/en/docs/about-claude/pricing
- All arithmetic:
gemini-3-6-flash-vs-3-5-flash-lite-price-change-reroute-math-math.py, assumptions editable inline.
Related: July 2026 LLM price list · GPT-5.6 migration teardown