VynarisEarly betaGet your API key

Gemini 3.6 Flash output falls to $7.50; Flash-Lite rises 67%

Google's Flash refresh moved two prices opposite ways: Gemini 3.6 Flash output fell 16.7% to $7.50 while 3.5 Flash-Lite output rose 67%. The per-workload re-route math, prices verified 2026-07-22.

Google's Flash-tier refresh (2026-07-21) moved two prices in opposite directions on the same page. Gemini 3.6 Flash keeps its $1.50 input rate but drops output from $9.00 to $7.50 per million tokens, a 16.7% cut. One row down, the new Gemini 3.5 Flash-Lite raises output from $1.50 to $2.50 (+66.7%) and input from $0.25 to $0.30 (+20%) against the 3.1 Flash-Lite it replaces. The mid-tier got cheaper; the budget tier got more expensive. Prices verified 2026-07-22.

That makes this a genuine re-route event, not a free upgrade. If you route output-heavy work to Flash, 3.6 pays you back. If you route cheap, high-volume traffic to Flash-Lite on the assumption a "refresh" holds the line, your next invoice is up to 67% higher on output. Below is the math on a real token shape, every price live-verified and every efficiency claim shown as a checkable derivation.

Verdict: who moves, who stays

Your workload                               What changed                                          Move to
------------------------------------------  ----------------------------------------------------  -------------------------------------------------------------------------------------------------------------------------------------
Output-heavy agent (coding, tool loops)     Output rate -16.7%, plus fewer output tokens claimed  Test 3.6 Flash; the double win compounds
Input-heavy (summarize, classify, extract)  Input unchanged at $1.50                              Little reason to migrate; bill barely moves
Cheap high-volume on 3.1 Flash-Lite         Output +66.7%, input +20%                             Re-price before migrating; 3.1 Flash-Lite is still listed
Sticker-shopping on price alone             3.6 Flash is not the cheapest per token               Compare [Luna](https://vynaris.com/models#gpt-5-6-luna) / [Haiku 4.5](https://vynaris.com/models#claude-haiku-4-5) on your own tokens

The price cut is real; the token claim is Google's

Two separate numbers drive 3.6 Flash, and they carry different weights of evidence.

The price is a fact, read live twice today on Google's pricing page: $1.50 input, $7.50 output. That output rate is 16.7% below the $9.00 the 3.5 Flash predecessor charged, at the same input.

The token efficiency is a vendor-and-benchmark claim, not a price. Google reports 3.6 Flash emits fewer output tokens per task: about 17% fewer on the independent Artificial Analysis Intelligence Index (28,000 down to 23,000 tokens on their reported run), and up to 65% fewer on DeepSWE v1.1, a long-horizon coding benchmark (276,000 down to 97,000 tokens). We print the absolute counts so you can check the percentages: 23,000 / 28,000 = 0.821, a 17.9% cut; 97,000 / 276,000 = 0.351, a 64.9% cut. These are Google's and a benchmark's numbers, not ours, and they will not generalize to every workload.

Stack the verified price cut with the claimed token cut and the effective output cost falls further than the sticker, because you pay a lower rate on fewer tokens:

effective output cost vs 3.5 Flash    output-cost cut
-----------------------------------------------------
sticker only (0.833x price)                 16.7%
x AA-Index tokens (0.821x)                  31.5%
x DeepSWE tokens  (0.351x)                  70.7%

Read that as a ceiling, not a promise. It touches output tokens only, and only if your traffic behaves like the benchmark.

The total-task saving depends on your input:output ratio

"About 30% cheaper" is an average, and the average hides the routing decision. The 3.6 Flash win comes almost entirely from output, so the saving scales with how output-heavy your task is, not with total token count. Two representative shapes, priced per 1,000 tasks:

Workload        Tokens (in / out)  3.5 Flash  3.6 Flash, same tokens  + AA-Index -17%  + DeepSWE -65%
--------------  -----------------  ---------  ----------------------  ---------------  ---------------
Summarization   12,000 / 600       $23.40     $22.50 (-3.8%)          $21.70 (-7.3%)   $19.58 (-16.3%)
Agentic coding  20,000 / 8,000     $102.00    $90.00 (-11.8%)         $79.29 (-22.3%)  $51.09 (-49.9%)
Two bar panels: summarization barely moves from 3.5 to 3.6 Flash, agentic coding drops up to 50%
Cost per 1,000 tasks: the 3.6 Flash saving scales with output share. Summarization is input-dominated and barely moves; the agentic loop captures the compounding cut. Prices verified 2026-07-22; token cuts vendor-reported.

Same refresh, opposite outcomes. The summarization job is input-dominated, input did not change, and even a generous token cut barely dents the bill. The agentic loop is output-dominated, so it captures both the rate cut and the fewer-tokens claim. Before you migrate anything, run your own token shape through the calculator rather than trusting a benchmark average.

Flash-Lite went the other way

The budget tier moved against you. On the new Gemini 3.5 Flash-Lite, the hike depends on output share too, and it cuts the opposite direction from Flash:

Cheap-tier workload  Tokens (in / out)  3.1 Flash-Lite  3.5 Flash-Lite  Change
-------------------  -----------------  --------------  --------------  ------
Classification       2,000 / 100        $0.65           $0.85           +30.8%
Short generation     1,000 / 1,000      $1.75           $2.80           +60.0%

If your budget tier does output-light classification, the +20% input hike dominates and you pay about 31% more. If it generates, the +67% output hike stings for 60%. The one piece of good news: the old 3.1 Flash-Lite is still listed at $0.25 / $1.50, so the cheapest call did not disappear, it just is not the new default. A "refresh" is a pricing event; re-price your actual traffic on both rows before you let a client library pick the newer string for you.

The honest part: 3.6 Flash is not the cheapest sticker

Credibility first. On identical token counts, 3.6 Flash is not the cheapest per-token option for output-heavy work. On the agentic shape (20,000 in / 8,000 out), priced with no efficiency credited to anyone, per 1,000 tasks:

Gemini 3.5 Flash-Lite      $ 26.00   (smaller model; quality tradeoff)
Claude Haiku 4.5           $ 60.00
GPT-5.6 Luna               $ 68.00
Gemini 3.6 Flash           $ 90.00
Claude Sonnet 5 (intro)    $120.00

GPT-5.6 Luna undercuts 3.6 Flash by $22.00 per 1,000 tasks (24.4%) on the same tokens. So the entire 3.6 Flash case rests on the fewer-tokens claim, which is the one number you cannot read off a pricing page. The break-even is concrete: 3.6 Flash must emit at most 5,067 output tokens against Luna's 8,000 on this shape, a 36.7% token cut, to draw level on price. That sits between the AA-Index 17% and the DeepSWE 65% figures, which means the answer is workload-specific and only your own eval settles it. Do not migrate on the vendor benchmark; measure your token count on your traffic first. This is the same lesson our GPT-5.6 migration teardown reached from the other side: token efficiency is real but small next to a deliberate model-routing decision.

What to actually do this week

For the full cross-provider rate sheet behind every model named here, see the July 2026 LLM price list; for the per-tier task math on the OpenAI foils, the GPT-5.6 Sol vs Terra vs Luna breakdown. A cheaper right-sized model only helps your agent if it clears your quality bar, so route the new tier onto steps where a wrong answer is caught cheaply, never onto planning.

FAQ

Is Gemini 3.6 Flash cheaper than 3.5 Flash? On output, yes: $7.50 vs $9.00 per million tokens, a 16.7% cut, with input unchanged at $1.50. Total-task savings depend on your input:output ratio and range from about 4% on summarization to double digits on output-heavy agents.

Did Gemini 3.5 Flash-Lite get cheaper? No. It rose versus the 3.1 Flash-Lite it succeeds: input +20% ($0.25 to $0.30), output +66.7% ($1.50 to $2.50). Real-workload increases run about 31% to 60% depending on output share.

Where does the "30% cheaper" figure come from? From stacking the verified 16.7% output-rate cut with Google's claim of roughly 17% fewer output tokens per task. The 30% is an effective output-cost figure, not a total-bill figure, and it is contingent on the token claim holding on your traffic.

Is 3.6 Flash the cheapest option? No. On identical tokens, GPT-5.6 Luna, Claude Haiku 4.5, and even the new Flash-Lite are cheaper per token. The 3.6 Flash argument rests on emitting fewer tokens per task, which you must verify on your own workload.

Are the token-efficiency numbers reliable? They are vendor-and-benchmark-reported (Google, Artificial Analysis Index, DeepSWE v1.1), not independently measured by us. We show the absolute token counts so you can check the percentages, and we recommend counting tokens on your own traces before crediting any model with a saving.

Sources

Related: July 2026 LLM price list · GPT-5.6 migration teardown