Blog · 2026-08-15 · Vynaris Team
Gemini 3.6 Flash is 50% cheaper until Dec 31: $13.50 per 1,000 tasks
Gemini 3.6 Flash now costs $13.50 per 1,000 8k/2k tasks, half its Jan. 1 rate. Every published service-tier row doubles in 2027.
Gemini 3.6 Flash now costs $13.50 per 1,000 tasks with 8,000 input and 2,000 output tokens. The same workload becomes $27.00 on 2027-01-01. Every published Standard, Batch, Flex, Priority and cache row doubles then. Prices verified 2026-08-15.
Updated 2026-08-16: Google added 3.7 Flash at the same rate. Our Gemini 3.7 benchmark-normalized cost analysis separates the unchanged token sticker from its higher Google-reported coding scores.
TL;DR
- Standard pricing is $0.75 per million input tokens and $3.75 per million output tokens through 2026-12-31. It becomes $1.50/$7.50 on 2027-01-01.
- The 8,000-input/2,000-output workload costs $13.50 per 1,000 tasks now and $27.00 after the reversion. One million tasks save $13,500 during the lower-price period.
- Treat $27.00 as the planning baseline. The lower row lasts 139 days from our capture date, and an architecture that only works at the temporary rate has a 2027 problem.
Verdict table
Service tier Through 2026-12-31, input/output per MTok From 2027-01-01, input/output per MTok 1,000 tasks now → 2027 Verdict
------------ ----------------------------------------- -------------------------------------- ---------------------- ------------------------------------------------
Standard $0.75 / $3.75 $1.50 / $7.50 $13.50 → $27.00 Use the lower rate; forecast the reversion
Batch $0.375 / $1.875 $0.75 / $3.75 $6.75 → $13.50 Cheapest listed tier when delay is acceptable
Flex $0.375 / $1.875 $0.75 / $3.75 $6.75 → $13.50 Same token price as Batch, with Flex constraints
Priority $1.35 / $6.75 $2.70 / $13.50 $24.30 → $48.60 Pay for priority when latency has business valueThe task shape is an editable assumption, not observed Vynaris traffic. Each task uses 8,000 input and 2,000 output tokens with no cache reads or tool fees. The Standard derivation is (8,000 × $0.75 + 2,000 × $3.75) / 1,000,000 × 1,000 = $13.50. Replace the rates with $1.50 and $7.50 to get $27.00.
Google's live Gemini API pricing page prints both date ranges on every row. This is not a comparison against a different model or a reseller discount. It is the same model ID and the same service tier at two dates.
The cut is shape-independent
Output-heavy jobs often react differently from input-heavy jobs because providers change one side of the price. That does not happen here. Every token rate in the published 3.6 Flash table is exactly half its 2027 value.
Bill component Through Dec 31 From Jan 1 Current share of future
------------------- --------------- --------------- -----------------------
Standard input $0.75/MTok $1.50/MTok 50%
Standard output $3.75/MTok $7.50/MTok 50%
Standard cache read $0.075/MTok $0.15/MTok 50%
Batch/Flex input $0.375/MTok $0.75/MTok 50%
Batch/Flex output $1.875/MTok $3.75/MTok 50%
Priority input $1.35/MTok $2.70/MTok 50%
Priority output $6.75/MTok $13.50/MTok 50%
Cache storage $0.50/MTok-hour $1.00/MTok-hour 50%That makes the per-request cost answer unusually clean: any fixed workload costs half as much before the deadline. Token shape still decides the absolute bill, but it does not change the 50% reduction.
At one million Standard tasks, the shown workload costs $13,500 now instead of $27,000 after the reversion. The saving is $27,000 - $13,500 = $13,500. This is the moment to replace our assumptions with your own token counts in the calculator, because a 20,000-input agent and a 2,000-input classifier will have different dollar bases even though both receive the same percentage cut.
Why Batch and Flex are not interchangeable
The token rows match, but the operating contracts do not. Google's Flex inference documentation positions Flex for lower-priority work that can tolerate variable service. Batch inference is asynchronous work submitted as a batch. Equal price does not make their latency, queueing or failure behavior equal.
Priority is the opposite trade. Google's Priority inference documentation describes a differentiated service path for workloads that value more predictable performance. On our task, Priority costs $24.30 per 1,000 through December, which is 3.6 times the $6.75 Batch/Flex line. The calculation is $24.30 / $6.75 = 3.6.
Pick the tier from the service requirement first. Then use the price row attached to that tier. Moving an interactive request to Batch just to quote the cheapest number is cost accounting by fiction.
The 139-day planning trap
The interval from 2026-08-15 to 2027-01-01 is 139 days. A temporary price belongs in a forecast as a dated rate card, not as the permanent unit cost.
For a simple budget model, split the forecast at midnight on 2027-01-01:
temporary-period cost = tasks_before_2027 × temporary cost per task
2027 cost = tasks_from_2027 × future cost per task
total forecast = temporary-period cost + 2027 costDo the same for context caching. Cache reads and storage also double. A cache-heavy workload does not escape the reversion; it only starts from a lower absolute input line.
Our July Gemini 3.6 Flash refresh analysis captured the price page that existed then. This article records a later page state with explicit promotional dates. The LLM price-multiplier guide remains useful for regional, priority and cache modifiers, but its stored rows should not be treated as today's source.
Honest tradeoff
Do not redesign a production path around a rate that expires. The temporary price can justify moving already-portable traffic or clearing a backlog. It should not justify a dependency whose margin disappears when the Standard bill doubles.
Keep the 2027 rate in unit-economics tests. If the feature still works at $27.00 per 1,000 representative tasks, the lower price is upside. If it only works at $13.50, you have a deadline disguised as a margin.
FAQ
How much does Gemini 3.6 Flash cost through 2026-12-31?
Standard costs $0.75 per million input tokens and $3.75 per million output tokens. Batch and Flex cost $0.375/$1.875. Priority costs $1.35/$6.75.
What changes on 2027-01-01?
Every published 3.6 Flash token, cache-read and cache-storage row doubles. Standard becomes $1.50/$7.50, while Batch and Flex become $0.75/$3.75.
How do you get $13.50 per 1,000 tasks?
The assumption is 8,000 input and 2,000 output tokens per task. At Standard rates, (8,000 × $0.75 + 2,000 × $3.75) / 1,000,000 × 1,000 = $13.50.
Should we switch to Batch because it is cheaper?
Only if the work is asynchronous. Batch costs $6.75 per 1,000 assumed tasks through December, but an interactive path cannot spend a lower token bill if its latency contract fails.
Sources
- Google Gemini API pricing, live capture and model ID verified 2026-08-15.
- Google Flex inference documentation, verified 2026-08-15.
- Google Priority inference documentation, verified 2026-08-15.
- Reproducible arithmetic:
artifacts/gemini-36-flash-50-percent-temporary-price-cut-december-2026-math.py.