Blog · 2026-09-05 · Vynaris Team
Gemini 3.8 Flash costs 2.1% more per DeepSWE pass than 3.7 Flash
Gemini 3.8 Flash scored 74% at $2.36 per DeepSWE task, yet cost per pass rose 2.1% versus 3.7 Flash. The break-even is 75.6%.
Gemini 3.8 Flash reached 74% on DeepSWE at $2.36 per task. Gemini 3.7 Flash reached 65% at $2.03. Yet cost per benchmark pass rose from $3.12 to $3.19, up 2.1%, because output grew 52%. At this price, 3.8 needs a 75.6% pass rate to win on unit cost.
Prices verified 2026-09-05 from Google's live Gemini API pricing. DeepSWE v1.1 results were updated 2026-09-03 and cover 113 tasks.
TL;DR
- Gemini 3.8 Flash scored nine points higher, 74% versus 65%. Its average cost per task also rose 16.3%, from $2.03 to $2.36.
- Dividing task cost by pass rate gives $3.19 per benchmark pass for 3.8 and $3.12 for 3.7. The stronger score did not make a successful result cheaper.
- Across 1,000 benchmark tasks, 3.8 buys 90 additional expected passes for $330 more. That is $3.67 per additional pass, before assigning business value.
Verdict table
Measure Gemini 3.8 Flash [high] Gemini 3.7 Flash [medium] Change
------------------------- ----------------------- ------------------------- ---------
DeepSWE pass rate 74%±1% 65%±3% +9 points
Average task cost $2.36 $2.03 +16.3%
Cost per benchmark pass $3.19 $3.12 +2.1%
Output tokens per task 143k 94k +52.1%
Agent steps per task 166 117 +41.9%
Spend per 1,000 tasks $2,360 $2,030 +$330
Expected passes per 1,000 740 650 +90The price winner is 3.7 by a narrow $0.07 per benchmark pass. The quality winner is 3.8 by nine score points. If an extra accepted coding task is worth more than $3.67, the higher total spend can still be rational.
What we measured and why
Datacurve's DeepSWE leaderboard measures long-horizon software work across 113 original tasks. Those tasks span 91 repositories and five languages. Every model runs through the same mini-swe-agent harness. The page reports pass rate, average dollar cost, output tokens, and tool-loop steps.
That public task denominator makes this comparison more useful than a rate-card comparison. Both Gemini versions have the same token prices. A fixed 8,000-input, 2,000-output call costs $13.50 per 1,000 requests on either model. DeepSWE shows that the newer model did not hold the trace fixed.
There is one important confound. Datacurve tested 3.8 at high thinking effort and 3.7 at medium effort. Google's model pages confirm those named effort levels exist, but the leaderboard does not provide this pair at the same setting. This is a comparison of the listed configurations, not a clean model-only experiment.
Assumptions and public inputs
Input Value Source or derivation
---------------------------- ------------------------------ -----------------------------
Benchmark version DeepSWE v1.1 113 tasks; updated 2026-09-03
Gemini 3.8 configuration high Datacurve leaderboard
Gemini 3.8 pass rate 74%±1% Datacurve leaderboard
Gemini 3.8 average task cost $2.36 Datacurve leaderboard
Gemini 3.8 output / steps 143k / 166 Datacurve leaderboard
Gemini 3.7 configuration medium Datacurve leaderboard
Gemini 3.7 pass rate 65%±3% Datacurve leaderboard
Gemini 3.7 average task cost $2.03 Datacurve leaderboard
Gemini 3.7 output / steps 94k / 117 Datacurve leaderboard
Cost per benchmark pass task cost ÷ pass-rate fraction shown below
Batch size 1,000 tasks editable scaling assumptionThe reported 143k and 94k figures are average output per task, not total benchmark volume. Google bills output at a rate that includes reasoning tokens. We do not know the corresponding average input volume from the leaderboard, so we do not pretend to reconstruct the $2.36 or $2.03 bills from output alone.
Reproducing the $3.19 and $3.12 result
The arithmetic uses expected benchmark passes as the denominator:
Gemini 3.8 Flash cost per pass
= $2.36 / 0.74
= $3.1892, rounded to $3.19
Gemini 3.7 Flash cost per pass
= $2.03 / 0.65
= $3.1231, rounded to $3.12
relative change
= ($3.1892 / $3.1231 - 1) × 100
= 2.1%This metric is an expectation, not an invoice line. You pay for failed trials too. Over 1,000 tasks, the expected totals make that explicit:
Gemini 3.8 Flash: 1,000 × $2.36 = $2,360; 1,000 × 74% = 740 passes
Gemini 3.7 Flash: 1,000 × $2.03 = $2,030; 1,000 × 65% = 650 passes
marginal cost for 90 additional expected passes
= ($2,360 - $2,030) / (740 - 650)
= $3.67 per additional passUse the LLM cost calculator to replace the rate-card token shape with your own input tokens and output. For agent workloads, also record total spend and accepted tasks outside the calculator. That second denominator catches retries and longer traces.

The trace grew much faster than the score. Pass rate increased 13.8% in relative terms. Output increased 52.1%, and steps increased 41.9%. More work bought more passes, but not a lower price per pass.
The break-even is 75.6%, not 74%
At $2.36 per task, 3.8 must clear a 75.6% pass rate to match 3.7's $3.1231 cost per pass:
required pass rate
= $2.36 / $3.1231
= 75.6%That is 1.6 points above the reported 74%. The other route is a lower task bill. At a 74% pass rate, 3.8 must cost no more than $2.31 per task. That requires a 2.1% cut from the observed $2.36.
This threshold is small enough to test. A shorter system prompt, fewer unnecessary tool results, or a lower effort setting could close a 2.1% gap. None is guaranteed. Measure the same tasks before changing the route.
Our GPT-6 Astra analysis showed the opposite pattern. Astra's token-efficiency gain nearly erased a 2.5x sticker premium. Here, two models share the same sticker, but the 3.8 configuration works longer and loses narrowly on quality-adjusted cost. Unit economics care about the trace, not the logo.
Same rate card, different realized bill
Google lists identical promotional prices for Gemini 3.8 Flash and 3.7 Flash through 2026-12-31:
Service tier Input per MTok Output per MTok 8k/2k call Per 1,000 calls
------------ -------------- --------------- ---------- ---------------
Standard $0.75 $3.75 $0.01350 $13.50
Batch $0.375 $1.875 $0.00675 $6.75
Flex $0.375 $1.875 $0.00675 $6.75
Priority $1.35 $6.75 $0.02430 $24.30Standard prices double to $1.50 input and $7.50 output on 2027-01-01. Batch and Flex also double. Priority becomes $2.70/$13.50. The relative model comparison remains unchanged only if their token shapes remain unchanged.
Batch inference halves the token rate, but it does not halve tokens. A 52% output increase still appears in the bill. The advertised 1,048,576-token context window is also identical across both model pages.
Our earlier Gemini 3.7 fixed-shape analysis showed why equal prices can improve normalized cost when token volume is held constant. DeepSWE now supplies the missing task-level receipt: the volume did not stay constant in these two configurations.
Confidence intervals erase the victory lap
The leaderboard reports 74%±1% for 3.8 and 65%±3% for 3.7. Holding task costs fixed, the interval endpoints produce these ranges:
Configuration Best endpoint Central estimate Worst endpoint
------------------------- ----------------- ---------------- --------------
Gemini 3.8 Flash [high] $3.15/pass at 75% $3.19 at 74% $3.23 at 73%
Gemini 3.7 Flash [medium] $2.99/pass at 68% $3.12 at 65% $3.27 at 62%The ranges overlap. The central estimate says 3.7 is 2.1% cheaper per pass. It does not prove that 3.7 has a durable cost advantage. A 113-task benchmark is strong test evidence and weak procurement certainty.
Run a paired evaluation on your repositories. Fix the harness, prompt, tools, retry policy, and effort setting. Measure accepted-task rate, total token cost, latency, and human review time. That produces the denominator your finance team can use.
What this means for routing
Route to 3.8 when the additional accepted result is worth more than its marginal $3.67 benchmark cost. That case is plausible for bug fixes, migrations, and security work. One accepted patch can be worth far more than four dollars.
Keep 3.7 when the workload has a firm quality floor that both models clear. It spends $330 less per 1,000 DeepSWE-shaped tasks. It also uses fewer steps and output tokens in the listed configuration.
Do not automate model routing from this table alone. First compare both at the same effort setting. Then route on your accepted-task cost, not DeepSWE's central estimate.
Honest tradeoff
Gemini 3.8 Flash may be the better economic choice even while its cost per benchmark pass is 2.1% higher. A nine-point pass-rate gain can reduce engineer waiting, reruns, and review work. Those costs are absent from the API bill.
The reverse is also true. High effort may explain part of both the quality gain and the 52% output growth. Without a same-effort comparison, we cannot assign the difference cleanly to model capability. If both models already clear your acceptance floor, 3.7 is the safer cost choice.
FAQ
Is Gemini 3.8 Flash more expensive than Gemini 3.7 Flash?
Not per token. Both cost $0.75/MTok input and $3.75/MTok output through 2026-12-31. On DeepSWE, the tested 3.8 configuration averaged $2.36 per task versus $2.03 for 3.7.
Why did cost per pass rise when the score improved?
Average task cost rose 16.3%, faster than pass rate's 13.8% relative increase. The 3.8 run also used 52.1% more output tokens and 41.9% more steps.
What is the break-even for Gemini 3.8 Flash?
At $2.36 per task, it needs a 75.6% pass rate to match 3.7's $3.12 per pass. At the reported 74%, average task cost must fall to $2.31.
Should we switch to Gemini 3.8 Flash?
Switch only if a paired test shows the extra accepted tasks are worth more than the extra spend. The DeepSWE central estimate prices each incremental pass at $3.67 across 1,000 tasks.
Sources
- Datacurve DeepSWE v1.1 leaderboard, 113 tasks, updated 2026-09-03 and captured 2026-09-05.
- DeepSWE changelog, Gemini 3.8 Flash added 2026-09-01; captured 2026-09-05.
- Google Gemini API pricing, rate card verified 2026-09-05.
- Google Gemini 3.8 Flash model page, stable model ID and effort levels verified 2026-09-05.
- Google Gemini 3.7 Flash model page, stable model ID and limits verified 2026-09-05.
All arithmetic is reproducible in artifacts/gemini-38-flash-deepswe-cost-per-pass-vs-37-flash-math.py.