VynarisEarly betaGet your API key

Gemini 3.8 Flash costs 2.1% more per DeepSWE pass than 3.7 Flash

Gemini 3.8 Flash scored 74% at $2.36 per DeepSWE task, yet cost per pass rose 2.1% versus 3.7 Flash. The break-even is 75.6%.

Gemini 3.8 Flash reached 74% on DeepSWE at $2.36 per task. Gemini 3.7 Flash reached 65% at $2.03. Yet cost per benchmark pass rose from $3.12 to $3.19, up 2.1%, because output grew 52%. At this price, 3.8 needs a 75.6% pass rate to win on unit cost.

Prices verified 2026-09-05 from Google's live Gemini API pricing. DeepSWE v1.1 results were updated 2026-09-03 and cover 113 tasks.

TL;DR

Verdict table

Measure                    Gemini 3.8 Flash [high]  Gemini 3.7 Flash [medium]  Change
-------------------------  -----------------------  -------------------------  ---------
DeepSWE pass rate          74%±1%                   65%±3%                     +9 points
Average task cost          $2.36                    $2.03                      +16.3%
Cost per benchmark pass    $3.19                    $3.12                      +2.1%
Output tokens per task     143k                     94k                        +52.1%
Agent steps per task       166                      117                        +41.9%
Spend per 1,000 tasks      $2,360                   $2,030                     +$330
Expected passes per 1,000  740                      650                        +90

The price winner is 3.7 by a narrow $0.07 per benchmark pass. The quality winner is 3.8 by nine score points. If an extra accepted coding task is worth more than $3.67, the higher total spend can still be rational.

What we measured and why

Datacurve's DeepSWE leaderboard measures long-horizon software work across 113 original tasks. Those tasks span 91 repositories and five languages. Every model runs through the same mini-swe-agent harness. The page reports pass rate, average dollar cost, output tokens, and tool-loop steps.

That public task denominator makes this comparison more useful than a rate-card comparison. Both Gemini versions have the same token prices. A fixed 8,000-input, 2,000-output call costs $13.50 per 1,000 requests on either model. DeepSWE shows that the newer model did not hold the trace fixed.

There is one important confound. Datacurve tested 3.8 at high thinking effort and 3.7 at medium effort. Google's model pages confirm those named effort levels exist, but the leaderboard does not provide this pair at the same setting. This is a comparison of the listed configurations, not a clean model-only experiment.

Assumptions and public inputs

Input                         Value                           Source or derivation
----------------------------  ------------------------------  -----------------------------
Benchmark version             DeepSWE v1.1                    113 tasks; updated 2026-09-03
Gemini 3.8 configuration      high                            Datacurve leaderboard
Gemini 3.8 pass rate          74%±1%                          Datacurve leaderboard
Gemini 3.8 average task cost  $2.36                           Datacurve leaderboard
Gemini 3.8 output / steps     143k / 166                      Datacurve leaderboard
Gemini 3.7 configuration      medium                          Datacurve leaderboard
Gemini 3.7 pass rate          65%±3%                          Datacurve leaderboard
Gemini 3.7 average task cost  $2.03                           Datacurve leaderboard
Gemini 3.7 output / steps     94k / 117                       Datacurve leaderboard
Cost per benchmark pass       task cost ÷ pass-rate fraction  shown below
Batch size                    1,000 tasks                     editable scaling assumption

The reported 143k and 94k figures are average output per task, not total benchmark volume. Google bills output at a rate that includes reasoning tokens. We do not know the corresponding average input volume from the leaderboard, so we do not pretend to reconstruct the $2.36 or $2.03 bills from output alone.

Reproducing the $3.19 and $3.12 result

The arithmetic uses expected benchmark passes as the denominator:

Gemini 3.8 Flash cost per pass
= $2.36 / 0.74
= $3.1892, rounded to $3.19

Gemini 3.7 Flash cost per pass
= $2.03 / 0.65
= $3.1231, rounded to $3.12

relative change
= ($3.1892 / $3.1231 - 1) × 100
= 2.1%

This metric is an expectation, not an invoice line. You pay for failed trials too. Over 1,000 tasks, the expected totals make that explicit:

Gemini 3.8 Flash: 1,000 × $2.36 = $2,360; 1,000 × 74% = 740 passes
Gemini 3.7 Flash: 1,000 × $2.03 = $2,030; 1,000 × 65% = 650 passes

marginal cost for 90 additional expected passes
= ($2,360 - $2,030) / (740 - 650)
= $3.67 per additional pass

Use the LLM cost calculator to replace the rate-card token shape with your own input tokens and output. For agent workloads, also record total spend and accepted tasks outside the calculator. That second denominator catches retries and longer traces.

Gemini 3.8 Flash relative changes versus Gemini 3.7 Flash on DeepSWE v1.1
Gemini 3.8 Flash pass rate rose 13.8%, while average task cost rose 16.3%, output tokens 52.1%, and agent steps 41.9%. Source: Datacurve DeepSWE v1.1, captured 2026-09-05. Configurations differ: 3.8 high versus 3.7 medium.

The trace grew much faster than the score. Pass rate increased 13.8% in relative terms. Output increased 52.1%, and steps increased 41.9%. More work bought more passes, but not a lower price per pass.

The break-even is 75.6%, not 74%

At $2.36 per task, 3.8 must clear a 75.6% pass rate to match 3.7's $3.1231 cost per pass:

required pass rate
= $2.36 / $3.1231
= 75.6%

That is 1.6 points above the reported 74%. The other route is a lower task bill. At a 74% pass rate, 3.8 must cost no more than $2.31 per task. That requires a 2.1% cut from the observed $2.36.

This threshold is small enough to test. A shorter system prompt, fewer unnecessary tool results, or a lower effort setting could close a 2.1% gap. None is guaranteed. Measure the same tasks before changing the route.

Our GPT-6 Astra analysis showed the opposite pattern. Astra's token-efficiency gain nearly erased a 2.5x sticker premium. Here, two models share the same sticker, but the 3.8 configuration works longer and loses narrowly on quality-adjusted cost. Unit economics care about the trace, not the logo.

Same rate card, different realized bill

Google lists identical promotional prices for Gemini 3.8 Flash and 3.7 Flash through 2026-12-31:

Service tier  Input per MTok  Output per MTok  8k/2k call  Per 1,000 calls
------------  --------------  ---------------  ----------  ---------------
Standard      $0.75           $3.75            $0.01350    $13.50
Batch         $0.375          $1.875           $0.00675    $6.75
Flex          $0.375          $1.875           $0.00675    $6.75
Priority      $1.35           $6.75            $0.02430    $24.30

Standard prices double to $1.50 input and $7.50 output on 2027-01-01. Batch and Flex also double. Priority becomes $2.70/$13.50. The relative model comparison remains unchanged only if their token shapes remain unchanged.

Batch inference halves the token rate, but it does not halve tokens. A 52% output increase still appears in the bill. The advertised 1,048,576-token context window is also identical across both model pages.

Our earlier Gemini 3.7 fixed-shape analysis showed why equal prices can improve normalized cost when token volume is held constant. DeepSWE now supplies the missing task-level receipt: the volume did not stay constant in these two configurations.

Confidence intervals erase the victory lap

The leaderboard reports 74%±1% for 3.8 and 65%±3% for 3.7. Holding task costs fixed, the interval endpoints produce these ranges:

Configuration              Best endpoint      Central estimate  Worst endpoint
-------------------------  -----------------  ----------------  --------------
Gemini 3.8 Flash [high]    $3.15/pass at 75%  $3.19 at 74%      $3.23 at 73%
Gemini 3.7 Flash [medium]  $2.99/pass at 68%  $3.12 at 65%      $3.27 at 62%

The ranges overlap. The central estimate says 3.7 is 2.1% cheaper per pass. It does not prove that 3.7 has a durable cost advantage. A 113-task benchmark is strong test evidence and weak procurement certainty.

Run a paired evaluation on your repositories. Fix the harness, prompt, tools, retry policy, and effort setting. Measure accepted-task rate, total token cost, latency, and human review time. That produces the denominator your finance team can use.

What this means for routing

Route to 3.8 when the additional accepted result is worth more than its marginal $3.67 benchmark cost. That case is plausible for bug fixes, migrations, and security work. One accepted patch can be worth far more than four dollars.

Keep 3.7 when the workload has a firm quality floor that both models clear. It spends $330 less per 1,000 DeepSWE-shaped tasks. It also uses fewer steps and output tokens in the listed configuration.

Do not automate model routing from this table alone. First compare both at the same effort setting. Then route on your accepted-task cost, not DeepSWE's central estimate.

Honest tradeoff

Gemini 3.8 Flash may be the better economic choice even while its cost per benchmark pass is 2.1% higher. A nine-point pass-rate gain can reduce engineer waiting, reruns, and review work. Those costs are absent from the API bill.

The reverse is also true. High effort may explain part of both the quality gain and the 52% output growth. Without a same-effort comparison, we cannot assign the difference cleanly to model capability. If both models already clear your acceptance floor, 3.7 is the safer cost choice.

FAQ

Is Gemini 3.8 Flash more expensive than Gemini 3.7 Flash?

Not per token. Both cost $0.75/MTok input and $3.75/MTok output through 2026-12-31. On DeepSWE, the tested 3.8 configuration averaged $2.36 per task versus $2.03 for 3.7.

Why did cost per pass rise when the score improved?

Average task cost rose 16.3%, faster than pass rate's 13.8% relative increase. The 3.8 run also used 52.1% more output tokens and 41.9% more steps.

What is the break-even for Gemini 3.8 Flash?

At $2.36 per task, it needs a 75.6% pass rate to match 3.7's $3.12 per pass. At the reported 74%, average task cost must fall to $2.31.

Should we switch to Gemini 3.8 Flash?

Switch only if a paired test shows the extra accepted tasks are worth more than the extra spend. The DeepSWE central estimate prices each incremental pass at $3.67 across 1,000 tasks.

Sources

All arithmetic is reproducible in artifacts/gemini-38-flash-deepswe-cost-per-pass-vs-37-flash-math.py.