VynarisEarly beta Kimi K3Get your API key

GPT-5.6's long-context cliff: what crossing 272k costs per agent turn vs Gemini's 200k re-rate

OpenAI long-context tier starts at 272k. Terra and Gemini share $2/$12→$4/$18 stickers but different cliffs. Growing-session math, prices verified 2026-08-04.

Cross 272,000 tokens on GPT-5.6 and the whole request re-prices: input doubles and output rises 1.5x. A Luna turn at 270k costs $0.0546; at 274k it costs $0.1105 (2.02x) for the same 500-token reply. Gemini 3.1 Pro does the same jump at 200k. Anthropic stays flat. Prices verified 2026-08-04.

TL;DR

What we computed, and why

In July we published the 200k-token cliff: Gemini whole-request re-rate versus Anthropic flat, with per-turn agent-loop math. OpenAI's short/long columns were already visible then, but the page did not print where "long" began, so we left GPT-5.6 out of the session totals rather than invent a threshold.

That gap is closed. The live OpenAI pricing markdown labels sibling flagship rows as (<272K context length) on the short column, and GPT-5.6 exposes the same Short / Long columns. We treat 272,000 tokens as the OpenAI line, stamp the capture date, and run the growing-session math OpenAI was missing from the July post. Post-cut Luna/Terra rates are used throughout.

This is not an "Updated" banner on the Gemini article. The load-bearing new pieces are the confirmed OpenAI threshold, the Terra-equals-Gemini sticker identity at different cliffs, and the crossover tables below.

The pricing rules (verified 2026-08-04)

All prices USD per 1M tokens. Output held at 500 tokens per turn to isolate the input cliff.

Model                                                                  Short / ≤ cliff  Long / > cliff   Cliff
---------------------------------------------------------------------  ---------------  ---------------  -----------------
[gpt-5.6-sol](https://vynaris.com/models#gpt-5-6-sol)                  $5.00 / $30.00   $10.00 / $45.00  ≥272k
[gpt-5.6-terra](https://vynaris.com/models#gpt-5-6-terra)              $2.00 / $12.00   $4.00 / $18.00   ≥272k
[gpt-5.6-luna](https://vynaris.com/models#gpt-5-6-luna)                $0.20 / $1.20    $0.40 / $1.80    ≥272k
[Gemini 3.1 Pro](https://vynaris.com/models#gemini-3-1-pro) Preview    $2.00 / $12.00   $4.00 / $18.00   >200k
[Claude Sonnet 5](https://vynaris.com/models#claude-sonnet-5) (intro)  $2.00 / $10.00   same             none (flat to 1M)
[Claude Opus 5](https://vynaris.com/models#claude-opus-5)              $5.00 / $25.00   same             none (flat to 1M)

Read the Terra and Gemini rows together. Identical stickers. Different thresholds. That is the structural finding of this post: for any prompt between 200,001 and 271,999 tokens, Gemini has already re-rated and Terra has not.

The word that matters on both OpenAI and Google is total. The higher rate is not a marginal surcharge on tokens above the line. Once the request crosses, every token in it bills at the higher rate. That is why a 1–2% size increase can roughly double the turn.

The assumptions (edit these)

Assumption                                                                Value          Note
------------------------------------------------------------------------  -------------  ---------------------------------------
Session start context                                                     20,000 tokens  system prompt + first file loads
Context growth per turn                                                   12,000 tokens  tool results, re-read code, prior turns
Turns in the session                                                      27 (0..26)     ends at 332,000 tokens re-sent
[Output tokens](https://vynaris.com/glossary/output-token-cost) per turn  500            held constant
Gemini cliff turn                                                         16             prompt reaches 212,000
OpenAI cliff turn                                                         21             prompt reaches 272,000

The input-token ratio is what makes this bite. An AI agent that re-sends a growing context window every turn is almost entirely an input workload. A surcharge on input lands on nearly the whole bill.

Results: two cliffs, 72k tokens apart

Per-turn cost at fixed prompt sizes, before caching:

Prompt size  Gemini 3.1 Pro  Claude Sonnet 5  gpt-5.6-luna  gpt-5.6-terra  gpt-5.6-sol
-----------  --------------  ---------------  ------------  -------------  -----------
199,000      $0.4040         $0.4030          $0.0404       $0.4040        $1.0100
201,000      $0.8130         $0.4070          $0.0408       $0.4080        $1.0200
270,000      $1.0890         $0.5450          $0.0546       $0.5460        $1.3650
274,000      $1.1050         $0.5530          $0.1105       $1.1050        $2.7625
320,000      $1.2890         $0.6450          $0.1289       $1.2890        $3.2225

Gemini's jump is the 199k→201k row: $0.404 → $0.813 (2.01x). OpenAI's jump is the 270k→274k row: Luna $0.0546 → $0.1105, Terra $0.546 → $1.105, Sol $1.365 → $2.763 (all 2.02x). Between those two lines, from 201k to 270k, Gemini is already on the expensive tier and Terra is still on the cheap one. At 270k that gap is $1.089 versus $0.546, almost exactly 2x, for the same sticker family.

Per-turn cost vs prompt size: Gemini steps at 200k, Terra and Luna step at 272k, Sonnet 5 rises linearly with no cliff
Per-turn cost vs prompt size, output fixed at 500 tokens. Gemini cliff at 200k; OpenAI GPT-5.6 cliff at 272k; Sonnet 5 flat. Prices verified 2026-08-04.

Results: the growing session

A fixed prompt size is a snapshot. The workload people actually run grows. Below is the 27-turn session, showing both cliffs.

Turn           Prompt   Gemini 3.1 Pro  Sonnet 5   Luna       Terra       Sol
-------------  -------  --------------  ---------  ---------  ----------  ----------
0              20,000   $0.0460         $0.0450    $0.0046    $0.0460     $0.1150
15             200,000  $0.4060         $0.4050    $0.0406    $0.4060     $1.0150
16             212,000  $0.8570         $0.4290    $0.0430    $0.4300     $1.0750
20             260,000  $1.0490         $0.5250    $0.0526    $0.5260     $1.3150
21             272,000  $1.0970         $0.5490    $0.1097    $1.0970     $2.7425
26             332,000  $1.3370         $0.6690    $0.1337    $1.3370     $3.3425
Session total           *$15.6830*      *$9.6390*  *$1.3308*  *$13.3080*  *$33.2700*

Turn 16 is Gemini's cliff: context grew 6%, per-turn cost grew 2.11x. Turn 21 is OpenAI's cliff: Terra and Luna roughly double in one step while Sonnet 5 tracks the token count (+4.6%). From turn 16 through turn 20, Terra undercuts Gemini by about half on every turn, despite matching stickers, because only Gemini has crossed.

Session ratios versus Sonnet 5: Gemini 1.63x, Terra 1.38x, Sol 3.45x, Luna 0.14x. Luna stays cheapest after its cliff because its long-tier stickers ($0.40/$1.80) are still far below Sonnet's flat $2/$10.

If you are sizing a real agent budget, plug your own growth curve into the calculator and mark where the session crosses 200k and 272k before you pick a provider.

At 1,000 turns held at a 320k context, the cliff tax alone (long rate minus what the short rate would have charged for the same request) is $643 on Gemini or Terra, $64 on Luna, and $1,608 on Sol. That money buys nothing. Same request, same output, priced by which side of a line the prompt fell on.

What it means for routing

The context window is a routing variable with two provider-specific tripwires now, not one.

An OpenAI-compatible gateway makes the test cheap: route requests whose context has grown past 200k away from Gemini 3.1 Pro, keep sub-272k OpenAI requests on the short tier, and send anything that must stay huge to a flat-rate model or to Luna if quality holds. That is model routing with two numeric triggers instead of one.

Where this model is wrong

When does the tiered model still win? When your context genuinely stays small. Under 200k, Gemini 3.1 Pro and Terra cost nearly the same as each other and sit next to Sonnet 5 on input. For a short-prompt workload that never approaches either cliff, the tiered model can be the right call. The cliff only bites workloads that grow into it, which is most agents and few chatbots.

FAQ

Where does OpenAI's long-context tier start?

At 272,000 tokens of context length, per the short-column label (<272K context length) on OpenAI's live pricing page (verified 2026-08-04). GPT-5.6 Sol, Terra, and Luna all expose Short / Long price columns at that boundary.

How is that different from Gemini's 200k cliff?

Same billing shape (whole request re-rates), different line. Gemini 3.1 Pro Preview jumps at >200k. OpenAI jumps at ≥272k. Terra and Gemini even share $2/$12 → $4/$18 stickers, so the 200k–272k band is where Gemini has already doubled and Terra has not.

Does Luna's cliff erase its price advantage?

No, on these numbers. Luna at 274k is $0.1105 per turn versus Sonnet 5 at $0.5530. The cliff doubles Luna; it does not put Luna above the flat mid-tier.

Should I always pick Anthropic for long context?

Only if you need flat rates and Sonnet/Opus quality. If Luna quality holds for your task, its long-tier stickers still undercut Sonnet 5 after 272k. Flat removes surprise; it does not automatically minimize the bill.

Sources

Prices change. We re-verify every figure monthly. Numbers are current as of 2026-08-04.

Related: 200k-token cliff (Gemini vs Anthropic) · August 2026 LLM price list · coding agent cost per task