Blog · 2026-08-04 · Vynaris Team
GPT-5.6's long-context cliff: what crossing 272k costs per agent turn vs Gemini's 200k re-rate
OpenAI long-context tier starts at 272k. Terra and Gemini share $2/$12→$4/$18 stickers but different cliffs. Growing-session math, prices verified 2026-08-04.
Cross 272,000 tokens on GPT-5.6 and the whole request re-prices: input doubles and output rises 1.5x. A Luna turn at 270k costs $0.0546; at 274k it costs $0.1105 (2.02x) for the same 500-token reply. Gemini 3.1 Pro does the same jump at 200k. Anthropic stays flat. Prices verified 2026-08-04.
TL;DR
- OpenAI's pricing page now labels the short tier as under 272k context length. GPT-5.6 Sol / Terra / Luna all carry matching long-context columns at 2.0x input and 1.5x output.
- Terra and Gemini 3.1 Pro share the same stickers ($2/$12 → $4/$18). The difference is the line: Gemini fires at 200k, OpenAI at 272k. A session that grows past 200k but stays under 272k pays Gemini's cliff and Terra's short rate.
- Over a 27-turn agent session that grows from 20k to 332k tokens, Gemini costs $15.68, Terra $13.31, Sonnet 5 (flat) $9.64, Luna $1.33. Luna stays cheapest even after its own cliff.
What we computed, and why
In July we published the 200k-token cliff: Gemini whole-request re-rate versus Anthropic flat, with per-turn agent-loop math. OpenAI's short/long columns were already visible then, but the page did not print where "long" began, so we left GPT-5.6 out of the session totals rather than invent a threshold.
That gap is closed. The live OpenAI pricing markdown labels sibling flagship rows as (<272K context length) on the short column, and GPT-5.6 exposes the same Short / Long columns. We treat 272,000 tokens as the OpenAI line, stamp the capture date, and run the growing-session math OpenAI was missing from the July post. Post-cut Luna/Terra rates are used throughout.
This is not an "Updated" banner on the Gemini article. The load-bearing new pieces are the confirmed OpenAI threshold, the Terra-equals-Gemini sticker identity at different cliffs, and the crossover tables below.
The pricing rules (verified 2026-08-04)
All prices USD per 1M tokens. Output held at 500 tokens per turn to isolate the input cliff.
Model Short / ≤ cliff Long / > cliff Cliff
--------------------------------------------------------------------- --------------- --------------- -----------------
[gpt-5.6-sol](https://vynaris.com/models#gpt-5-6-sol) $5.00 / $30.00 $10.00 / $45.00 ≥272k
[gpt-5.6-terra](https://vynaris.com/models#gpt-5-6-terra) $2.00 / $12.00 $4.00 / $18.00 ≥272k
[gpt-5.6-luna](https://vynaris.com/models#gpt-5-6-luna) $0.20 / $1.20 $0.40 / $1.80 ≥272k
[Gemini 3.1 Pro](https://vynaris.com/models#gemini-3-1-pro) Preview $2.00 / $12.00 $4.00 / $18.00 >200k
[Claude Sonnet 5](https://vynaris.com/models#claude-sonnet-5) (intro) $2.00 / $10.00 same none (flat to 1M)
[Claude Opus 5](https://vynaris.com/models#claude-opus-5) $5.00 / $25.00 same none (flat to 1M)Read the Terra and Gemini rows together. Identical stickers. Different thresholds. That is the structural finding of this post: for any prompt between 200,001 and 271,999 tokens, Gemini has already re-rated and Terra has not.
The word that matters on both OpenAI and Google is total. The higher rate is not a marginal surcharge on tokens above the line. Once the request crosses, every token in it bills at the higher rate. That is why a 1–2% size increase can roughly double the turn.
The assumptions (edit these)
Assumption Value Note
------------------------------------------------------------------------ ------------- ---------------------------------------
Session start context 20,000 tokens system prompt + first file loads
Context growth per turn 12,000 tokens tool results, re-read code, prior turns
Turns in the session 27 (0..26) ends at 332,000 tokens re-sent
[Output tokens](https://vynaris.com/glossary/output-token-cost) per turn 500 held constant
Gemini cliff turn 16 prompt reaches 212,000
OpenAI cliff turn 21 prompt reaches 272,000The input-token ratio is what makes this bite. An AI agent that re-sends a growing context window every turn is almost entirely an input workload. A surcharge on input lands on nearly the whole bill.
Results: two cliffs, 72k tokens apart
Per-turn cost at fixed prompt sizes, before caching:
Prompt size Gemini 3.1 Pro Claude Sonnet 5 gpt-5.6-luna gpt-5.6-terra gpt-5.6-sol
----------- -------------- --------------- ------------ ------------- -----------
199,000 $0.4040 $0.4030 $0.0404 $0.4040 $1.0100
201,000 $0.8130 $0.4070 $0.0408 $0.4080 $1.0200
270,000 $1.0890 $0.5450 $0.0546 $0.5460 $1.3650
274,000 $1.1050 $0.5530 $0.1105 $1.1050 $2.7625
320,000 $1.2890 $0.6450 $0.1289 $1.2890 $3.2225Gemini's jump is the 199k→201k row: $0.404 → $0.813 (2.01x). OpenAI's jump is the 270k→274k row: Luna $0.0546 → $0.1105, Terra $0.546 → $1.105, Sol $1.365 → $2.763 (all 2.02x). Between those two lines, from 201k to 270k, Gemini is already on the expensive tier and Terra is still on the cheap one. At 270k that gap is $1.089 versus $0.546, almost exactly 2x, for the same sticker family.

Results: the growing session
A fixed prompt size is a snapshot. The workload people actually run grows. Below is the 27-turn session, showing both cliffs.
Turn Prompt Gemini 3.1 Pro Sonnet 5 Luna Terra Sol
------------- ------- -------------- --------- --------- ---------- ----------
0 20,000 $0.0460 $0.0450 $0.0046 $0.0460 $0.1150
15 200,000 $0.4060 $0.4050 $0.0406 $0.4060 $1.0150
16 212,000 $0.8570 $0.4290 $0.0430 $0.4300 $1.0750
20 260,000 $1.0490 $0.5250 $0.0526 $0.5260 $1.3150
21 272,000 $1.0970 $0.5490 $0.1097 $1.0970 $2.7425
26 332,000 $1.3370 $0.6690 $0.1337 $1.3370 $3.3425
Session total *$15.6830* *$9.6390* *$1.3308* *$13.3080* *$33.2700*Turn 16 is Gemini's cliff: context grew 6%, per-turn cost grew 2.11x. Turn 21 is OpenAI's cliff: Terra and Luna roughly double in one step while Sonnet 5 tracks the token count (+4.6%). From turn 16 through turn 20, Terra undercuts Gemini by about half on every turn, despite matching stickers, because only Gemini has crossed.
Session ratios versus Sonnet 5: Gemini 1.63x, Terra 1.38x, Sol 3.45x, Luna 0.14x. Luna stays cheapest after its cliff because its long-tier stickers ($0.40/$1.80) are still far below Sonnet's flat $2/$10.
If you are sizing a real agent budget, plug your own growth curve into the calculator and mark where the session crosses 200k and 272k before you pick a provider.
At 1,000 turns held at a 320k context, the cliff tax alone (long rate minus what the short rate would have charged for the same request) is $643 on Gemini or Terra, $64 on Luna, and $1,608 on Sol. That money buys nothing. Same request, same output, priced by which side of a line the prompt fell on.
What it means for routing
The context window is a routing variable with two provider-specific tripwires now, not one.
- Know which line you will hit. If your re-sent context lives between 200k and 272k, Gemini has already doubled and OpenAI has not. Prefer Terra or a flat Anthropic model over Gemini 3.1 Pro in that band, all else equal.
- Trim before you cross, not after. Summarizing or dropping stale context to stay under the active cliff is worth ~2x on a tiered model and close to nothing on a flat one. The value of right-sizing depends on which model you are on.
- Luna changes the "just pay flat" instinct. After the 2026-07-30 cut, Luna's long-tier turn at 274k ($0.1105) is still 5x cheaper than Sonnet 5 at the same size ($0.5530). Flat is not automatically cheaper. Flat removes cliff risk; Luna's post-cut long tier can still win on dollars.
- [Prompt caching](https://vynaris.com/glossary/prompt-caching) does not dodge the cliff. Cached tokens still count toward the total that decides your tier. A large cached prefix can push you over the line on its own. Caching cuts the re-read bill; it does not move the threshold.
An OpenAI-compatible gateway makes the test cheap: route requests whose context has grown past 200k away from Gemini 3.1 Pro, keep sub-272k OpenAI requests on the short tier, and send anything that must stay huge to a flat-rate model or to Luna if quality holds. That is model routing with two numeric triggers instead of one.
Where this model is wrong
- Thresholds can move. We stamp 2026-08-04. OpenAI's short-column label is
(<272K context length)on sibling models; if the GPT-5.6 boundary shifts, the session totals shift with it. Re-verify before you budget. - Output was held constant. Reasoning-heavy turns generate far more output, and both OpenAI and Gemini also raise the output rate at the cliff (1.5x). Our numbers are the conservative case.
- Flat is not the same as cheap. Opus 5 is flat and expensive at every prompt size. Sonnet 5 is flat and mid. Luna is tiered and still under both after the cliff. Match the rule to the model to the task.
- Batch and caching change absolute levels. We modeled synchronous, uncached calls to keep the cliffs visible. Your effective rate can be lower on every row; the ratio between tiered and flat is the durable finding.
- Quality is out of scope. This post prices the cliff. It does not claim Luna at 300k matches Opus 5 at 300k on hard reasoning. Route on quality first, then on which side of 200k / 272k you sit.
When does the tiered model still win? When your context genuinely stays small. Under 200k, Gemini 3.1 Pro and Terra cost nearly the same as each other and sit next to Sonnet 5 on input. For a short-prompt workload that never approaches either cliff, the tiered model can be the right call. The cliff only bites workloads that grow into it, which is most agents and few chatbots.
FAQ
Where does OpenAI's long-context tier start?
At 272,000 tokens of context length, per the short-column label (<272K context length) on OpenAI's live pricing page (verified 2026-08-04). GPT-5.6 Sol, Terra, and Luna all expose Short / Long price columns at that boundary.
How is that different from Gemini's 200k cliff?
Same billing shape (whole request re-rates), different line. Gemini 3.1 Pro Preview jumps at >200k. OpenAI jumps at ≥272k. Terra and Gemini even share $2/$12 → $4/$18 stickers, so the 200k–272k band is where Gemini has already doubled and Terra has not.
Does Luna's cliff erase its price advantage?
No, on these numbers. Luna at 274k is $0.1105 per turn versus Sonnet 5 at $0.5530. The cliff doubles Luna; it does not put Luna above the flat mid-tier.
Should I always pick Anthropic for long context?
Only if you need flat rates and Sonnet/Opus quality. If Luna quality holds for your task, its long-tier stickers still undercut Sonnet 5 after 272k. Flat removes surprise; it does not automatically minimize the bill.
Sources
- OpenAI API pricing (GPT-5.6 Short/Long columns;
<272K context lengthshort-tier label), verified live 2026-08-04: https://developers.openai.com/api/docs/pricing - Google Gemini API pricing (Gemini 3.1 Pro Preview ≤200k / >200k), verified live 2026-08-04: https://ai.google.dev/gemini-api/docs/pricing
- Anthropic pricing (Sonnet 5 intro $2/$10 flat; Opus 5 $5/$25 flat), verified live 2026-08-04: https://platform.claude.com/docs/en/about-claude/pricing
- Prior Gemini-cliff analysis (OpenAI threshold then unknown): https://vynaris.com/blog/context-window-long-context-surcharge-costs
- All arithmetic:
gpt-56-long-context-cliff-vs-gemini-200k-cost-math.py
Prices change. We re-verify every figure monthly. Numbers are current as of 2026-08-04.
Related: 200k-token cliff (Gemini vs Anthropic) · August 2026 LLM price list · coding agent cost per task