VynarisEarly betaGet your API key

GPT-5.6 Sol API price cut: $100 to $72 per 1,000 tasks

GPT-5.6 Sol now costs $72 per 1,000 8k/2k tasks, down 28%. See cache, long-context and service-tier math. Verified 2026-08-22.

GPT-5.6 Sol now costs $72 per 1,000 tasks that use 8,000 input and 2,000 output tokens. The same fixed workload cost $100 yesterday. That is a 28% cut, but OpenAI calls the new $4/$20 rate promotional. Prices verified 2026-08-22.

TL;DR

Verdict table

Fixed workload                          Old cost / 1,000  New cost / 1,000  Reduction  Verdict
--------------------------------------  ----------------  ----------------  ---------  --------------------------------------
Balanced: 8k input, 2k output           $100.00           $72.00            28.0%      Re-run direct and routed margins
Input-heavy: 20k input, 1k output       $130.00           $100.00           23.1%      Input cut drives most savings
Output-heavy: 2k input, 8k output       $250.00           $168.00           32.8%      Output cut matters most
Fully cached: 8k cached, 2k output      $64.00            $43.20            32.5%      Cache hits keep the output-heavy shape
First cache write: 8k write, 2k output  $110.00           $80.00            27.3%      Write premium still applies
Long context: 300k input, 2k output     $3,090.00         $2,460.00         20.4%      Full request uses long-context rates

These workloads are editable assumptions, not observed traffic. The balanced row is (8,000 × $4 + 2,000 × $20) / 1,000,000 × 1,000 = $72.

GPT-5.6 Sol workload savings range from 20.4% to 32.8%
Source: OpenAI API pricing; derived fixed workloads. Prices verified 2026-08-22.

One price cut, six different savings rates

OpenAI reduced the two Standard rates by different amounts. Input is 20% cheaper. Output is 33% cheaper. Your per-request cost lands between them according to token shape.

The input-heavy row spends $80 + $20 = $100 per 1,000 attempts. Its old bill was $100 + $30 = $130. The decrease is $30 / $130 = 23.1%.

The output-heavy row moves from $10 + $240 = $250 to $8 + $160 = $168. Its $82 saving is 32.8%. This is close to the output-rate cut because output dominates the bill.

Use the cost calculator with your own input and output counts before changing a route. A single blended percentage hides the component that pays most of the invoice.

Caching got cheaper on both sides

Prompt caching reads fell from $0.50 to $0.40 per million tokens. Cache writes fell from $6.25 to $5. OpenAI prices a token as uncached input, cached input, or a cache write. The write price is not added to the normal input rate.

For 8,000 cached tokens plus 2,000 output tokens, the new bill is $3.20 + $40 = $43.20 per 1,000 tasks. The prior bill was $4 + $60 = $64.

The first write pass now costs $40 + $40 = $80 per 1,000. It was $50 + $60 = $110. Cache eligibility, prefix stability, and retention still decide whether these rows appear in production.

Our prompt-caching production guide covers that operational side. A lower cache rate cannot rescue a workload that misses its reusable prefix.

Batch, Flex and Fast moved too

The new short-context service-tier prices preserve a clean 0.5x/1x/2x ladder for this workload.

Tier      Input / output per MTok  8k/2k cost per 1,000  Operating constraint
--------  -----------------------  --------------------  -----------------------------------------------------------------------------------------------------------
Batch     $2 / $10                 $36                   Asynchronous [batch inference](https://vynaris.com/glossary/category/inference-and-serving#batch-inference)
Flex      $2 / $10                 $36                   Lower price with higher latency
Standard  $4 / $20                 $72                   Default online path
Fast      $8 / $40                 $144                  Premium latency path

Equal token prices do not make Batch and Flex interchangeable. Their request contracts differ. Pick the service behavior first, then price the eligible tier.

The long-context cliff also moved

OpenAI applies long-context pricing to the full request above 272,000 input tokens. The new multiplier is 2x for input and 1.5x for output. That makes the rates $8/$30, versus the old $10/$45.

A 300k-input/2k-output request now costs $2.40 + $0.06 = $2.46, or $2,460 per 1,000. The previous rate card produced $3.00 + $0.09 = $3.09, or $3,090 per 1,000. The reduction is only 20.4% because the input side dominates.

This is the same billing cliff audited in our GPT-5.6 long-context cost analysis. The threshold did not disappear. The rates on both sides changed.

The routing crossover changed overnight

On the balanced workload, Claude Opus 5 costs $40 input + $50 output = $90 per 1,000 attempts at Anthropic's $5/$25 rates.

Old Sol cost $100, or 11.1% more than Opus 5. New Sol costs $72, or 20% less. That does not prove equal task quality. It proves that any router using the old sticker has a stale break-even.

GPT-5.6 Cyber remains $12.50/$75. Its balanced bill is $100 + $150 = $250 per 1,000 tasks. Cyber was exactly 2.5x the old Sol bill. It is now 3.47x.

The earlier Cyber versus Sol security-task comparison still supplies the quality and escalation framework. Replace its Sol rate before using the decision boundary.

Honest tradeoff

Do not route to Sol because $72 is lower than $90. Model quality, retries, latency, and failed-task cost can reverse an $18 spread. A 20% retry rate turns $72 into $72 × 1.2 = $86.40 before failure penalties.

The other caveat is time. OpenAI says the promotional rate is available at least through 2026-11-21. That is a floor on availability, not a permanent-price promise. Keep the old $5/$30 row as a conservative forecast scenario until OpenAI publishes what follows.

Sources