Blog · 2026-08-22 · Vynaris Team
GPT-5.6 Sol API price cut: $100 to $72 per 1,000 tasks
GPT-5.6 Sol now costs $72 per 1,000 8k/2k tasks, down 28%. See cache, long-context and service-tier math. Verified 2026-08-22.
GPT-5.6 Sol now costs $72 per 1,000 tasks that use 8,000 input and 2,000 output tokens. The same fixed workload cost $100 yesterday. That is a 28% cut, but OpenAI calls the new $4/$20 rate promotional. Prices verified 2026-08-22.
TL;DR
- Standard input tokens fell from $5 to $4 per million. Output tokens fell from $30 to $20 per million.
- An 8k-input/2k-output task fell from $0.100 to $0.072. The reduction is 28%, not either headline percentage.
- The price is promotional through at least 2026-11-21. Budgeting beyond that date needs a second rate scenario.
Verdict table
Fixed workload Old cost / 1,000 New cost / 1,000 Reduction Verdict
-------------------------------------- ---------------- ---------------- --------- --------------------------------------
Balanced: 8k input, 2k output $100.00 $72.00 28.0% Re-run direct and routed margins
Input-heavy: 20k input, 1k output $130.00 $100.00 23.1% Input cut drives most savings
Output-heavy: 2k input, 8k output $250.00 $168.00 32.8% Output cut matters most
Fully cached: 8k cached, 2k output $64.00 $43.20 32.5% Cache hits keep the output-heavy shape
First cache write: 8k write, 2k output $110.00 $80.00 27.3% Write premium still applies
Long context: 300k input, 2k output $3,090.00 $2,460.00 20.4% Full request uses long-context ratesThese workloads are editable assumptions, not observed traffic. The balanced row is (8,000 × $4 + 2,000 × $20) / 1,000,000 × 1,000 = $72.

One price cut, six different savings rates
OpenAI reduced the two Standard rates by different amounts. Input is 20% cheaper. Output is 33% cheaper. Your per-request cost lands between them according to token shape.
The input-heavy row spends $80 + $20 = $100 per 1,000 attempts. Its old bill was $100 + $30 = $130. The decrease is $30 / $130 = 23.1%.
The output-heavy row moves from $10 + $240 = $250 to $8 + $160 = $168. Its $82 saving is 32.8%. This is close to the output-rate cut because output dominates the bill.
Use the cost calculator with your own input and output counts before changing a route. A single blended percentage hides the component that pays most of the invoice.
Caching got cheaper on both sides
Prompt caching reads fell from $0.50 to $0.40 per million tokens. Cache writes fell from $6.25 to $5. OpenAI prices a token as uncached input, cached input, or a cache write. The write price is not added to the normal input rate.
For 8,000 cached tokens plus 2,000 output tokens, the new bill is $3.20 + $40 = $43.20 per 1,000 tasks. The prior bill was $4 + $60 = $64.
The first write pass now costs $40 + $40 = $80 per 1,000. It was $50 + $60 = $110. Cache eligibility, prefix stability, and retention still decide whether these rows appear in production.
Our prompt-caching production guide covers that operational side. A lower cache rate cannot rescue a workload that misses its reusable prefix.
Batch, Flex and Fast moved too
The new short-context service-tier prices preserve a clean 0.5x/1x/2x ladder for this workload.
Tier Input / output per MTok 8k/2k cost per 1,000 Operating constraint
-------- ----------------------- -------------------- -----------------------------------------------------------------------------------------------------------
Batch $2 / $10 $36 Asynchronous [batch inference](https://vynaris.com/glossary/category/inference-and-serving#batch-inference)
Flex $2 / $10 $36 Lower price with higher latency
Standard $4 / $20 $72 Default online path
Fast $8 / $40 $144 Premium latency pathEqual token prices do not make Batch and Flex interchangeable. Their request contracts differ. Pick the service behavior first, then price the eligible tier.
The long-context cliff also moved
OpenAI applies long-context pricing to the full request above 272,000 input tokens. The new multiplier is 2x for input and 1.5x for output. That makes the rates $8/$30, versus the old $10/$45.
A 300k-input/2k-output request now costs $2.40 + $0.06 = $2.46, or $2,460 per 1,000. The previous rate card produced $3.00 + $0.09 = $3.09, or $3,090 per 1,000. The reduction is only 20.4% because the input side dominates.
This is the same billing cliff audited in our GPT-5.6 long-context cost analysis. The threshold did not disappear. The rates on both sides changed.
The routing crossover changed overnight
On the balanced workload, Claude Opus 5 costs $40 input + $50 output = $90 per 1,000 attempts at Anthropic's $5/$25 rates.
Old Sol cost $100, or 11.1% more than Opus 5. New Sol costs $72, or 20% less. That does not prove equal task quality. It proves that any router using the old sticker has a stale break-even.
GPT-5.6 Cyber remains $12.50/$75. Its balanced bill is $100 + $150 = $250 per 1,000 tasks. Cyber was exactly 2.5x the old Sol bill. It is now 3.47x.
The earlier Cyber versus Sol security-task comparison still supplies the quality and escalation framework. Replace its Sol rate before using the decision boundary.
Honest tradeoff
Do not route to Sol because $72 is lower than $90. Model quality, retries, latency, and failed-task cost can reverse an $18 spread. A 20% retry rate turns $72 into $72 × 1.2 = $86.40 before failure penalties.
The other caveat is time. OpenAI says the promotional rate is available at least through 2026-11-21. That is a floor on availability, not a permanent-price promise. Keep the old $5/$30 row as a conservative forecast scenario until OpenAI publishes what follows.
Sources
- OpenAI API pricing, current Standard, Batch, Flex, Fast, cache and long-context rows verified 2026-08-22.
- OpenAI GPT-5.6 Sol model page, reduction percentages, promotional window and 272k re-rate verified 2026-08-22.
- Anthropic API pricing, Claude Opus 5 $5/$25 verified 2026-08-22.
- Previous OpenAI rate card captured 2026-08-21 in
artifacts/gpt-56-sol-price-cut-72-per-1000-balanced-tasks-source-openai-pricing-2026-08-21.html. - Reproducible arithmetic:
artifacts/gpt-56-sol-price-cut-72-per-1000-balanced-tasks-math.py.