VynarisEarly beta Kimi K3Get your API key

LLM API Pricing August 2026: cost per task across 14 models — and what changed since July

GPT-5.6 Luna -80% and Terra -20% held through August. Cost per task across 14 models on a 2k/300 workhorse: 96x spread, prices verified 2026-08-04.

GPT-5.6 Luna fell 80% to $0.20/$1.20 and Terra fell 20% to $2/$12 on 2026-07-30; both held through August. On a 2,000-in / 300-out workhorse the 14-model spread is now 96x: Claude Fable 5 at $35.00 per 1,000 tasks versus deepseek-v4-flash at $0.364. Prices verified 2026-08-04.

This is the monthly series refresh of our July 2026 price table. The July→August story is almost entirely one provider's cut. Anthropic, DeepSeek, and Gemini stickers in this table are identical to last month. What moved is OpenAI's mid and budget GPT-5.6 tiers, and that reshaped the floor.

What changed since July

Model          July sticker ($/1M in/out)  August sticker  Workhorse $/1k tasks  Delta
-------------  --------------------------  --------------  --------------------  -----
gpt-5.6-luna   $1.00 / $6.00               $0.20 / $1.20   $3.80 → $0.76         -80%
gpt-5.6-terra  $2.50 / $15.00              $2.00 / $12.00  $9.50 → $7.60         -20%
gpt-5.6-sol    $5.00 / $30.00              $5.00 / $30.00  $19.00 → $19.00       held

Everything else in the 14-model set held: Claude Opus 5 at $5/$25, Sonnet 5 intro at $2/$10 through 2026-08-31, Haiku 4.5 at $1/$5, DeepSeek v4-flash at $0.14/$0.28, Gemini 3.6 Flash at $1.50/$7.50. The same-day Luna/Terra re-route forensics and the budget-tier floor comparison cover the cut in detail. This page is the dated full table you keep open.

One calendar item still pending: Claude Sonnet 5 intro pricing ends 2026-08-31. On 2026-09-01 it steps from $2/$10 to $3/$15, which moves the workhorse from $7.00 to $10.50 per 1,000 tasks (+50%). That is the next scheduled re-rate on this board.

The verdict table

One representative task: 2,000 input tokens plus 300 output tokens. That is a support-ticket triage, a JSON extraction from a document chunk, or one routing step inside an AI agent. The monthly column assumes 100,000 tasks. Math: cost per 1,000 tasks = (2,000 × input $/1M + 300 × output $/1M) / 1,000.

Model                    Input $/1M  Output $/1M  Per 1,000 tasks  Per 100k tasks/mo
-----------------------  ----------  -----------  ---------------  -----------------
Claude Fable 5           $10.00      $50.00       $35.00           $3,500
gpt-5.6-sol              $5.00       $30.00       $19.00           $1,900
gpt-5.5                  $5.00       $30.00       $19.00           $1,900
Claude Opus 5            $5.00       $25.00       $17.50           $1,750
Claude Sonnet 4.6        $3.00       $15.00       $10.50           $1,050
gpt-5.4                  $2.50       $15.00       $9.50            $950
gpt-5.6-terra            $2.00       $12.00       $7.60            $760
Claude Sonnet 5 (intro)  $2.00       $10.00       $7.00            $700
Claude Haiku 4.5         $1.00       $5.00        $3.50            $350
gpt-5.4-mini             $0.75       $4.50        $2.85            $285
deepseek-v4-pro          $0.435      $0.87        $1.131           $113
gpt-5.4-nano             $0.20       $1.25        $0.775           $78
gpt-5.6-luna             $0.20       $1.20        $0.760           $76
deepseek-v4-flash        $0.14       $0.28        $0.364           $36
Cost per 1,000 tasks across 14 models on a 2,000-in / 300-out workhorse, log scale, 96x spread from Fable 5 to DeepSeek v4-flash
Cost per 1,000 tasks (2,000 in / 300 out), log scale. Prices verified 2026-08-04 from first-party provider pages.

Three caveats sit under every row. Claude Sonnet 5 is introductory through 2026-08-31. OpenAI GPT-5.6 and gpt-5.5 carry a long-context tier at or above 272k tokens (Sol doubles to $10/$45). DeepSeek numbers are cache-miss prices; a cache hit on deepseek-v4-flash is $0.0028 per million input tokens.

The second shape: one agentic turn

Workhorse math understates agent bills, because agents re-send a larger prefix. Same 14 models, one multi-tool turn at 8,000 in / 1,500 out:

Model                    Per 1,000 agentic turns
-----------------------  -----------------------
Claude Fable 5           $155.00
gpt-5.6-sol              $85.00
gpt-5.5                  $85.00
Claude Opus 5            $77.50
Claude Sonnet 4.6        $46.50
gpt-5.4                  $42.50
gpt-5.6-terra            $34.00
Claude Sonnet 5 (intro)  $31.00
Claude Haiku 4.5         $15.50
gpt-5.4-mini             $12.75
deepseek-v4-pro          $4.785
gpt-5.4-nano             $3.475
gpt-5.6-luna             $3.400
deepseek-v4-flash        $1.540

Ranking holds. Absolute dollars stretch. Fable to flash is still ~100x. If your production mix is closer to the agentic row than the workhorse row, multiply the workhorse column by about 4.4 and stop pretending the triage table is your invoice.

Three clusters, one floor reset

The table is still three clusters, but the floor moved.

Flagship ($17.50–$35.00 per 1,000 workhorse tasks): Claude Fable 5, gpt-5.6-sol, gpt-5.5, Claude Opus 5. Use these where a wrong plan poisons every downstream step.

Mid ($7.00–$10.50): Sonnet 4.6, gpt-5.4, gpt-5.6-terra, Sonnet 5 intro. Strong general models. Roughly half to two-thirds of flagship.

Floor ($0.364–$3.50): gpt-5.6-luna, Haiku 4.5, gpt-5.4-mini, gpt-5.4-nano, both DeepSeek models. Classification, extraction, formatting, and most agent plumbing live here.

The August change is inside the floor. Luna now sits at $0.76 per 1,000 workhorse tasks, under Haiku ($3.50) by 4.6x and within 2.1x of deepseek-v4-flash ($0.364). Before the cut Luna was $3.80 on this shape, above Haiku. That is why the budget-tier floor post exists as a separate page: the ranking inside the cheap cluster changed, not just the absolute dollars.

When the expensive model still wins

Cost per token is not cost per correct task. Route a flagship when error compounds, when output ships without human review, or when a single miss costs more than the model premium. A code-review agent that misses a security bug did not save $12. It bought an incident.

The mistake is not using Fable 5 or gpt-5.6-sol. The mistake is using them as the default for every step because switching models means touching glue code. Model routing exists to stop that habit.

Don't route this

Honest limits of the cheap cluster, from the same pricing pages: long-context work gets expensive or unsupported at the bottom (gpt-5.4-mini and nano list no long-context tier), and DeepSeek's concurrency caps are lower than the US providers'. If your workload is 100 requests per day, the entire table above is noise. The spread between Fable and flash is about $3.46 per day at that volume, and your engineering time costs more. Cost routing starts mattering somewhere around a few thousand requests per day.

Also: do not treat this page as another Luna forensics write-up. The durable value is the dated full table and the July→August diff. For the cut itself, use the News analysis.

Doing the math on your own traffic

Take yesterday's request log, bucket by what each call actually does, multiply each bucket by the table once at your current model and once at the cheapest cluster that plausibly handles it. That delta, annualized, is the number for your next infra discussion. Run any token mix through the calculator when the shape differs from 2,000 / 300.

This is also what Vynaris does continuously: an OpenAI-compatible gateway that sends each request to the cheapest model that holds quality on that task class, one base URL swap, with per-request costs visible. Get an API key at vynaris.com.

FAQ

What is the cheapest LLM API in August 2026 among these 14?

deepseek-v4-flash at $0.14/$0.28 per million tokens (cache miss), or $0.364 per 1,000 workhorse tasks. Among frontier-lab models, gpt-5.6-luna at $0.20/$1.20 is cheapest at $0.76 per 1,000.

How much did GPT-5.6 Luna drop since July?

80% on both input and output stickers ($1/$6 → $0.20/$1.20), cutting workhorse cost from $3.80 to $0.76 per 1,000 tasks. The cut landed 2026-07-30 and held through this 2026-08-04 verify.

When does Claude Sonnet 5 intro pricing end?

2026-08-31. From 2026-09-01 the sticker is $3/$15, which moves the workhorse from $7.00 to $10.50 per 1,000 tasks (+50%).

Is the ranking the same on agentic workloads?

Yes on order, no on dollars. At 8,000 in / 1,500 out the same 14 models keep their relative order; absolute cost per 1,000 turns is roughly 4.4x the workhorse column.

Sources

Prices change. We re-verify every figure in this series monthly and stamp updates. Numbers are current as of 2026-08-04.

Related: July 2026 LLM price list · Luna/Terra cut re-route math · budget-tier floor after Luna