VynarisEarly beta Kimi K3Get your API key

GPT-5.6 Luna vs Haiku 4.5 vs Gemini 3.5 Flash-Lite vs DeepSeek v4-flash: the budget-tier cost floor

After Luna's 80% cut, the budget tier has two floors: DeepSeek v4-flash on raw sticker ($0.14/$0.28), GPT-5.6 Luna among frontier labs ($0.20/$1.20). Cost per 1,000 tasks vs Haiku 4.5 and Gemini 3.5 Flash-Lite across three shapes. Verified 2026-07-31.

For a balanced classification task (1,500 input, 150 output tokens), GPT-5.6 Luna costs $0.48 per 1,000 tasks against $0.83 for Gemini 3.5 Flash-Lite, $2.25 for Claude Haiku 4.5, and $0.25 for DeepSeek v4-flash. After OpenAI cut Luna 80% on 2026-07-30, the budget tier has two floors: DeepSeek on raw sticker, Luna among frontier labs. Prices verified 2026-07-31.

This is the durable version of that cut. The same-day re-route forensics live here; this page is the verdict you keep, because the ranking below holds until one of these four providers moves a price again.

Verdict table

Model                  $/1M in  $/1M out  $/1k tasks (class.)  Pick it when
---------------------  -------  --------  -------------------  --------------------------------------------------------------------------
DeepSeek v4-flash      $0.14    $0.28     $0.25                Raw sticker rules; output-heavy batch; its quality and data-governance fit
GPT-5.6 Luna           $0.20    $1.20     $0.48                You want a frontier lab near DeepSeek's price with OpenAI-native tooling
Gemini 3.5 Flash-Lite  $0.30    $2.50     $0.83                You are already on Vertex/Gemini and want multimodal input
Claude Haiku 4.5       $1.00    $5.00     $2.25                Anthropic tooling, extended thinking, or 5m/1h cache economics carry it

Classification shape = 1,500 input + 150 output tokens per task, first-party rates, no caching. On blended sticker (1M in + 1M out) the order is DeepSeek $0.42, Luna $1.40, Flash-Lite $2.80, Haiku $6.00.

What changed, and why Luna reset the tier

Before 2026-07-30, Luna listed at $1/$6, the same input sticker as Haiku 4.5. It was a mid-pack budget option. The 80% cut to $0.20/$1.20 dropped it below Flash-Lite and to within striking distance of DeepSeek. Luna's new $0.20 input equals OpenAI's own gpt-5.4-nano sticker ($0.20/$1.25) while being a far more capable model, which is the part that makes this a genuine tier reset rather than a rounding change.

Against the field, Luna now runs 4.7x cheaper than Haiku 4.5 and 1.7x cheaper than Flash-Lite on the classification shape. The only budget model still under it is DeepSeek v4-flash.

The cross-shape picture: where the gaps move

The ranking is stable across task shapes, but the size of each gap is not. We modeled three representative shapes so you can find yours. All figures are USD per 1,000 tasks, prices verified 2026-07-31.

Shape (in / out)               DeepSeek v4-flash  GPT-5.6 Luna  Gemini 3.5 Flash-Lite  Claude Haiku 4.5
-----------------------------  -----------------  ------------  ---------------------  ----------------
Extraction, 3,000 / 300        $0.50              $0.96         $1.65                  $4.50
Classification, 1,500 / 150    $0.25              $0.48         $0.83                  $2.25
Short generation, 400 / 1,200  $0.39              $1.52         $3.12                  $6.40
Grouped bars of cost per 1,000 tasks for four budget models across extraction, classification and short-generation shapes; DeepSeek lowest everywhere, Luna second, then Flash-Lite, Haiku highest
Cost per 1,000 tasks by shape, log scale. DeepSeek is the sticker floor; Luna the frontier-lab floor. Prices verified 2026-07-31.

The lever is output tokens. Luna's output price ($1.20/1M) is 4.3x DeepSeek's ($0.28/1M), while its input price ($0.20) is only 1.4x DeepSeek's ($0.14). So the two models are closest on input-heavy work and furthest apart on output-heavy work. On extraction (10% output share) Luna is 1.9x DeepSeek; on short generation (75% output share) Luna is 3.9x. If your workload leans output-heavy, DeepSeek's lead widens; if it leans input-heavy, Luna nearly catches it while giving you a frontier lab.

Haiku and Flash-Lite hold their positions above Luna on every shape. Their case is not price. It is right-sizing around a specific need: Anthropic's extended thinking and cache multipliers, or Gemini's multimodal input tokens and Vertex integration.

When the expensive sticker wins anyway

Cost per 1,000 tasks is verifiable to the cent. Cost per correct task is not, and that is where a cheaper model can lose. Route on cost-per-token only when quality is a tie; otherwise the sticker floor is a trap.

The honest rule: default the easy, high-volume bulk to the sticker floor, and reserve a stronger model for the hard, high-stakes minority. That is model routing, and it beats picking one model for everything. Drop your own in/out split into the calculator to see where your break-even sits.

Caching does not reorder this

Prompt caching helps most when a large prefix is stable across calls. These budget shapes are the opposite: small prompts, fresh input each call, little to cache. A cache hit trims each model roughly in proportion to its input share, so it lowers every bar without changing the order. Luna's cache read fell to $0.02/1M in the same cut, and DeepSeek's cache hit ($0.0028/1M) is lower still. Caching is a multiplier on the model you picked, not a way to rescue an expensive one into contention.

Bottom line

For the same-day cut forensics and the who-re-routes table, see the News analysis. For the wider rate sheet across 14 models, the July 2026 LLM price list.

FAQ

What is the cheapest budget LLM after Luna's cut? On raw sticker, DeepSeek v4-flash at $0.14/$0.28 per million tokens, verified 2026-07-31. Among frontier labs, GPT-5.6 Luna at $0.20/$1.20 is cheapest, sitting 1.9x to 3.9x above DeepSeek depending on output share.

Is GPT-5.6 Luna cheaper than Claude Haiku 4.5? Yes, by a wide margin now. On the classification shape (1,500 in / 150 out) Luna is $0.48 per 1,000 tasks against Haiku's $2.25, a 4.7x gap. Before the 80% cut they shared the same $1 input sticker.

When should I still pay for Haiku 4.5 or Gemini 3.5 Flash-Lite? When a capability decides the outcome: Anthropic's extended thinking and cache economics, or Gemini's multimodal input and Vertex integration. On price alone at this tier, both are now dominated by Luna and DeepSeek.

Does DeepSeek always beat Luna? On sticker, yes, across every shape here. The gap is smallest on input-heavy extraction (1.9x) and largest on output-heavy generation (3.9x). Luna wins the decision when frontier-lab quality, US residency, or OpenAI-native tooling is worth the premium.

Sources

Prices change. We re-verify every figure in this post monthly and stamp updates. Numbers are current as of 2026-07-31.

Related: Luna 80% cut re-route math · extraction workhorse comparison · July 2026 LLM price list