Blog · 2026-07-31 · Vynaris Team
GPT-5.6 Luna vs Haiku 4.5 vs Gemini 3.5 Flash-Lite vs DeepSeek v4-flash: the budget-tier cost floor
After Luna's 80% cut, the budget tier has two floors: DeepSeek v4-flash on raw sticker ($0.14/$0.28), GPT-5.6 Luna among frontier labs ($0.20/$1.20). Cost per 1,000 tasks vs Haiku 4.5 and Gemini 3.5 Flash-Lite across three shapes. Verified 2026-07-31.
For a balanced classification task (1,500 input, 150 output tokens), GPT-5.6 Luna costs $0.48 per 1,000 tasks against $0.83 for Gemini 3.5 Flash-Lite, $2.25 for Claude Haiku 4.5, and $0.25 for DeepSeek v4-flash. After OpenAI cut Luna 80% on 2026-07-30, the budget tier has two floors: DeepSeek on raw sticker, Luna among frontier labs. Prices verified 2026-07-31.
This is the durable version of that cut. The same-day re-route forensics live here; this page is the verdict you keep, because the ranking below holds until one of these four providers moves a price again.
Verdict table
Model $/1M in $/1M out $/1k tasks (class.) Pick it when
--------------------- ------- -------- ------------------- --------------------------------------------------------------------------
DeepSeek v4-flash $0.14 $0.28 $0.25 Raw sticker rules; output-heavy batch; its quality and data-governance fit
GPT-5.6 Luna $0.20 $1.20 $0.48 You want a frontier lab near DeepSeek's price with OpenAI-native tooling
Gemini 3.5 Flash-Lite $0.30 $2.50 $0.83 You are already on Vertex/Gemini and want multimodal input
Claude Haiku 4.5 $1.00 $5.00 $2.25 Anthropic tooling, extended thinking, or 5m/1h cache economics carry itClassification shape = 1,500 input + 150 output tokens per task, first-party rates, no caching. On blended sticker (1M in + 1M out) the order is DeepSeek $0.42, Luna $1.40, Flash-Lite $2.80, Haiku $6.00.
What changed, and why Luna reset the tier
Before 2026-07-30, Luna listed at $1/$6, the same input sticker as Haiku 4.5. It was a mid-pack budget option. The 80% cut to $0.20/$1.20 dropped it below Flash-Lite and to within striking distance of DeepSeek. Luna's new $0.20 input equals OpenAI's own gpt-5.4-nano sticker ($0.20/$1.25) while being a far more capable model, which is the part that makes this a genuine tier reset rather than a rounding change.
Against the field, Luna now runs 4.7x cheaper than Haiku 4.5 and 1.7x cheaper than Flash-Lite on the classification shape. The only budget model still under it is DeepSeek v4-flash.
The cross-shape picture: where the gaps move
The ranking is stable across task shapes, but the size of each gap is not. We modeled three representative shapes so you can find yours. All figures are USD per 1,000 tasks, prices verified 2026-07-31.
Shape (in / out) DeepSeek v4-flash GPT-5.6 Luna Gemini 3.5 Flash-Lite Claude Haiku 4.5
----------------------------- ----------------- ------------ --------------------- ----------------
Extraction, 3,000 / 300 $0.50 $0.96 $1.65 $4.50
Classification, 1,500 / 150 $0.25 $0.48 $0.83 $2.25
Short generation, 400 / 1,200 $0.39 $1.52 $3.12 $6.40
The lever is output tokens. Luna's output price ($1.20/1M) is 4.3x DeepSeek's ($0.28/1M), while its input price ($0.20) is only 1.4x DeepSeek's ($0.14). So the two models are closest on input-heavy work and furthest apart on output-heavy work. On extraction (10% output share) Luna is 1.9x DeepSeek; on short generation (75% output share) Luna is 3.9x. If your workload leans output-heavy, DeepSeek's lead widens; if it leans input-heavy, Luna nearly catches it while giving you a frontier lab.
Haiku and Flash-Lite hold their positions above Luna on every shape. Their case is not price. It is right-sizing around a specific need: Anthropic's extended thinking and cache multipliers, or Gemini's multimodal input tokens and Vertex integration.
When the expensive sticker wins anyway
Cost per 1,000 tasks is verifiable to the cent. Cost per correct task is not, and that is where a cheaper model can lose. Route on cost-per-token only when quality is a tie; otherwise the sticker floor is a trap.
- Schema-fragile extraction. If a malformed field breaks a downstream job, a model that fails validation 2% more often erases a 1.9x price gap in retries. Test schema-validity rate, not just the cost per token. This is the same lesson the extraction-workhorse comparison drew before Luna's cut, and it still holds: DeepSeek won on price there too, and quality-per-task decided the exceptions.
- Ambiguity that needs reasoning. When a task requires inferring an unstated value rather than copying a present one, Haiku 4.5's extended thinking or Luna's frontier training can justify the higher rate. Pure copy-out work does not.
- Data governance and stack fit. DeepSeek is the cheapest sticker, but its residency and provider profile do not clear every compliance bar. Luna gives a US frontier lab at 1.9x to 3.9x DeepSeek's price, which is often the cheaper decision once the governance cost is priced in.
- Latency-bound serving. None of these stickers include the latency SLA. If Flash-Lite or Haiku hits your p95 target and a cheaper model does not, the sticker gap is moot.
The honest rule: default the easy, high-volume bulk to the sticker floor, and reserve a stronger model for the hard, high-stakes minority. That is model routing, and it beats picking one model for everything. Drop your own in/out split into the calculator to see where your break-even sits.
Caching does not reorder this
Prompt caching helps most when a large prefix is stable across calls. These budget shapes are the opposite: small prompts, fresh input each call, little to cache. A cache hit trims each model roughly in proportion to its input share, so it lowers every bar without changing the order. Luna's cache read fell to $0.02/1M in the same cut, and DeepSeek's cache hit ($0.0028/1M) is lower still. Caching is a multiplier on the model you picked, not a way to rescue an expensive one into contention.
Bottom line
- Absolute cheapest bill: DeepSeek v4-flash, and its lead grows with output share.
- Cheapest frontier lab: GPT-5.6 Luna, now 4.7x under Haiku 4.5 and 1.7x under Flash-Lite, closest to DeepSeek on input-heavy work.
- Haiku 4.5 and Flash-Lite: no longer price-competitive at this tier; pick them for a capability, latency, or ecosystem reason, not the sticker.
For the same-day cut forensics and the who-re-routes table, see the News analysis. For the wider rate sheet across 14 models, the July 2026 LLM price list.
FAQ
What is the cheapest budget LLM after Luna's cut? On raw sticker, DeepSeek v4-flash at $0.14/$0.28 per million tokens, verified 2026-07-31. Among frontier labs, GPT-5.6 Luna at $0.20/$1.20 is cheapest, sitting 1.9x to 3.9x above DeepSeek depending on output share.
Is GPT-5.6 Luna cheaper than Claude Haiku 4.5? Yes, by a wide margin now. On the classification shape (1,500 in / 150 out) Luna is $0.48 per 1,000 tasks against Haiku's $2.25, a 4.7x gap. Before the 80% cut they shared the same $1 input sticker.
When should I still pay for Haiku 4.5 or Gemini 3.5 Flash-Lite? When a capability decides the outcome: Anthropic's extended thinking and cache economics, or Gemini's multimodal input and Vertex integration. On price alone at this tier, both are now dominated by Luna and DeepSeek.
Does DeepSeek always beat Luna? On sticker, yes, across every shape here. The gap is smallest on input-heavy extraction (1.9x) and largest on output-heavy generation (3.9x). Luna wins the decision when frontier-lab quality, US residency, or OpenAI-native tooling is worth the premium.
Sources
- OpenAI API pricing (GPT-5.6 Luna $0.20/$1.20, gpt-5.4-nano $0.20/$1.25, cache read $0.02), verified live 2026-07-31: https://developers.openai.com/api/docs/pricing
- Anthropic pricing (Claude Haiku 4.5 $1/$5, cache read $0.10, extended thinking), verified live 2026-07-31: https://platform.claude.com/docs/en/docs/about-claude/pricing
- Google Gemini API pricing (Gemini 3.5 Flash-Lite $0.30/$2.50), verified live 2026-07-31: https://ai.google.dev/gemini-api/docs/pricing
- DeepSeek API pricing (v4-flash $0.14/$0.28, cache hit $0.0028), verified live 2026-07-31: https://api-docs.deepseek.com/quick_start/pricing
- All arithmetic:
gpt-56-luna-vs-haiku-45-vs-flash-lite-vs-deepseek-budget-tier-floor-math.py, assumptions editable inline.
Prices change. We re-verify every figure in this post monthly and stamp updates. Numbers are current as of 2026-07-31.
Related: Luna 80% cut re-route math · extraction workhorse comparison · July 2026 LLM price list