Blog · 2026-08-14 · Vynaris Team
DeepSeek V4 price reset: off-peak costs 77–83% more
DeepSeek V4 Flash rises from $1.68 to $3.08 off-peak or $6.16 peak per 1,000 balanced tasks on August 16. Prices verified 2026-08-14.
DeepSeek-V4-Flash's “off-peak” rate will cost 83% more than today's price on an 8,000-input, 2,000-output task. After 2026-08-16 16:00 UTC, 1,000 tasks cost $3.08 off-peak or $6.16 peak, up from $1.68. DeepSeek-V4-Pro rises 77% off-peak. Prices verified 2026-08-14.
TL;DR
- Flash moves from $1.68 per 1,000 balanced tasks to $3.08 off-peak and $6.16 peak. Pro moves from $5.22 to $9.24 and $18.48.
- Peak hours total seven hours daily: 01:00–04:00 and 06:00–10:00 UTC. The remaining 17 hours use off-peak rates.
- Off-peak is half the new peak rate. It is not a discount against the price customers pay today.
Verdict table
Model Today /1M input-output Off-peak /1M input-output Peak /1M input-output 1,000-task bill: today → off-peak → peak
----------------- ---------------------- ------------------------- --------------------- ----------------------------------------
deepseek-v4-flash $0.14 / $0.28 $0.22 / $0.66 $0.44 / $1.32 $1.68 → $3.08 → $6.16
deepseek-v4-pro $0.435 / $0.87 $0.66 / $1.98 $1.32 / $3.96 $5.22 → $9.24 → $18.48The bill uses 8,000 input tokens and 2,000 output tokens per task, with no cache hits. For Flash today, the shown derivation is (8,000 × $0.14 + 2,000 × $0.28) / 1,000,000 × 1,000 = $1.68. Replace the two rates with $0.22 and $0.66 to get $3.08 off-peak.
DeepSeek's live pricing page supplies every rate and the exact cutover. Its 2026-08-13 release note verifies the current model release. The pricing discussion reached 122 points and 46 comments when we checked the Hacker News thread.
The lower new rate is still a price increase
The label is the trap. “Off-peak” compares with the new peak row, not with today's flat row.
For Flash, (3.08 / 1.68) - 1 = 83.3%. For Pro, (9.24 / 5.22) - 1 = 77.0%. Peak raises the same workload 266.7% on Flash and 254.0% on Pro. This is a change in per-request cost, not a cosmetic reshuffle of the table.
The output side drives much of the jump. Flash output rises from $0.28 per million to $0.66 off-peak, a 135.7% increase. Pro output goes from $0.87 to $1.98, up 127.6%. Input rises less: 57.1% on Flash and 51.7% on Pro.
That matters for agents that produce long plans, code or structured records. A growing context window pushes the bill closer to the input-rate change too.
What a scheduler can actually save
Peak covers seven of 24 UTC hours. If requests arrive uniformly, the new Flash blend is (7 × $6.16 + 17 × $3.08) / 24 = $3.98 per 1,000 tasks. Pro lands at $11.94. Those blended bills are 136.8% and 128.6% above today's rates.
Moving every deferrable peak request into the 17 off-peak hours cuts 22.6% from that uniform-arrival bill. It does not recover today's price. A scheduler turns $3.98 into $3.08 on Flash, not $1.68.
Batch queues, offline extraction and nightly evaluations are the obvious candidates. Keep timestamps in UTC and treat the 2026-08-16 16:00 UTC cutover as a configuration change. If you want to test your own token shape, use the LLM cost calculator once, then encode the result in the job's budget.
Interactive chat and synchronous tool calls are different. Delaying them converts a token saving into latency. A seven-hour peak window can easily cost more in abandoned work or missed service levels than the API delta.
Cache hits get a sharper reset
Prompt caching still helps, but the hit rate itself rises faster than the cache-miss rate. Flash cache-hit input moves from $0.0028 per million to $0.007 off-peak and $0.014 peak. Those are 2.5× and 5× today's rate. Pro moves from $0.003625 to $0.022 and $0.044, or 6.1× and 12.1×.
On a task with 6,000 cached input tokens, 2,000 uncached input tokens and 2,000 output tokens, Flash moves from $0.8568 per 1,000 tasks to $1.802 off-peak or $3.604 peak. Pro moves from $2.6318 to $5.412 or $10.824. The lower cost per token on a cache hit remains valuable. It simply no longer preserves the old unit cost.
Our earlier DeepSeek V4 production-cost analysis explains why concurrency and latency can outweigh sticker price. The prompt-cache write-churn model shows when changing a prefix destroys the saving before this price reset even enters the equation.
Honest tradeoff
Do not add time-based routing to a latency-sensitive path just to chase the lower row. You add queueing, clock logic and another failure mode. For steady interactive traffic, accept the new blended rate or route by measured quality and model routing. Schedule only work that already tolerates delay.
FAQ
Is DeepSeek V4 off-peak cheaper than today's API price?
No. On the 8,000-input, 2,000-output shape, Flash is 83.3% more expensive off-peak and Pro is 77.0% more expensive.
When do DeepSeek's new prices start?
At 16:00 UTC on 2026-08-16. Peak hours are 01:00–04:00 and 06:00–10:00 UTC each day.
Does prompt caching avoid the increase?
No. It still reduces the input bill, but cache-hit rates also rise. Pro's new off-peak cache-hit rate is about 6.1× today's rate.
Vynaris option
Vynaris is an OpenAI-compatible inference gateway. Use it when a workload needs quality-aware provider choice, not merely a clock. One base URL swap keeps the scheduling and routing decision outside application code.
Sources
- DeepSeek API pricing, verified 2026-08-14.
- DeepSeek 2026-08-13 release note, verified 2026-08-14.
- Hacker News pricing discussion, counts verified 2026-08-14.