DeepSeek V4 API Price Increase (Aug 16, 2026): New Peak/Off-Peak Rates
What happened. DeepSeek's new peak/off-peak API pricing for the V4 model family took effect at 16:00 UTC on August 16, 2026 — the schedule DeepSeek published on its official pricing page and confirmed in its Aug 13, 2026 changelog entry. Off-peak rates are half of peak rates; peak hours are 01:00–04:00 and 06:00–10:00 UTC. The old "significant increase expected" warning is gone — the concrete numbers are now live on api-docs.deepseek.com/quick_start/pricing (verified August 16, 2026).
What the new DeepSeek V4 rates are (per 1M tokens)
| Model | Input, cache hit | Input, cache miss | Output | Window |
|---|---|---|---|---|
| V4-Flash | $0.007 | $0.22 | $0.66 | OFF-PEAK |
| V4-Flash | $0.014 | $0.44 | $1.32 | PEAK |
| V4-Pro | $0.022 | $0.66 | $1.98 | OFF-PEAK |
| V4-Pro | $0.044 | $1.32 | $3.96 | PEAK |
Previous (pre-hike, effective until 16:00 UTC Aug 16): V4-Flash $0.0028 (cache hit) / $0.14 (cache miss) / $0.28 output; V4-Pro $0.003625 / $0.435 / $0.87.
How much did prices actually go up?
- V4-Flash: off-peak ≈ 2.5x (cache hit) / 1.6x (cache miss) / 2.4x (output); peak ≈ 5x / 3.1x / 4.7x.
- V4-Pro: off-peak ≈ 6.1x (cache hit) / 1.5x (cache miss) / 2.3x (output); peak ≈ 12.1x / 3.0x / 4.6x.
- The single largest jump is V4-Pro cache-hit input at peak — roughly 12x the old rate.
- TechStartups frames the overall increase as 50% to 1,100% above the old schedule (Aug 13, 2026); Bloomberg's paywalled report described it as "quadrupling peak prices."
Alongside the pricing change, the Aug 13 changelog entry also marked DeepSeek-V4-Pro GA (APP/Web/API, model name unchanged, native Responses API, three thinking-effort levels: low/high/max).
What the increase costs an agency
DeepSeek was the default "cheapest stack" assumption in agency cost models. That assumption does not survive the new table:
- At peak, V4-Flash output ($1.32/1M) now sits above Gemini 3.7 Flash intro pricing ($0.75/$3.75 per 1M through 2026-12-31), Muse Glimmer hosted on Together AI ($0.35/$1.50), and within reach of Grok 4.6 ($2/$6) and Qwen 3.8 Max ($2/$6).
- At peak, V4-Pro ($1.32 input / $3.96 output) overlaps Claude Sonnet 5 ($2/$10) on input and approaches premium frontier on output — the "cheap reasoning model" positioning is gone for peak-hour work.
- Cache hits and off-peak scheduling are now the levers. Off-peak cache-hit input is $0.007 (Flash) / $0.022 (Pro) — the only rates that stay near the old price level.
- Margin impact: for an agency whose retainer was priced on the old Flash schedule, a peak-hour agentic workload now costs ~3–5x on the compute line. That changes project margins unless the estimate is re-baselined.
What to do now
- Re-run every estimate that assumed DeepSeek pricing with the new table (use this calculator's model strategy selector for the updated comparison).
- Schedule batch and non-urgent work off-peak (outside 01:00–04:00 and 06:00–10:00 UTC) to lock the half-price tier.
- Exploit cache hits. Prompt-cached workloads at off-peak are the cheapest live option in the family.
- Re-benchmark alternatives: Gemini 3.7 Flash intro ($0.75/$3.75, through 2026-12-31), Grok 4.6 ($2/$6, 500K context), Qwen 3.8 Max ($2/$6 open weights), and self-hosted V4 (weights are MIT-licensed — sustained workloads can be served directly via vLLM-class tooling).
Re-run your estimate with the new DeepSeek rates
Open the AI Agency Pricing Calculator →Model DeepSeek V4, Gemini 3.7 Flash, Grok 4.6, Qwen 3.8 Max, Claude Sonnet 5, and local self-host — with current 2026 rates.
Frequently asked questions
When did the DeepSeek V4 API price increase take effect?
The new peak/off-peak rates took effect at 16:00 UTC on August 16, 2026, per DeepSeek's official pricing page (footnote 1) and the Aug 13, 2026 changelog entry. Peak hours are 01:00–04:00 and 06:00–10:00 UTC; all other hours are off-peak at half the peak rate.
How much did DeepSeek V4 API prices increase?
V4-Flash peak pricing is $0.014 (cache hit) / $0.44 (cache miss) / $1.32 output per 1M tokens vs $0.0028/$0.14/$0.28 before — roughly 5x on cache-hit input, 3.1x on cache-miss input, and 4.7x on output. V4-Pro peak is $0.044/$1.32/$3.96 vs $0.003625/$0.435/$0.87 — roughly 12.1x on cache-hit input, 3.0x on cache-miss input, and 4.6x on output. Off-peak rates are half of peak.
Is DeepSeek still the cheapest AI API after the increase?
Only off-peak cache-miss input remains near the old ballpark, and even that is 1.5–1.6x higher. At peak rates, V4-Flash output ($1.32/1M) now sits above Gemini 3.7 Flash intro pricing ($0.75/$3.75), Muse Glimmer on Together AI ($0.35/$1.50), and close to Grok 4.6 ($2/$6) and Qwen 3.8 Max ($2/$6). DeepSeek is no longer the default "cheapest stack" assumption — re-run your routing math with the new table.
What should agencies do about the DeepSeek price increase?
Update cost assumptions immediately: use peak/off-peak rates in estimates, schedule non-urgent batch work off-peak, treat cache hits as the lever (off-peak cache-hit input is $0.007 Flash / $0.022 Pro), and re-benchmark alternatives (Gemini 3.7 Flash intro, Grok 4.6, Qwen 3.8 Max, self-hosted MIT-licensed V4 weights for sustained workloads).
Sources
- DeepSeek official pricing page (live verified Aug 16, 2026 — new peak/off-peak table active): api-docs.deepseek.com/quick_start/pricing
- DeepSeek changelog, "DeepSeek-V4-Pro Update" (Aug 13, 2026 — GA + API Pricing Adjustment): api-docs.deepseek.com/updates
- TechStartups (Aug 13, 2026 — 50% to 1,100% above current): techstartups.com
- Elser AI (Aug 14, 2026 — 16:00 UTC confirmation + full table): elser.ai
- Reuters + Bloomberg (Aug 13, 2026 — paywalled; registered-uncited corroboration)
Accuracy note: All new rates, peak hours, and the effective date (16:00 UTC Aug 16, 2026) are verified against DeepSeek's official pricing page and changelog on Aug 16, 2026. Multiplier math (e.g., 5x, 12.1x) is derived from the published before/after rates, not a DeepSeek statement. The 50%–1,100% range is TechStartups' framing; "quadruples peak prices" is Bloomberg's framing (paywalled). DeepSeek reserves the right to adjust prices — re-verify before quoting. V4 weights remain MIT-licensed for self-host escape hatch purposes.