DeepSeek V4 API Price Increase (Aug 16, 2026): New Peak/Off-Peak Rates

Published August 16, 2026By ABD Legacy LLC
DeepSeek pricing AI model cost API price hike

What happened. DeepSeek's new peak/off-peak API pricing for the V4 model family took effect at 16:00 UTC on August 16, 2026 — the schedule DeepSeek published on its official pricing page and confirmed in its Aug 13, 2026 changelog entry. Off-peak rates are half of peak rates; peak hours are 01:00–04:00 and 06:00–10:00 UTC. The old "significant increase expected" warning is gone — the concrete numbers are now live on api-docs.deepseek.com/quick_start/pricing (verified August 16, 2026).

What the new DeepSeek V4 rates are (per 1M tokens)

ModelInput, cache hitInput, cache missOutputWindow
V4-Flash$0.007$0.22$0.66OFF-PEAK
V4-Flash$0.014$0.44$1.32PEAK
V4-Pro$0.022$0.66$1.98OFF-PEAK
V4-Pro$0.044$1.32$3.96PEAK

Previous (pre-hike, effective until 16:00 UTC Aug 16): V4-Flash $0.0028 (cache hit) / $0.14 (cache miss) / $0.28 output; V4-Pro $0.003625 / $0.435 / $0.87.

How much did prices actually go up?

Alongside the pricing change, the Aug 13 changelog entry also marked DeepSeek-V4-Pro GA (APP/Web/API, model name unchanged, native Responses API, three thinking-effort levels: low/high/max).

What the increase costs an agency

DeepSeek was the default "cheapest stack" assumption in agency cost models. That assumption does not survive the new table:

What to do now

  1. Re-run every estimate that assumed DeepSeek pricing with the new table (use this calculator's model strategy selector for the updated comparison).
  2. Schedule batch and non-urgent work off-peak (outside 01:00–04:00 and 06:00–10:00 UTC) to lock the half-price tier.
  3. Exploit cache hits. Prompt-cached workloads at off-peak are the cheapest live option in the family.
  4. Re-benchmark alternatives: Gemini 3.7 Flash intro ($0.75/$3.75, through 2026-12-31), Grok 4.6 ($2/$6, 500K context), Qwen 3.8 Max ($2/$6 open weights), and self-hosted V4 (weights are MIT-licensed — sustained workloads can be served directly via vLLM-class tooling).

Re-run your estimate with the new DeepSeek rates

Open the AI Agency Pricing Calculator →

Model DeepSeek V4, Gemini 3.7 Flash, Grok 4.6, Qwen 3.8 Max, Claude Sonnet 5, and local self-host — with current 2026 rates.

Frequently asked questions

When did the DeepSeek V4 API price increase take effect?

The new peak/off-peak rates took effect at 16:00 UTC on August 16, 2026, per DeepSeek's official pricing page (footnote 1) and the Aug 13, 2026 changelog entry. Peak hours are 01:00–04:00 and 06:00–10:00 UTC; all other hours are off-peak at half the peak rate.

How much did DeepSeek V4 API prices increase?

V4-Flash peak pricing is $0.014 (cache hit) / $0.44 (cache miss) / $1.32 output per 1M tokens vs $0.0028/$0.14/$0.28 before — roughly 5x on cache-hit input, 3.1x on cache-miss input, and 4.7x on output. V4-Pro peak is $0.044/$1.32/$3.96 vs $0.003625/$0.435/$0.87 — roughly 12.1x on cache-hit input, 3.0x on cache-miss input, and 4.6x on output. Off-peak rates are half of peak.

Is DeepSeek still the cheapest AI API after the increase?

Only off-peak cache-miss input remains near the old ballpark, and even that is 1.5–1.6x higher. At peak rates, V4-Flash output ($1.32/1M) now sits above Gemini 3.7 Flash intro pricing ($0.75/$3.75), Muse Glimmer on Together AI ($0.35/$1.50), and close to Grok 4.6 ($2/$6) and Qwen 3.8 Max ($2/$6). DeepSeek is no longer the default "cheapest stack" assumption — re-run your routing math with the new table.

What should agencies do about the DeepSeek price increase?

Update cost assumptions immediately: use peak/off-peak rates in estimates, schedule non-urgent batch work off-peak, treat cache hits as the lever (off-peak cache-hit input is $0.007 Flash / $0.022 Pro), and re-benchmark alternatives (Gemini 3.7 Flash intro, Grok 4.6, Qwen 3.8 Max, self-hosted MIT-licensed V4 weights for sustained workloads).

Sources

Accuracy note: All new rates, peak hours, and the effective date (16:00 UTC Aug 16, 2026) are verified against DeepSeek's official pricing page and changelog on Aug 16, 2026. Multiplier math (e.g., 5x, 12.1x) is derived from the published before/after rates, not a DeepSeek statement. The 50%–1,100% range is TechStartups' framing; "quadruples peak prices" is Bloomberg's framing (paywalled). DeepSeek reserves the right to adjust prices — re-verify before quoting. V4 weights remain MIT-licensed for self-host escape hatch purposes.