DeepSeek V4.1-Flash: the 100x cached-price spread and the Sept 14 reroute that lasted a day

Published September 11, 2026 · Updated September 11, 2026By ABD Legacy LLC
DeepSeek V4.1 Flash pricing deepseek-flash rate Is deepseek-v4-pro retired Off-peak cached pricing

The short answer. DeepSeek V4.1-Flash is live on the API as deepseek-flash, and the two legacy Flash names - deepseek-v4-flash and deepseek-v4-flash-vision-exp - are still accepted: DeepSeek serves those requests with the V4.1-Flash model and bills them at the Flash price. The previously announced reroute of deepseek-v4-pro to V4.1-Flash, which was due to begin at 04:00 UTC, Sept 14, 2026, was withdrawn on 2026-09-11, so V4 Pro keeps serving at unchanged billing. What survives the retraction is the spread inside the Flash row itself - $0.003 per 1M cached input off-peak against $0.30 per 1M cache-miss input at peak, a 100x per-token rate gap.

What shipped: DeepSeek-V4.1-Flash as deepseek-flash

DeepSeek shipped DeepSeek-V4.1-Flash on 2026-09-10. The launch note says to set your model to deepseek-flash, and the API docs repeat that as the model name.

One launch claim we do not repeat as a finding: the news page says tests put V4.1-Flash ahead of V4 Pro on performance, cost, speed and total runtime, but the parties who ran them are not named anywhere, so the claim cannot be checked.

The DeepSeek V4.1-Flash price sheet (USD per 1M tokens)

First-party DeepSeek rates and third-party marketplace rates are separated below, because they are not the same product: $0.30/$1.20 is DeepSeek's own peak pair, and it is also what competing hosts charge - it is not the model's headline price.

Model / endpointInput /1MOutput /1MCached input /1MNotes
deepseek-flash - DeepSeek first-party, off-peak$0.15 (cache miss)$0.60$0.003Official sheet. Off-peak is every hour outside the peak windows below.
deepseek-flash - DeepSeek first-party, peak$0.30$1.20$0.006Peak = Mon-Fri 01:00-04:00 and 06:00-10:00 UTC; exactly 2x off-peak.
deepseek-v4-pro - DeepSeek first-party, off-peak$0.66$1.98$0.022Still served after 2026-09-14; model version DeepSeek-V4-Pro-0813.
deepseek-v4-pro - DeepSeek first-party, peak$1.32$3.96$0.044Unchanged rates.
deepseek/deepseek-v4.1-flash - OpenRouter headline$0.15$0.60$0.003OpenRouter's own "In / Out Price"; equals DeepSeek off-peak. Not a max-effort variant price.
Via SiliconFlow, Modal, Wafer, GMICloud, io.net, NovitaAI, Morph (third-party hosts)$0.30$1.20$0.006-$0.03Third-party hosts on OpenRouter; cache read varies by host.
Via DeepInfra$0.20$0.60$0.006OpenRouter provider row, context 1,048,576.
Via Fireworks$0.22$0.66$0.007OpenRouter provider row.
Via Venice$0.375$1.50$0.0075OpenRouter provider row.
Artificial Analysis listing - "V4.1 Flash (Reasoning, Max Effort)"$0.30$1.20cache discount 98%Artificial Analysis states these are based on DeepSeek's API - i.e. DeepSeek peak rates, not an OpenRouter price.

The 100x, stated precisely. The headline multiple is a per-token rate comparison, not a reduction in your total bill, and it depends on which two rates you put side by side: $0.003 per 1M cache-hit input off-peak vs $0.30 per 1M cache-miss input at peak = 100x. Compare the cached rate with off-peak cache-miss input ($0.15 per 1M) instead and the same pair is 50x. A mixed workload that misses the cache most of the time lands far below either number - the cache-hit ratio, not the list price, decides what you pay.

Where $0.30/$1.20 actually comes from. It is DeepSeek's own peak cache-miss input and peak output pair, and it is the pair Artificial Analysis lists for "V4.1 Flash (Reasoning, Max Effort)". On OpenRouter it is what third-party hosts charge; OpenRouter's own headline for the model is $0.15 in / $0.60 out with $0.003 cache reads.

Predecessor Flash rates, for anyone still carrying them in a spreadsheet: the retired Flash rows were off-peak $0.007 cache-hit / $0.22 cache-miss / $0.66 output, now $0.003 / $0.15 / $0.60. Those old rows are no longer on the live price sheet, so this comparison comes from two independent third-party pricing trackers, not from DeepSeek.

Cost-per-task reference, Artificial Analysis: $0.27 per Intelligence Index task and $476.89 to run the full Intelligence Index evaluation on this model.

The Sept 14 reroute that was withdrawn after a day

This is the part most coverage still gets wrong, because DeepSeek's own two sets of pages still disagree. The announced reroute of deepseek-v4-pro is not in force.

Two first-party pages, still contradicting each other. DeepSeek's marketing news page continues to advertise the reroute; the API docs - the pages that govern what the API actually bills - retract it. If you quote the news page as current, you will be wrong.

Why almost nobody noticed. The reversal was inserted into the existing 2026-09-10 changelog entry instead of being published as a new dated entry, so the changelog's newest heading still reads Sept 10 and a reader checking "what is new since the launch" sees nothing new. The withdrawal also came with no notice period, and DeepSeek publishes no versioned deprecation policy - the announcement gave four days, the reversal gave none. Independent trackers caught it: a pricing tracker published a before-and-after of the footnote, and a same-week analysis piece reached the same conclusion from the two first-party pages.

The operational lesson: on this platform, a dated first-party announcement is not a commitment until the API docs carry it, and even then it can be edited in place without a new changelog entry. Pin nothing to a future date you have not re-checked that week.

Migration checklist

Six checks, in order. The first four are the ones that cost money if you skip them.

Benchmarks: what was measured, and by whom

Every figure below names the party that produced it. DeepSeek's own table was run in DeepSeek's harness at maximum reasoning effort - those numbers are vendor-reported, not independent, and they are labelled that way throughout.

MeasureV4.1-FlashComparatorsMeasured by
Intelligence Index (composite)40 (Reasoning, Max Effort; #6 of 113 in its class)GPT-5.6 Sol (max) 47; Claude Opus 5 (adaptive reasoning, max effort) 51; DeepSeek V4 Pro 0813 36Artificial Analysis
Output speed198.6 output tokens/sec (#5 of 113)-Artificial Analysis
Cost per Intelligence Index task$0.27 (#19 of 113)-Artificial Analysis
Intelligence Index, same-day third-party test40 - ahead of V4 Pro 0813 (36), behind Kimi K3 (44) and GLM-5.3 (45)Kimi K3 44; GLM-5.3 45KuCoin news report (third party)
DeepSWE v1.174.2 (DeepSeek-run)Claude Opus 5 74.0; GPT-5.6 Sol 73.0 (both as reported in DeepSeek's table)VentureBeat, reporting DeepSeek-run evaluations
Terminal-Bench 3.030.0Claude Opus 5 43.3VentureBeat, reporting DeepSeek's table
Terminal-Bench 4.031.2Claude Opus 5 51.8VentureBeat, reporting DeepSeek's table
GPQA Diamond / SEC-Bench ProGPT-5.6 Sol leads V4.1-Flash in DeepSeek's own tableGPT-5.6 SolVentureBeat, reporting DeepSeek's table

What Artificial Analysis supports, and what it does not. It does support a narrower claim: V4.1-Flash is fast and cheap, and it beats its own predecessor and V4 Pro (40 vs 36). It does not support the idea that V4.1-Flash tops GPT-5.6 Sol or Claude Opus 5 - on the composite it sits below both, 40 against 47 and 51. VentureBeat's own body makes the same point: it lists benchmarks where Opus 5 and Sol lead, calls the results "not uniformly dominant", and frames the evidence as a price-performance thesis rather than a clean intelligence lead. Only the headline goes further than the numbers do.

Vendor-reported table (DeepSeek's own harness, max reasoning effort, temperature 1.0, top_p 0.95): GPQA Diamond 90.9; HLE 36.8 (39.1 text-only); Codeforces 3471; MathArena Apex 65.6; Terminal-Bench 2.1 90.6; Terminal-Bench 3.0 30.0; Terminal-Bench 4.0 31.2; DeepSWE v1.1 74.2; ProgramBench 20.3; NL2Repo-Bench 65.4; CyberGym 88.1; SEC-Bench Pro 62.8; ExploitGym 15.3; HLE with tools 63.9; Automation-Bench 54.8; Agents' Last Exam 31.8; Chartography with tools 78.9; BabyVision with tools 89.6; ZeroBench-main with tools 49.0. All of these are DeepSeek's numbers about DeepSeek's model.

Effort caveat on any quoted score: DeepSeek ran its comparison table at its maximum effort setting of 100. Raising effort from 25 to 100 lifts DeepSWE v1.1 from 66.0% to 74.2% and Terminal-Bench 2.1 from 82.4% to 90.6%, but consumes roughly 2.5 times as many output tokens. The public API presets are low / high / max = 50 / 75 / 100 (VentureBeat, quoting DeepSeek). Benchmark a score and a token bill together, or you are comparing two different products.

Frequently asked questions

DeepSeek V4.1 Flash pricing

DeepSeek's own API prices deepseek-flash at $0.15 per 1M input tokens and $0.60 per 1M output tokens off-peak, with cache hits at $0.003 per 1M. Peak rates are exactly double - $0.30, $1.20 and $0.006 - during Monday to Friday 01:00-04:00 and 06:00-10:00 UTC; every other hour is off-peak. deepseek-v4-pro keeps its own live row at $0.66 input, $1.98 output and $0.022 cached input off-peak (peak $1.32, $3.96, $0.044). On OpenRouter the headline for deepseek/deepseek-v4.1-flash is the same off-peak pair as DeepSeek's - $0.15 in, $0.60 out, $0.003 cache read - while third-party hosts on OpenRouter list $0.30 in and $1.20 out.

is deepseek-v4-pro retired

No. DeepSeek announced on 2026-09-10 that deepseek-v4-pro requests would be routed to V4.1-Flash from 04:00 UTC on Sept 14, 2026, then withdrew that plan on 2026-09-11. The current wording on DeepSeek's own API docs, pricing page and changelog is that DeepSeek will continue providing API services for DeepSeek V4 Pro after September 14, 2026 with the billing method remaining unchanged. deepseek-v4-pro keeps its own pricing row (model version DeepSeek-V4-Pro-0813) at $0.66 input, $1.98 output and $0.022 cached input off-peak. There is no Sept 14 migration and no retirement date in force.

DeepSeek V4.1 Flash vs GPT-5.6 Sol

On Artificial Analysis's Intelligence Index, V4.1-Flash at Reasoning, Max Effort scores 40 while GPT-5.6 Sol at max effort scores 47 (Claude Opus 5 scores 51) - so the composite index favors the rivals, not V4.1-Flash. VentureBeat's own body reported the same picture, listing benchmarks where Claude Opus 5 leads V4.1-Flash 43.3 to 30.0 on Terminal-Bench 3.0 and 51.8 to 31.2 on Terminal-Bench 4.0, and where GPT-5.6 Sol leads on GPQA Diamond and SEC-Bench Pro in DeepSeek's own table; its headline says further than its body does. Where V4.1-Flash leads is price: $0.003 per 1M cached input off-peak against the $0.40 per 1M cache-read rate press reporting attributes to GPT-5.6 Sol, and $0.15/$0.60 against Sol's $4/$20 headline. Artificial Analysis also measures V4.1-Flash as fast - 198.6 output tokens per second, 5th of 113 models - and cheap per task, at $0.27 per Intelligence Index task.

cheapest 1M context model 2026

Scope first: this compares the 1M-context models whose published rates we were able to verify on 2026-09-11, at list price, and it is not a global cheapest-model claim. DeepSeek V4.1-Flash has a 1M-token context and the lowest cached-input rate in that set - $0.003 per 1M off-peak, $0.006 at peak - against press-reported cache-read rates of $0.40 per 1M for GPT-5.6 Sol, $0.50 for Claude Opus 5 and $0.30 for Kimi K3. Headline pairs in the same set: V4.1-Flash $0.15/$0.60 off-peak ($0.30/$1.20 at peak), GPT-5.6 Sol $4/$20, Claude Opus 5 $5/$25, Kimi K3 $3/$15. Hosted variants can price differently - third-party hosts on OpenRouter list V4.1-Flash at $0.30/$1.20 - so verify the endpoint you actually call.

DeepSeek Flash off-peak cached pricing

A cached input token on DeepSeek's own API costs $0.003 per 1M off-peak and $0.006 per 1M at peak - off-peak is exactly half of peak. Off-peak means every hour outside Monday to Friday 01:00-04:00 and 06:00-10:00 UTC. That $0.003 rate is the low end of the 100x per-token spread against the $0.30 per 1M peak cache-miss input rate; the same cached token costs more through third-party hosts on OpenRouter, whose cache-read rates vary by host and sit above DeepSeek's own. The spread is a per-token rate comparison, not a discount on your whole bill - it only pays out on the share of your traffic that actually hits the cache.

Put the Flash row next to the models you already run

See the V4.1-Flash row in the AI Agency Pricing Calculator

The calculator carries the DeepSeek rows - V4.1-Flash through the legacy deepseek-v4-flash name, plus V4 Pro - alongside the other models you mix in production.

Sources

Accuracy note: every figure on this page traces to the research fact sheet compiled on 2026-09-11 (sources fetched 2026-09-11 between 22:00 and 22:12 UTC; fact sheet sha256 9f3fbbe21866d9f9ab40c433fc50bb246e94f0c6ce06a2541944111e32b48703). No hands-on model testing was performed for this page: benchmark numbers are reproduced from the party that measured them and are attributed in the tables above, and DeepSeek's own harness results are labelled vendor-reported. One claim carried by other coverage is deliberately absent here: the assertion that V4.1-Flash beats GPT-5.6 Sol and Claude Opus 5 - Artificial Analysis measures the opposite order on its composite index, and a separate reported AutomationBench figure could not be verified on Artificial Analysis's own page, so no part of it is repeated. Rates move frequently - re-verify before quoting, and note that DeepSeek has edited first-party pricing text in place within a single day.