The DeepSeek V4 family
This family has two active versions, and the engineering choice between them inverts what a standard flash/pro structure normally implies: Flash runs at a rate three times lower, its concurrency ceiling is five times higher, and the Responses API is still not active on Pro. The default configuration should be Flash.
- Pro costs three times more
- Flash concurrency ceiling is five times higher
- a one-million-token window on both versions
Last checked: This is a reference page. It is re-checked against the vendor sources and updated when a new version ships. Change log
Which version belongs in your architecture
Flash. DeepSeek own table logs the same one-million-token context window and the same 384K-token output ceiling for both models, so raw capacity is not the differentiator; rate is. Flash charges $0.14 for uncached input and $0.28 for output; Pro charges $0.435 and $0.87, roughly three times as much. The concurrency ceiling is 2500 requests on Flash and 500 on Pro.
So the cheaper version carries five times the parallel-processing capacity. In architectures where traffic arrives in bursts, this parameter sometimes matters more than the rate itself, and any analysis that treats Pro as simply the higher tier has usually not accounted for this column.
Documentation the vendor published that gets quoted less often
The change log, entry dated 31 July 2026, states the official Flash release shipped with sharply enhanced agentic capability and that its benchmark results are far beyond V4-Pro-Preview. The benchmark data logged beside that claim: Terminal Bench 2.1 at 82.7, NL2Repo at 54.2, Cybergym at 76.7, DeepSWE at 54.4, and Toolathlon verified at 70.3.
This point requires precision, because the distinction changes the conclusion. That comparison runs against the preview of Pro, not against the Pro sitting in today price table, and the same change log states the official Pro release is coming soon. So the precise claim is not that Flash outperforms Pro; the precise claim is that the vendor has not demonstrated Pro superiority in any document, while its rate is three times higher. For a deployment decision, that is sufficient.
One further technical detail is logged in the same entry: Flash 0731 did not change the preview architecture or model size and was only re-post-trained. The jump in benchmark data came from the training process, not from added hardware resources.
Two capabilities active only on Flash
The capability table carries a column rarely covered in summaries: the Responses API is active on Flash and not on Pro. The footnote to that same table states Pro support will be added in early August 2026. This page was checked on 12 August 2026 and that sentence was still logged as current, so either the capability has not shipped or it has shipped and the documentation was not updated. Either reading carries the same implication for planning: do not rely on it without direct confirmation.
Flash is also natively adapted for Codex, per the same change log. If your tooling is Codex, Claude Code, or OpenCode, this version can be deployed as the backend model with no code changes required.
The timeline, the clearest record in this group
24 April 2026: V4 released, and the API accepts both the V4 Pro and V4 Flash identifiers. The same entry announces that the two legacy identifiers, deepseek-chat and deepseek-reasoner, will be decommissioned three months later, on 24 July 2026, and until then point respectively at Flash non-thinking and thinking modes. 31 July 2026: the official Flash release, self-logged as public beta.
This level of date logging is unmatched among the Chinese vendors in this reference. Moonshot and Z.ai log no dates for any version, which is exactly why the GLM-5 family page carries a release order but no timeline.
Release timeline
Each version with its own release date, and what changed against the one before it.
- DeepSeek V4 Flash Vision preview
- DeepSeek V4 Flash current The lowest price point in the coding table, and the highest output ceiling in it.
- DeepSeek V4 Pro current
Using it from Iran
This column carries a date and says how it was checked, because most listicles guess it. Where we have only read a vendor policy page, the note below says exactly that.
- Reachable
- not checked
- Payment
- not checked
- Free tier
- no
How we checked: DeepSeek publishes no official supported-countries list, so the access field is logged as unverified. No test has been run from an Iranian connection either.
Documented strengths
- The cheapest row in our coding table, even at the uncached input rate
- A one-million-token context window and up to 384K-token output on both versions
- A 2500-request concurrency ceiling on Flash, five times that of Pro
- A precisely dated change log, down to logging the shutdown date of legacy identifiers
Known weaknesses and limits
- The pricing page states the overall API rate will rise in the near future with a significant increase expected, so a long-term budget should not be locked to today figure.
- The official V4 Pro release has not shipped, and the vendor benchmark comparison runs against the Pro preview rather than the version being sold today.
- The Responses API is not active on Pro, and the documentation still logs the passed early-August date.
- Flash itself is logged as public beta in the change log, a factor that must be weighed for any sensitive workload.
- FIM Completion is operational in non-thinking mode only, on both versions.
- Access from Iran is unverified, and DeepSeek publishes no country list.
Technical verdict
Set Flash as the default configuration and deploy Pro only after measuring your own workload directly and observing a real difference, because the vendor has not demonstrated one in any document. Do not fix today rate into a long-term contract either, since an increase has already been announced.
Frequently raised questions, with documented answers
Does V4 Pro outperform V4 Flash
DeepSeek has published no benchmark data placing Pro ahead. The only published comparison puts Flash far beyond the Pro preview. Pro runs at three times the rate, carries a fifth of the concurrency ceiling, and the Responses API is not active on it.
What is the status of the deepseek-chat and deepseek-reasoner identifiers
They were decommissioned on 24 July 2026, an announcement logged three months earlier in the 24 April change log entry. Until that date they pointed respectively at the non-thinking and thinking modes of deepseek-v4-flash.