Flash Killed Pro: The Day DeepSeek Retired V4-Pro
On September 14, 2026, DeepSeek redirected all requests from V4-Pro to V4.1-Flash, making its Flash model the new flagship after outperforming older versions.

Stock photo for illustration only, not from the actual event
- DeepSeek is phasing out V4-Pro and routing all traffic to V4.1-Flash.
- V4.1-Flash scored 90.6 on Terminal-Bench 2.1, beating top competitor models.
- SimpleQA factual knowledge score dropped from 55.2 down to 42.3.
- Three major AI labs (DeepSeek, Z.ai, Moonshot) are pursuing completely different architecture strategies.
September 14, 2026, 04:00 UTC marks an unusual milestone in artificial intelligence history. From that moment onward, every incoming request directed to deepseek-v4-pro gets automatically routed to V4.1-Flash, remaining that way until V4.1-Pro officially ships.
Put plainly, DeepSeek has retired its own flagship moniker and handed primary duties over to the model labeled Flash. Multi-party testing confirmed that V4.1-Flash surpasses V4-Pro across capability, cost, speed, and total runtime.

Stock photo for illustration only, not from the actual event
Looking strictly at execution work, the performance gap is significant. Scoring 90.6 on Terminal-Bench 2.1 outperforms its in-house sibling while also edging out Claude Opus 5 at 89.1 and GPT-5.6 Sol at 88.8, all while priced under a dollar per million tokens.
"We're phasing out V4-Pro."
DeepSeek Announcement
However, capabilities did not improve uniformly. SimpleQA scores dropped from 55.2 to 42.3, demonstrating a regression in answering factual world knowledge questions even as agentic capabilities improved.
The decision by DeepSeek to replace its flagship Pro model with a Flash variant highlights a major shift in the AI industry. Prioritizing architectural efficiency, reduced KV cache memory, and lower serving costs over traditional naming conventions suggests that optimized mid-tier models may increasingly usurp top-tier roles, prompting competitors to rethink their deployment strategies.
The three major labs are betting on completely different roadmaps:
- Z.ai: Keeps both tiers, building a 753B parameter main model alongside a 320B GLM-5.3-Flash with a significantly smaller KV cache.
- DeepSeek: All-in on a single lane, moving its entire API to Flash with a new 552B parameter architecture that activates less per token.
- Moonshot: Focuses purely on size with Kimi K3 at 2.8 trillion parameters, omitting a lower-priced scaled-down variant.

Stock photo for illustration only, not from the actual event
Data from Artificial Analysis shows that while overall scores vary narrowly, task costs can differ by up to eight times. Users must also account for mandatory reasoning modes in models like Kimi K3, which permanently generate billable reasoning tokens for every single request.
Source: Dev.to
Found something wrong in this article? Report an issue with this article
Comments
Leave a Comment