一个要在 9 月 14 日取代 V4 Pro 的模型,有两项分数没打过 V4 ProThe Model Set to Replace V4 Pro on September 14 Loses to It on Two Benchmarks
官方口径原话(价格页注释、新闻页、更新日志三处一致):
The official wording, verbatim (consistent across three places: the pricing-page note, the news page, and the changelog):
一、先把话说在前面:不是全面超越
官方口径原话(价格页注释、新闻页、更新日志三处一致):
经多方测试,V4.1 Flash 在性能、费用、速度、总用时等各项指标上已全面超越 DeepSeek V4 Pro,因此我们计划有序下线 V4 Pro。北京时间 2026 年 9 月 14 日 12:00 之后,至未来 V4.1 Pro 上线之前,用户访问 deepseek-v4-pro 的请求将全部路由到 V4.1 Flash,并按 V4.1 Flash 单价计费。
以上是官方说法,我们原样引用,不做二次背书。
而在官方自己那份 51 页技术报告里,附了一张完整对照表。表里 V4.1 Flash 有两项是低于 V4 Pro 的:
| 基准 | V4.1-Flash | V4-Pro(0813) | 差值 |
|---|---|---|---|
| --- | --- | --- | --- |
| GPQA Diamond | 90.9 | 92.4 | -1.5 |
| HLE | 36.8(39.1*) | 42.7* | -5.9 |
那个星号必须说清楚:42.7 与 39.1 都带星号,通常表示"纯文本子集"口径,即 HLE 这一项两组数字并非完全同一口径下的直接对比。写出星号的存在,是因为省略它就成了选择性引用。
它真正大幅领先的是 Agent 与终端操作类:
| 基准 | V4.1-Flash | V4-Pro | V4-Flash(0731) | Opus5 | GPT5.6-Sol |
|---|---|---|---|---|---|
| --- | --- | --- | --- | --- | --- |
| Terminal-Bench 2.1 | 90.6 | 87.9 | 82.7 | 89.1 | 88.8 |
| Terminal-Bench 3.0 | 30.0 | 11.8 | 7.6 | 43.3 | 51.8 |
| Terminal-Bench 4.0 | 31.2 | 12.4 | 7.0 | 39.9 | 51.8 |
| DeepSWE v1.1 | 74.2 | 62.7 | 54.4 | 73.0 | 74.0 |
| CyberGym | 88.1 | 83.3 | 76.7 | 84.5 | 84.5 |
| Automation-Bench | 54.8 | 43.2 | 37.7 | 50.3 | 45.8 |
| Agents' Last Exam | 31.8 | 25.7 | 25.2 | 28.6 | 26.7 |
| Codeforces Rating | 3471 | 3348 | 3289 | — | — |
所以更准确的描述是:在 Agent 与终端操作类任务上大幅超越,在研究生级推理类任务上小幅落后。而官方正是用"全面超越"这六个字,为 9 月 14 日的全量路由与"有序下线 V4 Pro"提供依据的。
三个陷阱必须点破。
陷阱一:Terminal-Bench 版本分裂。 2.1 拿到 90.6,这是官方对外新闻稿放的数字。但同一基准家族,3.0 只有 30.0、4.0 只有 31.2——落后 Opus5 的 43.3 和 39.9,也落后 GPT5.6-Sol 的 51.8。新旧两版结论完全相反,只引任何一个都是半个事实。
陷阱二:DeepSWE 的"0.2 分"要写准。 V4.1-Flash 是 74.2,对照表里 GPT5.6-Sol 是 74.0、Opus5 是 73.0。所以准确说法是:领先 GPT-5.6 Sol 仅 0.2 分,领先 Opus 5 1.2 分。这不是"险胜 Opus 5 0.2 分"——部分二手转述搞错了这个对照关系。
陷阱三:基准是谁的。 DeepSWE、CyberGym、Automation-Bench 这几个拿高分的项目,都是 DeepSeek 自家或合作方的基准,不是独立第三方榜单。不是说分数无效,而是目前缺一份外部审计。
二、第一方:官方自报的架构,以及"为什么长这样"
V4.1 Flash 于 9 月 10 日发布,552B 参数 MoE(骨干口径;另有 196B Engram 条件记忆,故总量高于 552B)。
架构上最关键的变化是全新的 Causal-Encoder-Decoder(CED)结构:40 层 = 20 层因果编码器 + 20 层解码器。它带来非对称激活——预填充阶段每 token 激活 80 亿参数,解码阶段每 token 激活 160 亿参数。
官方给出的设计动机值得记下来:长程 Agent 的工作负载是"输入重"的,读工具输出、读文件、读历史的量远大于自己生成的量,所以把读写两阶段拆开分别降本。这段话回答了"这个模型为什么长这样"——它不是通用能力提升,是为 Agent workload 定制的取舍。
其他公开事实:384 个路由专家加 1 个共享专家、每 token 激活 6 个;最大位置编码 1,048,576,即 1M 上下文;FP4 主 KV 缓存。训练走 45 万亿 token 多模态语料从零训练,稀疏注意力先在 64K 序列上训练、到 34 万亿 token 时把上下文扩展到 100 万,后训练为 SFT、RL、在线蒸馏(OPD)。视觉侧为 DeepSeek-ViT 从零训练加两层 MLP 投影器,从预训练初期即与文本联合处理——原生多模态,区别于上一代 vision-exp 的外挂式方案。另支持 1 到 100 整数级推理强度调节。
KV 缓存那三个数字,参照系各不相同,最容易引混:
- 相对上一代 V4-Flash,HBM 占用降到 1/4,SSD 占用降到 1/8;
- 相对初代 V1(2023 年 11 月),降幅是 437 倍。
演进链条:V1 389,120 字节/token,到 V3.2 的 48,068,到 V4-Flash 的 3,514,再到 V4.1-Flash 的 890。百万 token 时活跃状态小于 1GB。
反面也要记一笔:有第三方指出部署摩擦——没有标准 chat template,只有参考的 Python encoder 与 Rust 工具。开源不等于好上手。
三、第二方:平台实测,最冷的一组数字
OpenRouter 在 9 月 10 日上架该模型,定价为输入 0.30 美元/百万、输出 1.20 美元/百万、cache-read 0.006 美元/百万。更值得注意的是生产实测:P50 延迟 0.76 秒,吞吐 134 tokens/秒。
这个数字要和内测期反馈放在一起看:内测期社区流传的速度散落在 284、355、420、507 tok/s 各处,出处与场景各不相同;到了生产环境,是 134。原因不难理解:内测期单账号限流 20 并发,路上基本没车;正式版并发上限 2500。内测峰值不等于生产常态,用内测数字做容量规划会翻车。
四、第三方:社区真跑,最有血肉也最矛盾
新浪那位开发者的实测,同一双手、同一天、同一个模型,跑出了两个极端。
正面:"我的世界"体素世界(地形、方块、第一人称、背包、光照、水、生存循环),模型干满 52 分钟,交付 7550 行完整游戏,53 个 Node.js 单测与 38 个浏览器集成测试全过。更有意思的是它自己做代码审查,找出 15 个 bug 逐个修复、还补了回归测试。
反面:同一位开发者测"怪兽城市大战",跑两个多小时没出结果,无奈放弃。
中间态也有:3D 鹈鹕骑车动画,12 分钟交付完整 Three.js 项目(6 个源文件),但车架穿帮、车把手缺失、鹈鹕双腿像两截骨头拼接——能用,有瑕疵。
成本方面,含未完成测试在内,总计 1.56 亿 tokens、825 次 API 调用、1.93 美元(约 13 元),他本人的评价是"1.5 亿 tokens、13 块,还是挺香的"。
Reddit r/LocalLLaMA 上的反馈是实测约 2.24 倍加速、最高约 30% 的 token 效率提升;也有人推测速度提升部分来自内测期的低并发而非架构本身,这个质疑应当保留。
以下为内测版数据,正式版尚无同场景复测:
- 3D S 形停车:V4.1 零碰撞 25.4 秒,V4 Flash 用时 11 分 27 秒且画面上下颠倒;
- 小店任务:1 分 41 秒 vs 5 分 54 秒;
- 14 组任务:端到端 34.5 分钟 vs 30.5 分钟,V4.1 反而更慢,其中等工具 18.7 分钟 vs 6.6 分钟。吐字快不等于交活快;
- 自认知混乱:问"你是谁"答"Claude"(发布后是否修复未查到专门验证);
- 暂不支持视频,只测了图片。
关于成本,本篇不下定论。"1.5 亿 tokens 花 13 元"与"5 分钟 10 块钱",任务规模与缓存命中率都不可比。能否真省钱取决于缓存命中率与输入输出比;也有媒体把这轮降价解读为针对此前涨价致用户流失的补救动作。请用你自己的任务结构去测。
五、第四方:流程与动机的质疑
通知流程过于仓促。 路由通知只走了 9 月 9 日的一个官方群帖,没有 API 变更日志,没有正式弃用公告,从通知到生效不到 48 小时。作为对照,OpenAI 关停 Sora API 提前 6 个月发正式通知。Hacker News 上有一条评论值得原样引用:"如果你在 V4 Pro 上验证过工作流,你未必想突然在生产环境里测 V4.1 Flash。"
版本号有误导性。 官方说这是"新架构家族中尺寸最小的模型",却给了它一个".1"的版本号。开创全新架构家族的模型,与只调整后训练配方的模型,在版本号上只差一个小数点——开发者没法据此判断该投多少评估预算。
至今没有第三方审计。 截至 9 月 11 日,Artificial Analysis 未发布实测,模型页仍标注"尚未被任何推理服务商部署";LMArena、SWE-bench 官方榜、LiveCodeBench 均检索不到该模型。
把这三件事合起来,就得到了一句话:一个 Flash 档位的模型全面超越了自家的旗舰——这是一个被用来退休该旗舰的承重主张,目前却没有任何外部审计。
还有一个实用风险:可复现性。 旧模型名 deepseek-v4-flash、deepseek-v4-flash-vision-exp 已下线,请求被静默路由到 V4.1 Flash;若测试套件仍用旧名调用,跑的可能是新模型。
调用名与实际模型名是两个层级,不要混为一谈:
- 请求时填的调用名是 <code>deepseek-flash</code>——这是官方要求使用的名字;
- 服务端解析后,实际提供服务的是 <code>deepseek-v4.1-flash</code>,响应里 model 字段返回的也是后者;
- 官方版本字符串则写作 DeepSeek-V4.1-Flash。
所以如果你在响应的 model 字段里看到 deepseek-v4.1-flash,不必困惑,那不是你填错了,而是调用名被解析后的结果。真实的风险在另一头:用旧名 deepseek-v4-flash 或 deepseek-v4-flash-vision-exp 调用的测试套件,跑的其实已经是新模型。
六、DSH 用户:这次是硬绑定,我们的读者最该看这段
DeepSeek Harness v0.1.5 与 V4.1 Flash 同步发布,而且官方对这个模型在 DSH 里做了专项训练优化。
官方 Harness 更新说明的原话是:"新版本与 DeepSeek V4.1 Flash 模型训练深度结合,模型在 DSH 不同配置中都进行了专项训练和优化"——覆盖标准模式、程序化工具调用(PTC)模式、极简模式三种配置。
最有价值的一条在这里: 使用 V4.1 Flash 时,DSH 新版本支持在保留已有 KV Cache 的情况下更新系统提示词。对连续运行的 Agent 任务,这意味着中途可调整任务方向而无需重算整个上下文。
请把它和第二节的架构放在一起看:CED 把读写拆开、KV 缓存压到 890 字节/token、百万 token 活跃状态小于 1GB——这些技术设计落到产品上,就是"改系统提示词不重算上下文"这一个功能。技术不是炫技,它最后落成一个用户能感知的动作,这是全文最漂亮的一处闭环。
其他更新:Web 界面支持上传图片与 PDF;Sidebar 提供文件树与预览(Markdown、HTML、PDF、代码、图片);两侧开放标准化插件扩展入口;多 Agent 支持父子双向通信与插话停止,主 Agent 可为子 Agent 选模型与推理强度,另有实验性 Agent Teams(共享任务列表,默认关闭)。安装方式为 npx @deepseek-ai/dsh web。
身份要区分: DSH 是 DeepSeek 自有的开源 Agent 框架,未被列入官方合作方名单(该标签给的是第三方产品)。但 V4.1 Flash 是首个在 Harness 多配置下做专项训练优化的模型,这比合作方名单硬得多。
七、规格、价格与几个容易写错的事实
规格:上下文 1M、输出上限 384K、并发 2500(V4 Pro 为 500)、支持图像理解(V4 Pro 不支持)、思考模式默认开启且可切换。
价格(元/百万 tokens,9 月 10 日 12:00 生效):
| 档位 | V4.1-Flash 空闲 | V4.1-Flash 高峰 | V4-Pro 空闲 | V4-Pro 高峰 |
|---|---|---|---|---|
| --- | --- | --- | --- | --- |
| 输入缓存命中 | 0.02 | 0.04 | 0.15 | 0.30 |
| 输入缓存未命中 | 1.0 | 2.0 | 4.5 | 9.0 |
| 输出 | 4.0 | 8.0 | 13.5 | 27.0 |
高峰时段为北京时间周一至周五 9:00–12:00、14:00–18:00,其余为空闲,空闲价为高峰的一半。注意:V4 Pro 本次价格未变。 此前我们报道的 -60%、-33%、-11% 是 Flash 新价相对 Flash 旧价,不是相对 Pro。
许可证:多家技术媒体报道为 MIT 许可,但我们未能直连核对仓库 LICENSE 原件,故只能写到"据多家技术媒体报道为 MIT 许可,商用前请以仓库 LICENSE 原件为准",不建议据此推定商用无限制。
多模态:官方口径只有视觉与图像理解。有第三方称含音频,无官方依据,本文不采用。
V4.1 Pro 上线时间:官方未公布,本文不作推测。
八、所以,该怎么用它
第一,Agent 类长任务,值得认真评估。 架构、基准、社区实测三个来源方向一致:读写非对称的设计就是为它做的,Terminal-Bench 2.1、Automation-Bench、Codeforces 的跃升不是噪声。
第二,研究生级推理,别急着换。 GPQA Diamond 与 HLE 两项,官方自己的表摆在那里,分数是低的。
第三,所有基准分数先打个折,等你自己的结果。 生产吞吐 134 tok/s 与内测峰值的落差、缺席的第三方审计、静默路由带来的可复现性风险,都是真实摩擦。
最后回到开头:9 月 14 日 12:00 之后,V4 Pro 的请求全量路由到 V4.1 Flash 并按 Flash 价格计费,直到 V4.1 Pro 上线。无论你认不认同"全面超越"这四个字,时点是确定的——在那之前,把自己最关键的那条工作流跑一遍。
数据口径说明
本文数据采集于 2026 年 9 月 11 日;模型于 9 月 10 日发布;路由生效时点为 9 月 14 日 12:00(北京时间)。官方口径来自价格页注释、新闻页、更新日志与技术报告;平台实测来自 OpenRouter 上架信息;社区实测来自开发者公开分享与 Reddit 讨论。内测期数据已标注"内测版",不与正式版混用;竞品分数仅作陈述;"全面超越 V4 Pro"为官方说法,本文不做二次背书。
1. Let's Get This Straight Up Front: It's Not a Comprehensive Sweep
The official wording, verbatim (consistent across three places: the pricing-page note, the news page, and the changelog):
After testing by multiple parties, V4.1 Flash has comprehensively surpassed DeepSeek V4 Pro across every metric, including performance, cost, speed, and total elapsed time, so we plan to retire V4 Pro in an orderly manner. After 12:00 on September 14, 2026, Beijing time, and until a future V4.1 Pro launches, user requests to deepseek-v4-pro will all be routed to V4.1 Flash and billed at V4.1 Flash rates.
That's the official line, quoted as-is, without a second endorsement from us.
And in the company's own 51-page technical report, there's a full comparison table attached. In it, V4.1 Flash is lower than V4 Pro on two items:
| Benchmark | V4.1-Flash | V4-Pro(0813) | Difference |
|---|---|---|---|
| --- | --- | --- | --- |
| GPQA Diamond | 90.9 | 92.4 | -1.5 |
| HLE | 36.8 (39.1*) | 42.7* | -5.9 |
That asterisk needs spelling out: both 42.7 and 39.1 carry it, which usually indicates a "text-only subset" basis — meaning the two HLE numbers aren't a straight apples-to-apples comparison. We note the asterisk's existence because omitting it would make this a selective citation.
Where it genuinely leads by a wide margin is Agent and terminal-operation tasks:
| Benchmark | V4.1-Flash | V4-Pro | V4-Flash(0731) | Opus5 | GPT5.6-Sol |
|---|---|---|---|---|---|
| --- | --- | --- | --- | --- | --- |
| Terminal-Bench 2.1 | 90.6 | 87.9 | 82.7 | 89.1 | 88.8 |
| Terminal-Bench 3.0 | 30.0 | 11.8 | 7.6 | 43.3 | 51.8 |
| Terminal-Bench 4.0 | 31.2 | 12.4 | 7.0 | 39.9 | 51.8 |
| DeepSWE v1.1 | 74.2 | 62.7 | 54.4 | 73.0 | 74.0 |
| CyberGym | 88.1 | 83.3 | 76.7 | 84.5 | 84.5 |
| Automation-Bench | 54.8 | 43.2 | 37.7 | 50.3 | 45.8 |
| Agents' Last Exam | 31.8 | 25.7 | 25.2 | 28.6 | 26.7 |
| Codeforces Rating | 3471 | 3348 | 3289 | — | — |
So the more accurate description is: it greatly surpasses V4 Pro on Agent and terminal-operation tasks, and falls slightly behind on graduate-level reasoning tasks. And it is precisely with the phrase "comprehensively surpasses" that the company justifies the full routing on September 14 and the "orderly retirement of V4 Pro."
Three traps have to be called out.
Trap one: Terminal-Bench's version split. It scores 90.6 on 2.1 — the number the company put in its press release. But within the same benchmark family, it scores only 30.0 on 3.0 and 31.2 on 4.0 — behind Opus5's 43.3 and 39.9, and behind GPT5.6-Sol's 51.8. The old and new versions lead to opposite conclusions, so citing only one of them is half a fact.
Trap two: get DeepSWE's "0.2 points" right. V4.1-Flash scores 74.2; in the comparison table GPT5.6-Sol is 74.0 and Opus5 is 73.0. So the accurate statement is: it leads GPT-5.6 Sol by just 0.2 points, and Opus 5 by 1.2. It is not "a narrow 0.2-point win over Opus 5" — some secondhand accounts get that comparison wrong.
Trap three: whose benchmarks are these? The items where it scores high — DeepSWE, CyberGym, Automation-Bench — are all DeepSeek's own or partner benchmarks, not independent third-party leaderboards. That doesn't make the scores invalid, but there's currently no external audit.
2. Party One: The Architecture the Company Reports, and "Why It Looks Like This"
V4.1 Flash was released on September 10: a 552B-parameter MoE (backbone figure; there's also a 196B Engram conditional memory, so the total is higher than 552B).
The most important architectural change is an entirely new Causal-Encoder-Decoder (CED) structure: 40 layers = 20 causal encoder layers + 20 decoder layers. It produces asymmetric activation — 8 billion parameters activated per token in the prefill stage, 16 billion per token in the decode stage.
The stated design motivation is worth writing down: long-horizon Agent workloads are "input-heavy" — the volume of tool output, files, and history they read far exceeds what they generate — so the read and write stages are split apart to cut costs on each separately. That passage answers "why this model looks the way it does" — it isn't a general capability uplift, it's a trade-off custom-built for Agent workloads.
Other public facts: 384 routed experts plus 1 shared expert, 6 activated per token; maximum position encoding 1,048,576, i.e. a 1M context; FP4 main KV cache. Training ran from scratch on a 45-trillion-token multimodal corpus; sparse attention was first trained on 64K sequences, then the context was extended to 1 million at 34 trillion tokens; post-training used SFT, RL, and online distillation (OPD). On the vision side, DeepSeek-ViT is trained from scratch plus a two-layer MLP projector, processed jointly with text from the very start of pretraining — native multimodality, unlike the previous generation's bolt-on vision-exp approach. It also supports integer reasoning-effort adjustment from 1 to 100.
Those three KV cache numbers each use a different baseline, and they're the easiest thing to mix up:
- Relative to the previous generation, V4-Flash, HBM footprint drops to 1/4 and SSD footprint to 1/8;
- Relative to the first generation, V1 (November 2023), the reduction is 437×.
The evolution chain: V1 at 389,120 bytes/token, to 48,068 in V3.2, to 3,514 in V4-Flash, and then 890 in V4.1-Flash. At a million tokens, the active state is under 1GB.
A negative point also deserves a note: third parties have flagged deployment friction — there's no standard chat template, only a reference Python encoder and Rust tooling. Open source doesn't mean easy to get started with.
3. Party Two: Platform Measurements, the Coldest Set of Numbers
OpenRouter listed the model on September 10, priced at $0.30 per million input tokens, $1.20 per million output tokens, and $0.006 per million cache-read tokens. More notable are the production measurements: P50 latency 0.76 seconds, throughput 134 tokens/second.
That number has to be read alongside the beta-period feedback: during the beta, community speed figures were scattered across 284, 355, 420, and 507 tok/s, each with different sources and scenarios; in production, it's 134. The reason isn't hard to grasp: the beta capped accounts at 20 concurrent requests, so the road was basically empty; the release version's concurrency ceiling is 2500. Beta peaks are not production norms, and capacity planning off beta numbers will blow up in your face.
4. Party Three: Real Community Runs — the Most Human and the Most Contradictory
That developer from Sina ran the same model with the same hands on the same day and got two extremes.
The upside: a Minecraft-style voxel world (terrain, blocks, first-person view, inventory, lighting, water, survival loop) — the model worked a full 52 minutes and delivered a complete 7,550-line game, with all 53 Node.js unit tests and 38 browser integration tests passing. More interesting still, it reviewed its own code, found 15 bugs, fixed them one by one, and added regression tests.
The downside: the same developer tested a "monster city battle" and ran for over two hours with no result, then gave up.
There's a middle case too: a 3D pelican-riding-a-bike animation delivered a complete Three.js project (6 source files) in 12 minutes, but the frame clips through, the handlebars are missing, and the pelican's legs look like two pieces of bone stitched together — usable, with flaws.
On cost, including the unfinished tests, the total was 156 million tokens, 825 API calls, and $1.93 (about ¥13). His own verdict: "156 million tokens for ¥13 — still a pretty good deal."
Feedback on Reddit's r/LocalLLaMA reports roughly a 2.24× speedup in testing and up to about 30% better token efficiency; some also speculate that part of the speed gain comes from the beta's low concurrency rather than the architecture itself, and that doubt should be kept on the record.
The following is beta-build data; the release version has no repeat tests in the same scenarios:
- 3D S-shaped parking: V4.1 finished in 25.4 seconds with zero collisions; V4 Flash took 11 minutes 27 seconds and rendered the view upside down;
- Small shop task: 1 minute 41 seconds vs 5 minutes 54 seconds;
- 14 task sets: 34.5 minutes end to end vs 30.5 minutes — V4.1 was actually slower, with 18.7 minutes of that spent waiting on tools vs 6.6 minutes. Fast tokens don't mean fast delivery;
- Self-identity confusion: asked "who are you," it answered "Claude" (we found no dedicated verification of whether this was fixed after release);
- Video isn't supported yet; only images were tested.
On cost, this piece reaches no verdict. "156 million tokens for ¥13" and "¥10 in 5 minutes" involve task scales and cache hit rates that aren't comparable. Whether you really save money depends on your cache hit rate and your input-to-output ratio; some media also read this price cut as a remedial move after an earlier hike drove users away. Test it against your own task structure.
5. Party Four: Questions About Process and Motive
The notification process was far too rushed. The routing notice went out only as a single official group post on September 9 — no API changelog, no formal deprecation announcement — and less than 48 hours passed from notice to effect. By comparison, OpenAI gave six months' formal notice before shutting down the Sora API. One Hacker News comment is worth quoting as-is: "If you've validated workflows on V4 Pro, you may not want to suddenly be testing V4.1 Flash in production."
The version number is misleading. The company calls this "the smallest model in the new architecture family," yet gives it a ".1" version number. A model that opens an entirely new architecture family, and a model that only tweaks the post-training recipe, differ by a single decimal point in version numbering — developers have no way to judge from that how much evaluation budget to spend.
There's still no third-party audit. As of September 11, Artificial Analysis had published no measurements, and the model page still said "not yet deployed by any inference provider"; the model can't be found on LMArena, the official SWE-bench leaderboard, or LiveCodeBench either.
Put those three things together and you get one sentence: a Flash-tier model comprehensively surpassing its own flagship is a load-bearing claim being used to retire that flagship — and it currently has no external audit behind it.
There's also a practical risk: reproducibility. The old model names deepseek-v4-flash and deepseek-v4-flash-vision-exp have been taken offline, and requests are silently routed to V4.1 Flash; if a test suite still calls the old names, it may be running the new model.
The call name and the actual model name are two different levels — don't conflate them:
- The call name you put in a request is <code>deepseek-flash</code> — that's the name the company requires you to use;
- After server-side resolution, the model actually serving you is <code>deepseek-v4.1-flash</code>, and that's what the model field in the response returns;
- The official version string is written DeepSeek-V4.1-Flash.
So if you see deepseek-v4.1-flash in a response's model field, don't be confused — you didn't fill it in wrong; that's the result of the call name being resolved. The real risk is on the other end: test suites calling the old names deepseek-v4-flash or deepseek-v4-flash-vision-exp are in fact already running the new model.
6. DSH Users: This Time It's a Hard Binding — Our Readers Should Read This Section Most
DeepSeek Harness v0.1.5 shipped in step with V4.1 Flash, and the company did dedicated training optimization for this model inside DSH.
The official Harness update notes say, verbatim: "the new version is deeply integrated with DeepSeek V4.1 Flash model training; the model was specially trained and optimized across different DSH configurations" — covering three configurations: standard mode, programmatic tool calling (PTC) mode, and minimal mode.
The most valuable item is here: when using V4.1 Flash, the new DSH version can update the system prompt while preserving the existing KV Cache. For continuously running Agent tasks, that means you can change the task direction mid-flight without recomputing the entire context.
Read that alongside the architecture in section 2: CED splits reads from writes, KV cache is compressed to 890 bytes/token, and the active state at a million tokens is under 1GB — and what all that technical design lands as, in the product, is one feature: "changing the system prompt doesn't recompute the context." The technology isn't showing off; it ultimately lands as one action a user can feel — the most elegant closed loop in this piece.
Other updates: the web interface supports uploading images and PDFs; the Sidebar provides a file tree and previews (Markdown, HTML, PDF, code, images); standardized plugin extension entry points open on both sides; multi-Agent support includes parent-child two-way communication and interrupt-to-stop, the main Agent can pick the model and reasoning effort for sub-agents, plus an experimental Agent Teams (shared task list, off by default). Installation is npx @deepseek-ai/dsh web.
Distinguish the identity: DSH is DeepSeek's own open-source Agent framework and is not on the official partner list (that label goes to third-party products). But V4.1 Flash is the first model with dedicated training optimization across multiple Harness configurations, which is a much harder binding than a partner list.
7. Specs, Prices, and a Few Facts That Are Easy to Get Wrong
Specs: 1M context, 384K output cap, 2500 concurrency (V4 Pro is 500), image understanding supported (V4 Pro doesn't support it), thinking mode on by default and switchable.
Prices (yuan/million tokens, effective 12:00 on September 10):
| Tier | V4.1-Flash off-peak | V4.1-Flash peak | V4-Pro off-peak | V4-Pro peak |
|---|---|---|---|---|
| --- | --- | --- | --- | --- |
| Input cache hit | 0.02 | 0.04 | 0.15 | 0.30 |
| Input cache miss | 1.0 | 2.0 | 4.5 | 9.0 |
| Output | 4.0 | 8.0 | 13.5 | 27.0 |
Peak hours are Monday–Friday 9:00–12:00 and 14:00–18:00 Beijing time, everything else is off-peak, and off-peak prices are half of peak. Note: V4 Pro's prices did not change this time. The -60%, -33%, and -11% we reported earlier are Flash's new prices versus Flash's old prices, not versus Pro.
License: multiple tech media report an MIT license, but we could not directly verify the LICENSE file in the repository, so we can only write "multiple tech media report an MIT license; before commercial use, defer to the LICENSE file in the repository" — we don't recommend inferring unlimited commercial use from it.
Multimodality: the official line covers only vision and image understanding. Some third parties claim audio is included, with no official basis; this article doesn't adopt it.
V4.1 Pro launch timing: not announced officially; this article makes no guesses.
8. So How Should You Use It?
First, long-running Agent tasks are worth serious evaluation. Architecture, benchmarks, and community tests all point the same way: the read/write asymmetric design was built for exactly this, and the jumps on Terminal-Bench 2.1, Automation-Bench, and Codeforces aren't noise.
Second, don't rush to switch for graduate-level reasoning. On GPQA Diamond and HLE, the company's own table is right there, and the scores are lower.
Third, discount every benchmark score and wait for your own results. The gap between 134 tok/s production throughput and beta peaks, the missing third-party audit, and the reproducibility risk from silent routing are all real friction.
Finally, back to the opening: after 12:00 on September 14, V4 Pro requests are fully routed to V4.1 Flash and billed at Flash prices, until V4.1 Pro launches. Whether or not you accept "comprehensively surpasses," the timing is certain — before then, run your single most critical workflow through it once.
Data and Sourcing Notes
Data in this article was collected on September 11, 2026; the model was released on September 10; the routing takes effect at 12:00 on September 14 (Beijing time). Official figures come from the pricing-page note, the news page, the changelog, and the technical report; platform measurements come from OpenRouter's listing; community tests come from developers' public posts and Reddit discussions. Beta-period data is labeled "beta build" and isn't mixed with release-version data; competitor scores are stated as-is; "comprehensively surpasses V4 Pro" is the official line, and this article doesn't second it.