你的 V4 Pro 请求,明天起会被自动换成另一个模型,而且更便宜Tomorrow Your V4 Pro Requests Get Auto-Swapped to a Cheaper Model
昨天我们报了 Flash 系列的调价。当时只说了一半——那不是一次孤立的降价,而是新模型上线的配套动作。
Yesterday we reported the Flash series price change. At the time we only told half of it — that wasn't an isolated price cut, but a companion move to a new model launch.

昨天我们只报了一半
昨天我们报了 Flash 系列的调价。当时只说了一半——那不是一次孤立的降价,而是新模型上线的配套动作。
完整的三件事,都发生在 9 月 10 日:
- DeepSeek 发布 V4.1 Flash 模型(官方口径为"北京时间 9 月 10 日前后");
- 12:00 起 Flash 系列执行新价格;
- V4 Pro 的请求被自动路由到 V4.1 Flash,并按 Flash 单价计费。
第三件最反常:你一行代码都不用改,服务端就会把你原本发给 V4 Pro 的请求,交给一个更新的模型,然后按更便宜的价格收钱。升级还降价,这在行业里确实不多见。
官方怎么说的
以下是 DeepSeek 开放平台通知的原文:
DeepSeek 计划于北京时间 2026 年 9 月 10 日前后正式发布 V4.1 Flash 模型。经内部、外部多方测试,V4.1 Flash 在性能、费用、速度、总用时等各项指标上已全面超越 V4 Pro。秉持着对用户负责的态度,在 V4.1 Flash 正式上线之后、V4.1 Pro 上线之前,我们会将对 V4 Pro 的请求全部路由到 V4.1 Flash,并按 V4.1 Flash 单价计费。如您在 V4 Pro 和 V4.1 Flash 的对比测试中发现任何问题,请及时向我们反馈,感谢您的支持!
需要提醒的是,"在各项指标上全面超越 V4 Pro"是官方说法,我们原样转述,不做二次背书。官方至今没有发布任何 V4.1 Flash 的基准分数表——这件事本身也是个信息。
自动路由的边界,别理解错了
先把容易误会的地方划清楚:
- 这是一段过渡期政策,只在"V4.1 Flash 上线之后、V4.1 Pro 上线之前"成立。V4.1 Pro 发布之后怎么办,官方没有说明。
- 用户无需任何代码改动,路由在服务端完成。也就是说,你调用时写的模型名还是 V4 Pro,实际跑的是 V4.1 Flash,账单按 Flash 算。
- 官方给的理由是"秉持着对用户负责的态度"——措辞本身很克制,没有任何营销语。
- 这项政策没有出现在任何公开文档页面上,来源为登录态开放平台通知与官方交流群,由掘金、上海证券报、IT之家、科创板日报等转引。
换句话说,这是一个"先做后说"的动作:钱已经在省了,说明文档还没跟上。
这笔账,才是最有说服力的部分
同一条请求,从 V4 Pro 切到 V4.1 Flash,价格差了多少?
| 档位(元 / 百万 tokens) | V4 Pro 现行空闲价 | V4.1 Flash 新空闲价 | 降幅 |
|---|---|---|---|
| --- | --- | --- | --- |
| 输入(缓存命中) | 0.15 | 0.02 | 约 -87% |
| 输入(缓存未命中) | 4.5 | 1.0 | 约 -78% |
| 输出 | 13.5 | 4.0 | 约 -70% |
高峰时段按各自空闲价 ×2 计算,倍数关系不变(高峰时段为周一至周五 9:00–12:00、14:00–18:00)。V4 Pro 现行价来自官方定价页,V4.1 Flash 新价来自官方调价通知,降幅为依据两者推算。
把三档合成一个具体场景(以下为按上表价格做的测算,不代表任何实际账单):
一个 agent 任务,输入 100 万 tokens,其中 80% 命中缓存,输出 5 万 tokens。
- 走 V4 Pro,现行空闲价:0.8 × 0.15 + 0.2 × 4.5 + 0.05 × 13.5 ≈ 1.70 元
- 走 V4.1 Flash,新空闲价:0.8 × 0.02 + 0.2 × 1.0 + 0.05 × 4.0 ≈ 0.42 元
同一个任务,账单少掉约四分之三。这就是"模型变强"和"钱变少"能同时成立的原因:便宜的那部分不是来自挤出利润,而是来自两代模型本身的价差——V4 Pro 的输出单价一直是 13.5 元,Flash 一直是几块钱的量级,路由一把把 Pro 的活按 Flash 的价结了。
顺带把昨天那张 Flash 调价表也贴在这里,方便对照(单位:元 / 百万 tokens):
| 档位 | 现行空闲价 | 新空闲价 | 降幅 | 新高峰价 |
|---|---|---|---|---|
| --- | --- | --- | --- | --- |
| 输入(缓存命中) | 0.05 | 0.02 | -60% | 0.04 |
| 输入(缓存未命中) | 1.5 | 1.0 | -33% | 2.0 |
| 输出 | 4.5 | 4.0 | -11% | 8.0 |
注意两件事:一是调价表里的 -60% 只指缓存命中档;二是这张表讲的是 Flash 用户自己的变化,而上面那张 V4 Pro 对照表,才是路由政策真正的冲击面。
社区实测:有惊喜,也有冷水
先说正面的(均为社区用户在各自场景下的实测,非官方口径):
- 3D"S 形路线停车"任务:V4.1 约 25.4 秒零碰撞完成;V4 用时 11 分 27 秒,且主画面上下颠倒。
- 小店任务(Node.js + 购物车 + 优惠码 + 幂等下单):V4.1 用时 1 分 41 秒,V4 用时 5 分 54 秒。
- 多模态幻觉:有博主发了一张西装照,模型答"条纹西装",放大确认属实。
- 速度:社区普遍反馈比 V4-Flash 明显更快,但流传的数字散布很宽(200 到 600 tok/s 的说法都有,284 / 355 / 420 / 507 各有出处),且都来自不同人的不同场景,无法互相印证,也不代表官方性能。
再说反面的,这部分同样重要:
- 有开发者做了 14 组任务实测,标题就叫"居然比上代贵":同款小游戏,V4.1 端到端 34.5 分钟,V4 是 30.5 分钟。原因很朴素——吐字快不等于交活快,省下来的生成时间被工具调用与验证环节吃掉了(等工具 18.7 分钟 vs 6.6 分钟)。
- 第一财经报道,有开发者反馈"5 分钟 10 块钱""瞬间几十块没了",对"成本更低"这个结论存疑,推测更快的生成速度意味着单位时间吞掉更多 token。
- 过度思考的问题仍在:有人烧掉 4M tokens 还没跑完一个 Frogger 游戏;七奇迹基准跑了 2 小时、花费 2.6 美元。
- 自认知混乱:有人问"你是谁",它回答"Claude"。
- 内测版局限:暂不支持视频,只测了图片;部分第三方适配还没跟上。
这两面并不矛盾。模型单次反应变快,和一次任务的总成本变低,是两件不同的事——如果你的瓶颈在等模型吐字,新模型是净收益;如果瓶颈在工具链往返和反复验证,那吐字快反而可能让单位时间烧得更多。
DSH 用户该关注什么
生态侧目前有两个确定信息:网易有道 LobsterAI 在 9 月 8 日宣布率先集成 V4.1 Flash 内测版,是这轮里动作最快的;尚未查到 DSH 官方针对 v4.1-flash 的专属适配公告。
但有一件事值得 DSH 用户提前留意:在 V4-native 协议下,DSH 默认使用 deepseek-v4-flash 作为高吞吐通道,遇到难题自动升级到 deepseek-v4-pro。V4.1 Flash 上线、且 Pro 请求被服务端路由之后,这条默认通道和升级路径大概率会发生变化——具体怎么变,取决于后续的模型标识与协议调整。
务实的建议:等 9 月 10 日之后跑一轮自己的真实任务,对比切换前后的耗时与账单,再决定要不要动配置。别人的基准分和你的工作流,往往是两回事。
口径与可信度提醒
- 截至 9 月 9 日夜间,DeepSeek 官方公开文档中没有任何 V4.1 Flash 条目:api-docs 价格页仍只有 v4-flash、v4-pro、v4-flash-vision-exp 三个模型,更新日志最新一条为 8 月 21 日,官网首页头条仍是 V4-Pro。
- 因此本文所有发布与路由信息,均来自 DeepSeek 开放平台通知与官方交流群,经第三方媒体转引,并非取自官方公开文档页面。请以官方最终公告与平台实际行为为准。
- 发布时间为官方口径的"9 月 10 日前后",不是确定时点;确定时点只有一个:调价自 9 月 10 日 12:00 起生效。
- V4.1 Flash 的上下文长度、输出上限、并发限制、是否支持思考模式,官方均未说明,本文不做任何推测,也不沿用上一代参数套用。
- 社区实测数据均标注了测试者与场景,属个人环境下的观察,不代表官方性能,也不构成任何选型结论。
最后一句:这轮最值得记住的不是降了多少钱,而是"服务端可以替你换模型"这件事本身。模型迭代速度已经快到,靠用户自己改代码是跟不上的。
Yesterday We Only Told Half the Story

Yesterday we reported the Flash series price change. At the time we only told half of it — that wasn't an isolated price cut, but a companion move to a new model launch.
All three things happened on September 10:
- DeepSeek released the V4.1 Flash model (officially worded as "around September 10, Beijing time");
- Starting at 12:00, the Flash series moved to new pricing;
- V4 Pro requests are automatically routed to V4.1 Flash and billed at Flash rates.
The third is the most unusual: without changing a single line of code, the server hands your V4 Pro request to a newer model and then charges you the cheaper price. An upgrade that also cuts the price — that's genuinely rare in this industry.
What the Official Notice Says
Here is the original text of the DeepSeek open platform notice:
DeepSeek plans to officially release the V4.1 Flash model around September 10, 2026, Beijing time. After extensive internal and external testing, V4.1 Flash has comprehensively surpassed V4 Pro across every metric, including performance, cost, speed, and total elapsed time. In the spirit of an attitude of responsibility toward users, after V4.1 Flash officially goes live and before V4.1 Pro launches, we will route all requests to V4 Pro to V4.1 Flash and bill them at V4.1 Flash rates. If you find any problems in comparative testing between V4 Pro and V4.1 Flash, please report them to us promptly. Thank you for your support!
One caveat: "comprehensively surpassed V4 Pro across every metric" is the official wording, which we relay verbatim without seconding it. To date, the company has published no benchmark score table for V4.1 Flash at all — and that absence is itself information.
The Boundaries of Auto-Routing: Don't Misread Them
Let's draw the lines where confusion is most likely:
- This is an interim policy, valid only "after V4.1 Flash goes live and before V4.1 Pro launches." What happens after V4.1 Pro ships, the company has not said.
- Users need no code changes at all; routing happens server-side. In other words, the model name you call is still V4 Pro, what actually runs is V4.1 Flash, and the bill is calculated at Flash rates.
- The stated reason is "an attitude of responsibility toward users" — the wording itself is notably restrained, with no marketing language whatsoever.
- This policy appears on no public documentation page; it comes from a logged-in open platform notice and an official community group, relayed by Juejin, Shanghai Securities News, ITHome, STAR Market Daily, and others.
In other words, this is an "act first, document later" move: the savings are already live, the docs haven't caught up.
The Math Is the Most Persuasive Part
For the same request, how much does the price differ when it moves from V4 Pro to V4.1 Flash?
| Tier (yuan / million tokens) | V4 Pro current off-peak price | V4.1 Flash new off-peak price | Reduction |
|---|---|---|---|
| --- | --- | --- | --- |
| Input (cache hit) | 0.15 | 0.02 | ~-87% |
| Input (cache miss) | 4.5 | 1.0 | ~-78% |
| Output | 13.5 | 4.0 | ~-70% |
Peak-hour prices are each tier's off-peak price ×2, and the multiple is unchanged (peak hours are Monday–Friday 9:00–12:00 and 14:00–18:00). V4 Pro's current prices come from the official pricing page; V4.1 Flash's new prices come from the official repricing notice; the reductions are derived from the two.
Combine the three tiers into one concrete scenario (the figures below are estimates based on the table above and do not represent any actual bill):
An agent task with 1 million input tokens, 80% of which hit the cache, and 50,000 output tokens.
- On V4 Pro, current off-peak price: 0.8 × 0.15 + 0.2 × 4.5 + 0.05 × 13.5 ≈ ¥1.70
- On V4.1 Flash, new off-peak price: 0.8 × 0.02 + 0.2 × 1.0 + 0.05 × 4.0 ≈ ¥0.42
For the same task, the bill drops by about three quarters. This is why "a stronger model" and "less money" can both be true at once: the savings don't come from squeezed margins, but from the price gap between the two model generations — V4 Pro's output price has always been ¥13.5, and Flash has always been in the single-digit yuan range. Routing simply settles Pro's work at Flash's rates.
While we're at it, here's yesterday's Flash repricing table again for comparison (unit: yuan / million tokens):
| Tier | Current off-peak price | New off-peak price | Reduction | New peak price |
|---|---|---|---|---|
| --- | --- | --- | --- | --- |
| Input (cache hit) | 0.05 | 0.02 | -60% | 0.04 |
| Input (cache miss) | 1.5 | 1.0 | -33% | 2.0 |
| Output | 4.5 | 4.0 | -11% | 8.0 |
Two things to note: first, the -60% in the repricing table refers only to the cache-hit tier; second, that table describes what changes for Flash users themselves, whereas the V4 Pro comparison above is the real blast radius of the routing policy.
Community Tests: Some Pleasant Surprises, Some Cold Water
First the positives (all from community users testing in their own scenarios, not official figures):
- 3D "S-shaped path parking" task: V4.1 finished in about 25.4 seconds with zero collisions; V4 took 11 minutes 27 seconds and rendered the main view upside down.
- Small shop task (Node.js + shopping cart + coupon codes + idempotent checkout): V4.1 took 1 minute 41 seconds, V4 took 5 minutes 54 seconds.
- Multimodal hallucination: a blogger posted a photo of a suit, and the model answered "pinstripe suit" — zooming in confirmed it was correct.
- Speed: the community broadly reports it's noticeably faster than V4-Flash, but the circulating numbers are all over the place (claims range from 200 to 600 tok/s, with 284 / 355 / 420 / 507 each having their own source), and they all come from different people in different scenarios, so they can't corroborate each other and don't represent official performance.
Now the negatives — this part matters just as much:
- One developer ran 14 sets of task tests under the headline "it's actually more expensive than the previous generation": for the same small game, V4.1 took 34.5 minutes end to end, versus 30.5 minutes for V4. The reason is plain — fast token output doesn't mean fast delivery; the generation time saved gets eaten by tool calls and verification (18.7 minutes waiting on tools vs 6.6 minutes).
- Yicai reported that developers said "¥10 in 5 minutes" and "dozens of yuan gone in an instant," casting doubt on the "lower cost" conclusion and speculating that faster generation means burning more tokens per unit of time.
- Overthinking is still a problem: someone burned 4M tokens without finishing a single Frogger game; the Seven Wonders benchmark ran 2 hours and cost $2.60.
- Self-identity confusion: asked "who are you," it answered "Claude."
- Beta limitations: video isn't supported yet — only images were tested; some third-party integrations haven't caught up.
These two sides don't contradict each other. A model reacting faster on a single turn and a task costing less overall are two different things — if your bottleneck is waiting for tokens, the new model is a net win; if your bottleneck is tool-chain round trips and repeated verification, then fast token output may actually burn more per unit of time.
What DSH Users Should Watch
On the ecosystem side there are two confirmed data points so far: NetEase Youdao's LobsterAI announced on September 8 that it was the first to integrate the V4.1 Flash beta, the fastest move in this cycle; we have not found any DSH official announcement of dedicated support for v4.1-flash.
But one thing is worth DSH users' early attention: under the V4-native protocol, DSH uses deepseek-v4-flash by default as its high-throughput channel and automatically escalates to deepseek-v4-pro when it hits a hard problem. Once V4.1 Flash ships and Pro requests are routed server-side, that default channel and escalation path will most likely change — exactly how depends on subsequent model identifiers and protocol adjustments.
The practical advice: after September 10, run a round of your own real tasks, compare time and bills before and after the switch, and only then decide whether to touch your configuration. Someone else's benchmark scores and your workflow are usually two different things.
Notes on Sourcing and Reliability
- As of the night of September 9, there was no V4.1 Flash entry anywhere in DeepSeek's official public documentation: the api-docs pricing page still listed only three models — v4-flash, v4-pro, and v4-flash-vision-exp; the newest changelog entry was dated August 21; and the homepage headline was still V4-Pro.
- Therefore all release and routing information in this article comes from DeepSeek open platform notices and the official community group, relayed by third-party media, not from official public documentation pages. Please treat the official final announcement and the platform's actual behavior as authoritative.
- The release timing is the official phrasing "around September 10," not a fixed moment; there is only one fixed moment: the repricing takes effect at 12:00 on September 10.
- V4.1 Flash's context length, output cap, concurrency limit, and whether it supports a thinking mode are all unstated by the company. This article makes no guesses and does not carry over the previous generation's specs.
- All community test data names the tester and the scenario; it reflects observations in individual environments, does not represent official performance, and does not constitute any model-selection conclusion.
One last point: the most memorable thing about this round isn't how much the price dropped, but the fact that "the server can swap your model for you." Model iteration is now moving so fast that users changing their own code can't keep up.