梁圣回来了,缓存价格直降60%Saint Liang Is Back: Cache Prices Drop a Straight 60%
先说结论。
Let's start with the bottom line.

号外:flash 系列调价,9 月 10 日中午十二点生效
先说结论。
北京时间 2026 年 9 月 10 日 12:00 起,DeepSeek flash 系列调整定价。新的空闲时段价格为(单位:元 / 百万 tokens):
- 输入缓存命中:0.02 元
- 输入缓存未命中:1 元
- 输出:4 元
高峰时段价格为空闲时段价格的 2 倍,即 0.04 元 / 2 元 / 8 元。
最受益的是多轮对话、长上下文、agent 工作流这类场景——它们的账单里,缓存命中那部分通常占大头,而这一档直接降了六成。
官方通知原文如下:
我们将于北京时间2026年9月10日12:00起,调整flash系列定价:空闲时段输入缓存命中单价0.02元、输入缓存未命中单价1元,输出单价4元;高峰时段价格为空闲时段价格的2倍。请合理安排您的使用。
新旧价目对比
据官方定价页现行牌价(deepseek-v4-flash,核验于 2026-09-09 凌晨)与本次公告推算,三档价格的变动如下:
| 档位(元 / 百万 tokens) | 现行空闲价 | 新空闲价 | 降幅 | 新高峰价 |
|---|---|---|---|---|
| --- | --- | --- | --- | --- |
| 输入(缓存命中) | 0.05 | 0.02 | -60% | 0.04 |
| 输入(缓存未命中) | 1.5 | 1.0 | -33% | 2.0 |
| 输出 | 4.5 | 4.0 | -11% | 8.0 |
这里必须点明一个口径问题:标题里的"直降 60%",说的是缓存命中这一档(0.05 元降到 0.02 元),不是所有档位都降 60%。 缓存未命中降三成,输出降约一成。三档降幅差别很大,照着"六折"去做预算会算错。
另一个没变的东西也值得留意:高峰仍是空闲的 2 倍,倍数关系维持原样。现行规则里,高峰时段为周一至周五 9:00–12:00、14:00–18:00,空闲价格是高峰的一半——新公告的表述与此一致,说明峰谷机制本身没有调整。
实际能省多少,取决于你的缓存命中率
降幅不是一个固定数,取决于你的请求里命中缓存的比例。举个例子(以下为按上表价格做的测算,仅用于说明结构,不代表任何实际账单):
假设某 agent 工作流每处理 100 万 tokens 的输入,其中 90% 命中缓存、10% 未命中,同时产生 5 万 tokens 输出。
- 按现行空闲价:0.9 × 0.05 + 0.1 × 1.5 + 0.05 × 4.5 ≈ 0.42 元
- 按新空闲价:0.9 × 0.02 + 0.1 × 1.0 + 0.05 × 4.0 ≈ 0.32 元
整体下降约 24%。命中率再高一些,比如 95%,整体降幅会更接近三成;反过来,如果任务每次都在喂全新的长文档、命中率很低,那你能拿到的主要是缓存未命中档 33% 的降幅。
结论很直接:这套新价目对"反复读取同一批上下文"的工作流最友好。 agent 多轮迭代、长文档反复问答、固定知识库检索、批量跑同一套 prompt 的任务,都属于这一类。而一次性、上下文每次都变的任务,获益相对有限。
一个老建议,现在更值钱了:能挪的任务挪到空闲
高峰是空闲的 2 倍,这个倍数在新价目下依然成立。也就是说,把一个原本跑在周二上午十点的批量任务挪到晚上,效果等同于再打一次五折,而且这个折扣和本次降价是可以叠加的。
按上面的例子算:同一个任务在高峰时段,现行价约 0.84 元,新价约 0.64 元;挪到空闲时段,新价约 0.32 元。降价和错峰叠在一起,才是这次调价能吃满的姿势。
具体到 DSH 生态里,有两件事本号上期报道过,正好能对上:
一是小鲸鱼余额挂件,它本身就支持按峰谷定价换算,配了令牌可以精确到每小时的 token 用量——调价生效后,拿它对着新价目看一眼实时消耗,比事后翻账单直观。二是 dsh-web 的"梁神模式"agent 预设,正是围绕 flash 系列做实测调优的(社区评测均值 98.5),两阶段锚定的设计初衷之一就是省掉首轮不必要的工具开销。这两条这里不展开,回头看上期的详细报道即可。
关于"梁圣回来了"
调价消息一出,社区里已经有人喊出"梁圣回来了"。需要说明的是,这是社区对本次调价的情绪表达与玩梗;官方通知全文只涉及价格与时段安排,未提及任何人事信息。本文不对调价原因做任何推测,也不把它与任何个人决策建立因果关系。
生效与口径提醒
- 新价格自北京时间 2026 年 9 月 10 日 12:00 起生效,在此之前仍按现行价格计费。
- 本文价格信息的来源有两处:生效时间、三档新价与峰谷倍数,来自官方通知;现行价格与峰谷时段规则,来自官方定价页(deepseek-v4-flash,核验于 2026-09-09 凌晨)。表中"降幅"与"新高峰价"为依据上述两处信息推算得出。
- 最终计费以官方通知与平台实际账单为准。若你的用量较大,建议生效后先跑一两轮小额任务对一下账单,再放量。
最后提醒一句:价格是长期变量,不是一次性红包。真要省钱,缓存命中率和错峰这两件事,比任何一次降价都更长久。
News Flash: Flash Series Repriced, Effective Noon on September 10

Let's start with the bottom line.
Starting at 12:00 on September 10, 2026, Beijing time, DeepSeek is adjusting pricing for the flash series. The new off-peak prices are (unit: yuan / million tokens):
- Input cache hit: ¥0.02
- Input cache miss: ¥1
- Output: ¥4
Peak-hour prices are 2× the off-peak prices, i.e. ¥0.04 / ¥2 / ¥8.
The biggest winners are multi-turn conversation, long-context, and agent workflow scenarios — in their bills, the cache-hit portion usually dominates, and that tier just dropped a full 60%.
The original official notice reads:
Starting at 12:00 on September 10, 2026, Beijing time, we will adjust pricing for the flash series: during off-peak hours, the unit price for input cache hits is ¥0.02, for input cache misses ¥1, and for output ¥4; peak-hour prices are 2× the off-peak prices. Please plan your usage accordingly.
Old vs. New Pricing
Based on the current published rates on the official pricing page (deepseek-v4-flash, verified in the early hours of 2026-09-09) and this announcement, the three tiers change as follows:
| Tier (yuan / million tokens) | Current off-peak price | New off-peak price | Reduction | New peak price |
|---|---|---|---|---|
| --- | --- | --- | --- | --- |
| Input (cache hit) | 0.05 | 0.02 | -60% | 0.04 |
| Input (cache miss) | 1.5 | 1.0 | -33% | 2.0 |
| Output | 4.5 | 4.0 | -11% | 8.0 |
One framing point has to be spelled out here: the "straight 60% cut" in the headline refers to the cache-hit tier (¥0.05 down to ¥0.02) — not every tier drops 60%. Cache misses fall by a third, and output by about a tenth. The three reductions differ widely, so budgeting off "40% of the old price" will get the math wrong.
Another thing that didn't change is worth noting: peak is still 2× off-peak, and the multiple holds exactly as before. Under the current rules, peak hours are Monday–Friday 9:00–12:00 and 14:00–18:00, and off-peak prices are half of peak — the new notice says the same thing, which means the peak/off-peak mechanism itself wasn't touched.
How Much You Actually Save Depends on Your Cache Hit Rate
The reduction isn't a fixed number; it depends on the share of your requests that hit the cache. Take an example (the figures below are estimates based on the table above, purely to illustrate the structure, and do not represent any actual bill):
Suppose an agent workflow processes 1 million tokens of input per run, 90% of which hit the cache and 10% miss, and produces 50,000 output tokens.
- At current off-peak prices: 0.9 × 0.05 + 0.1 × 1.5 + 0.05 × 4.5 ≈ ¥0.42
- At new off-peak prices: 0.9 × 0.02 + 0.1 × 1.0 + 0.05 × 4.0 ≈ ¥0.32
That's an overall drop of about 24%. With a higher hit rate — say 95% — the overall reduction gets closer to 30%; conversely, if a task keeps feeding in brand-new long documents with a very low hit rate, what you mainly get is the 33% cut on the cache-miss tier.
The conclusion is straightforward: this new price sheet is friendliest to workflows that re-read the same batch of context. Multi-turn agent iteration, repeated Q&A over long documents, retrieval from a fixed knowledge base, and batch runs over the same prompt set all fall into that category. One-shot tasks whose context changes every time benefit comparatively little.
An Old Piece of Advice, Now Worth More: Move What You Can to Off-Peak
Peak is 2× off-peak, and that multiple still holds under the new price sheet. Which means moving a batch job that used to run at ten on a Tuesday morning to the evening is effectively another 50% off — and that discount stacks with this price cut.
Using the example above: the same task at peak hours costs about ¥0.84 at current prices and about ¥0.64 at the new ones; moved to off-peak, the new price is about ¥0.32. Stacking the price cut with off-peak scheduling is the only way to get the full benefit of this repricing.
In the DSH ecosystem specifically, two things we covered in our last issue line up neatly here:
First, the Little Whale balance widget, which already supports peak/off-peak price conversion — with a token configured it can pin token usage down to the hour. Once the new prices take effect, glancing at real-time consumption against the new price sheet is far more intuitive than digging through the bill afterward. Second, dsh-web's "Liang God mode" agent preset, which is built around hands-on tuning for the flash series (community benchmark average 98.5); one of the design goals behind its two-stage anchoring is to eliminate unnecessary tool overhead in the first round. We won't expand on either here — see last issue's detailed coverage.
About "Saint Liang Is Back"
The moment the repricing news broke, people in the community were already shouting "Saint Liang is back." To be clear, this is the community's emotional reaction to the price cut and a running joke; the full official notice covers only pricing and time-slot arrangements and mentions no personnel matters whatsoever. This article does not speculate about the reason for the repricing, nor does it draw any causal link between it and any individual's decisions.
Effective Date and Sourcing Notes
- The new prices take effect at 12:00 on September 10, 2026, Beijing time; before that, billing continues at current prices.
- The price information in this article comes from two places: the effective time, the three new tier prices, and the peak/off-peak multiple come from the official notice; the current prices and the peak/off-peak time rules come from the official pricing page (deepseek-v4-flash, verified in the early hours of 2026-09-09). The "Reduction" and "New peak price" columns in the table are derived from those two sources.
- Final billing is subject to the official notice and the platform's actual bills. If your usage is heavy, we suggest running one or two small tasks after the change takes effect and checking them against your bill before scaling up.
One last reminder: price is a long-term variable, not a one-off red envelope. If you really want to save money, cache hit rate and off-peak scheduling will outlast any single price cut.