NeuroAI NEUROAINEUROAI.SITE
ESC

DeepSeek Cut API Prices 60% — and the Cheapest Part Is Memory, Not Thinking

From 10 September 2026, DeepSeek's Flash series gets off-peak cache-hit inputs at ¥0.02 per million tokens. The cut targets repeated, high-volume work — a subsidy for AI that works all day, not for AI that chats.

2026-09-09 · 643 words · NeuroAI
DeepSeek Cut API Prices 60% — and the Cheapest Part Is Memory, Not Thinking

A friend who runs an AI writing tool complained last month: the model had not changed, the user count had not grown, and the bill had doubled since January.

"What am I actually buying," he asked, "intelligence or electricity?"

An announcement from DeepSeek's open platform on 9 September 2026 may rewrite that complaint.

Key takeaways

  • Effective 12:00 Beijing time on 10 September 2026, DeepSeek adjusted Flash-series API pricing, continuing peak/off-peak time-of-use billing.
  • New off-peak rates: cache-hit input ¥0.02 per million tokens; cache-miss input ¥1; output ¥4.
  • Peak hours (Monday–Friday, 09:00–12:00 and 14:00–18:00) are double.
  • Versus the old schedule: cache-hit input down 60%, cache-miss input down about 33%, output down about 11%.
  • Context: Anthropic's Claude Fable 5.1 cut cache-read cost by 75% while holding its price tiers. Rivals are cutting in the same place, on purpose.

Why the biggest cut landed on "cache hit"

Three technical words, one very ordinary idea. A cache hit means: you are asking about something the model has just processed something similar to.

In practice that is: uploading one product manual and asking a hundred questions about it. An agent that reads your daily report every morning and summarises it. A support bot quoting the same script library thousands of times.

These workloads share three properties — repetitive, high-frequency, large-volume — and historically they were the most expensive kind, because every turn re-read the context.

Cutting that by 60% is not a discount for chat users. It is a subsidy for people who put AI to work continuously.

Output fell only 11%, which tells you the other half of the story: the part where the model actually thinks has not got much cheaper. What fell was memory, not reasoning.

Why cut again, one month after the last change?

DeepSeek had already refined its peak/off-peak rules in August 2026 and introduced uniform low weekend pricing. Moving the main inference model again a month later is not a small gesture.

The background is competitive pressure arriving from several directions at once: Zhipu released GLM-5.3 as a public download; Tencent open-sourced the 770B MoE Hunyuan Hy4 preview; ModelBest and OpenBMB open-sourced MiniCPM5-2B on 8 September, ranking first among open models under 4B parameters.

When model quality gaps narrow visibly, price becomes the simplest switching valve — and the fact that multiple labs are cutting in the same place says the industry has done the arithmetic. The volume of the future is not a human asking one question. It is a machine working for a day.

What it means for you

Be realistic first: this will probably not show up in your AI subscription immediately. API pricing is wholesale, for developers. Consumer subscriptions are retail, and there is an entire pricing strategy in between.

It will still reach you, by three routes:

  • AI features get less rationed. Limits on daily turns and document length exist to control cost. As cost falls, the limits loosen.
  • Persistent assistants become affordable. An assistant that remembers your preferences from three months ago and handles routine tasks each morning was absurdly expensive before. It is starting to make sense.
  • Small tools get cheaper. The utilities you pay a few dollars a month for have more room to cut.

The line worth remembering: model price cuts are never charity. They are paving for the next increment of usage.

What to do now

Do not top up early for "cheaper." This round has a specific effective time and time-of-use rules — check whether your tools actually run on the Flash series.

Hand AI your repetitive work. Batch document processing, fixed daily reporting, tagging thousands of items: these are the direct beneficiaries, and they cost materially less this month than last.

Figures from DeepSeek's open platform announcement and public reporting, September 2026. Confirm current rates and time-of-use rules on the official pricing page.

More in “Foundation Models” → · Back to home · Markdown version