---
title: "DeepSeek Cut API Prices 60% — and the Cheapest Part Is Memory, Not Thinking"
date: 2026-09-09
category: Foundation Models
site: NeuroAI
canonical: https://neuroai.site/a/na-deepseek-cuts-cache-prices
language: en
---

# DeepSeek Cut API Prices 60% — and the Cheapest Part Is Memory, Not Thinking

> From 10 September 2026, DeepSeek's Flash series gets off-peak cache-hit inputs at ¥0.02 per million tokens. The cut targets repeated, high-volume work — a subsidy for AI that works all day, not for AI that chats.

A friend who runs an AI writing tool complained last month: the model had not changed, the user count had not grown, and the bill had doubled since January.

"What am I actually buying," he asked, "intelligence or electricity?"

An announcement from DeepSeek's open platform on **9 September 2026** may rewrite that complaint.

## Key takeaways

- **Effective 12:00 Beijing time on 10 September 2026**, DeepSeek adjusted Flash-series API pricing, continuing peak/off-peak time-of-use billing.

- **New off-peak rates:** cache-hit input **¥0.02 per million tokens**; cache-miss input **¥1**; output **¥4**.

- **Peak hours** (Monday–Friday, 09:00–12:00 and 14:00–18:00) are double.

- **Versus the old schedule:** cache-hit input down **60%**, cache-miss input down about **33%**, output down about **11%**.

- **Context:** Anthropic's Claude Fable 5.1 cut cache-read cost by **75%** while holding its price tiers. Rivals are cutting in the same place, on purpose.

## Why the biggest cut landed on "cache hit"

Three technical words, one very ordinary idea. A cache hit means: you are asking about something the model has just processed something similar to.

In practice that is: uploading one product manual and asking a hundred questions about it. An agent that reads your daily report every morning and summarises it. A support bot quoting the same script library thousands of times.

These workloads share three properties — **repetitive, high-frequency, large-volume** — and historically they were the most expensive kind, because every turn re-read the context.

Cutting that by 60% is not a discount for chat users. It is a **subsidy for people who put AI to work continuously**.

Output fell only 11%, which tells you the other half of the story: the part where the model actually thinks has not got much cheaper. **What fell was memory, not reasoning.**

## Why cut again, one month after the last change?

DeepSeek had already refined its peak/off-peak rules in August 2026 and introduced uniform low weekend pricing. Moving the main inference model again a month later is not a small gesture.

The background is competitive pressure arriving from several directions at once: Zhipu released GLM-5.3 as a public download; Tencent open-sourced the 770B MoE Hunyuan Hy4 preview; ModelBest and OpenBMB open-sourced MiniCPM5-2B on 8 September, ranking first among open models under 4B parameters.

When model quality gaps narrow visibly, **price becomes the simplest switching valve** — and the fact that multiple labs are cutting in the same place says the industry has done the arithmetic. The volume of the future is not a human asking one question. It is a machine working for a day.

## What it means for you

Be realistic first: **this will probably not show up in your AI subscription immediately.** API pricing is wholesale, for developers. Consumer subscriptions are retail, and there is an entire pricing strategy in between.

It will still reach you, by three routes:

- **AI features get less rationed.** Limits on daily turns and document length exist to control cost. As cost falls, the limits loosen.

- **Persistent assistants become affordable.** An assistant that remembers your preferences from three months ago and handles routine tasks each morning was absurdly expensive before. It is starting to make sense.

- **Small tools get cheaper.** The utilities you pay a few dollars a month for have more room to cut.

The line worth remembering: model price cuts are never charity. They are paving for the next increment of usage.

## What to do now

**Do not top up early for "cheaper."** This round has a specific effective time and time-of-use rules — check whether your tools actually run on the Flash series.

**Hand AI your repetitive work.** Batch document processing, fixed daily reporting, tagging thousands of items: these are the direct beneficiaries, and they cost materially less this month than last.

*Figures from DeepSeek's open platform announcement and public reporting, September 2026. Confirm current rates and time-of-use rules on the official pricing page.*

---

Published by NeuroAI (https://neuroai.site/) — https://neuroai.site/a/na-deepseek-cuts-cache-prices
Free to quote with attribution and a link to the original.
