A model that activates only a fraction of its 552 billion parameters per token is now free to download, fine-tune and ship inside your own stack. On 10 September 2026, DeepSeek (深度求索) published the weights of V4.1-Flash on Hugging Face under an MIT licence. The headline number is not the size — it is the architecture that lets a model that big stay cheap to run.
The release that actually happened
- DeepSeek-V4.1-Flash, released 10 September 2026, served on the API as
deepseek-flash. - Open weights (开源权重) on Hugging Face under an MIT licence, with a technical report published alongside.
- A 552B-parameter transformer backbone plus a 196B "Engram" conditional-memory module — about 748B parameters in total (roughly 763B if you count a separate speculative-decoding module).
- It activates only about 8B parameters per token in prefill and 16B per token in decode.
- A 1,048,576-token context (1M), with up to 384K output tokens.
- Multimodal: it reads images and text and writes text — vision is native to the base design, not bolted on.
- Trained from scratch on 45 trillion tokens.
Why "read cheaper than write" is the clever bit
Most large language models are decoder-only: the same uniform stack processes your input and generates the reply, charged at the same rate. V4.1-Flash splits the job.
It uses a Causal Encoder-Decoder (CED): a 20-layer encoder that reads the prompt, then a 20-layer decoder that produces the answer. Combined with Compressed Sparse Attention 2 (CSA2) and an FP4 KV cache, the global cache shrinks to about 890 bytes per token — roughly a quarter of the earlier V4-Flash and about 1/437 of the original V1. A separate Engram module is a 196B lookup table the network consults sparsely: think of it as a colleague who remembers where something is written, versus one who works the answer out from scratch.
The payoff is in long documents and agent workloads, where a model spends far more tokens reading context and calling tools than it does replying. Reading gets cheap; the huge parameter count becomes a memory aid rather than a per-token tax.
The benchmarks DeepSeek published
On its own model card, DeepSeek reports:
- Terminal-Bench 2.1: 90.6
- DeepSWE v1.1: 74.2
- CyberGym: 88.1
- GPQA Diamond: 90.9
- Humanity's Last Exam: 36.8
Pricing on the platform is US$0.15 / US$0.60 per million input / output tokens off-peak (cache reads at US$0.003), with peak hours at double that. These are DeepSeek's vendor numbers, and independent reproduction has not yet landed.
V4.1-Pro, and a walked-back retirement
A larger V4.1-Pro is the coming variant. At launch, DeepSeek said it would route all deepseek-v4-pro API traffic to Flash from 14 September and bill at Flash rates. Within a day it reversed course, confirming that V4-Pro continues to be served. The signal is clear anyway: DeepSeek is now happy to make a cheaper, smaller-activation model its default flagship.
The open-weight playbook behind it
DeepSeek has made a habit of giving frontier-class weights away for free — V3 at the end of 2024, R1 in early 2025 — and then undercutting the market on API price. V4.1-Flash follows the same script: hand developers the model, charge for the managed endpoint. The ripple goes past China. DeepSeek was reported to be preparing a STAR Market (科创板) IPO, and its price cuts have repeatedly moved the Hong Kong shares of rivals such as MiniMax and Z.ai on release days. The strategic point is ecosystem continuity, not any single benchmark number.
What readers can do now
- If you build: self-host Flash for long-context agent and coding work — the MIT licence allows commercial use, and the thin KV cache is what makes the 1M window affordable.
- If you use the API: the default
deepseek-flashis now the V4.1 generation; tunereasoning_effort(an integer from 1 to 100) to trade cost against accuracy per request. - If you track the field: watch whether V4.1-Pro actually ships and whether third parties confirm the agentic leads — the architecture claims are more interesting than the parameter count.
Honest limitations
Sources: the Hugging Face model card deepseek-ai/DeepSeek-V4.1-Flash (primary), and aiwiki.ai, llmreference.com, groundtruth.day, newdecoded.com and ai-tldr.dev (all citing that card: 10 Sep 2026 release, 552B+196B parameters, MIT licence, 1M context, 8B/16B activation, 510.3 GB of published weights across 48 shards, CED/CSA2 details, V4-Pro reversal). The three circulating totals — 552B, 748B, 763B — describe different parts of the checkpoint (backbone; backbone plus Engram; backbone plus Engram plus the speculative module). Every benchmark figure is DeepSeek's own vendor number, not independently reproduced. Pricing is stated in USD, so no RMB conversion applies. This is not investment advice.
