A model that can both snap out a quick reply and quietly plan a ten-step task used to require two separate deployments. DeepSeek V3.1, released on 21 August 2025, collapsed that into one.
The company called it "our first step toward the agent era (迈向智能体时代的第一步)" — and Bloomberg, reporting the launch the same day, framed it as DeepSeek keeping pace while the industry waits for its next flagship. The practical change is small to describe and large to feel: a single model now switches between fast answers and deep reasoning on demand.
Two modes, one weight set
DeepSeek V3.1 keeps the same Mixture-of-Experts (混合专家) backbone as V3 — about 671 billion total parameters with roughly 37 billion active per token. What changed is the inference design:
- Thinking mode (deepseek-reasoner) — chain-of-thought reasoning for hard problems
- Non-thinking mode (deepseek-chat) — instant responses for ordinary queries
Users toggle it through the "DeepThink" button in the app and web UI; developers pick the behaviour at the API level. The previous generation had split this across the V3 chat model and the R1 reasoning model. V3.1 folds both into one.
Both modes share a 128K-token context window, up from the shorter windows of earlier versions — enough to hold a long contract, a codebase, or a multi-document research session in one pass.
Stronger agent skills
The headline capability is not raw scores but agent readiness. DeepSeek says post-training boosted tool use and multi-step agent tasks, and reports gains on:
- SWE-bench Verified — 66.0 for the agent coding benchmark
- Terminal-Bench — improved terminal/tool operation
- BrowseComp and complex search — stronger multi-step retrieval
On efficiency, V3.1-Think reaches answers faster than the older R1-0528 while producing 20–50% fewer output tokens at comparable task quality. For anyone paying per token, that compression is the real feature: the model thinks harder in fewer words.
Open weights, domestic-chip tailoring
Like its siblings, V3.1 ships as open weights under an MIT licence on Hugging Face (deepseek-ai/DeepSeek-V3.1) and ModelScope, base and post-trained versions alike. The base model added 840 billion tokens of continued pre-training on top of V3, much of it aimed at extending long-context behaviour.
One quiet but strategic detail: V3.1 uses a UE8M0 FP8 precision format that DeepSeek says was designed for next-generation domestic AI chips. In other words, the model was tuned to run efficiently on Chinese accelerators, not just Nvidia hardware — a direct response to export controls.
The long-context detail
Part of the 128K window is real engineering, not just a config bump. DeepSeek extended long-context training in phases — a 32K phase grown to about 630 billion tokens and a 128K phase of roughly 209 billion tokens on top of the base. The point is stability: the model is less likely to lose the thread halfway through a long document or a long agent trace, which is exactly where earlier long-context models tended to degrade.
API and compatibility
The API was upgraded in step:
deepseek-chatnow maps to non-thinking mode;deepseek-reasonerto thinking mode- Anthropic API format supported, so Claude Code-style tooling can point at DeepSeek
- Strict function calling in beta, enforcing that tool outputs match the declared schema
- Pricing was revised from 6 September 2025, ending off-peak discounts
Why it matters
Agentic AI is the industry's next wedge: systems that don't just answer but act — booking, coding, querying, filing. DeepSeek's bet is that the model layer should make that native, not require glue code and a separate reasoner. A 128K window plus strict tool-calling plus open weights is a credible developer package.
Honest limitations
- SWE-bench Verified 66.0 and the 20–50% token reductions are DeepSeek's own reported figures; independent reproduction was not re-verified here.
- "First step toward the agent era" is DeepSeek's marketing framing; the model is not a full autonomous agent by itself, only the engine for one.
- The 671B/37B active count follows the established V3 architecture; DeepSeek did not publish a fresh full parameter audit for V3.1 specifically.
- UE8M0 FP8 support depends on hardware and software stacks that, for next-gen domestic chips, were described as "upcoming" at launch; real-world availability varies by vendor.
- This article covers V3.1, not later V3.x revisions that may have shipped after August 2025.
What readers can do now
- Flip the switch — in the DeepSeek app or API, toggle DeepThink on a hard task and off for a quick one, and compare token cost; the savings on reasoning-heavy work are the point.
- Wire it to your tools — because the API speaks the Anthropic format and supports strict function calling, point an existing Claude Code-style agent at DeepSeek V3.1 to test a lower-cost reasoning backend.
- Self-host the open weights — if you run on domestic accelerators, the UE8M0 FP8 design is worth a benchmark on your own hardware before committing to a closed API.
