A model with 321 billion parameters usually lives behind a paid API, on someone else's servers, where you cannot see how it works. StepFun (阶跃星辰) did the opposite: it put the weights on Hugging Face and told anyone with enough GPUs to take it home.
The release is Step 3, and the bet behind it is that the future of useful AI is not the biggest model you can rent, but the most efficient one you can own.
The company behind the weights
StepFun was founded in April 2023 in Shanghai's Xuhui district by Jiang Daxin (姜大昕), a former global vice-president at Microsoft, and has positioned itself as a multimodal specialist rather than a general chat competitor. Its model family — text, vision, video, speech — has been released in waves, and open-sourcing is now a stated strategy, not a one-off.
Step 3 is the headline, but it sits inside a broader open catalogue:
- Step-Video-T2V, a 30-billion-parameter text-to-video model open-sourced in February 2025 with Geely.
- Step-Audio 2 mini, an end-to-end speech model open-sourced on 1 September 2025.
- Step-2, the company's trillion-parameter MoE large model (大模型) that reached a public release in February 2025.
What Step 3 actually is
Step 3 is a multimodal mixture-of-experts (MoE) model. The headline spec is its economy of motion:
- 321 billion total parameters, but only 38 billion activated per token.
- A 64,000-token context window.
- Released 31 July 2025 as open weights under Apache 2.0, downloadable from Hugging Face and ModelScope.
The trick is in the architecture. StepFun built a custom attention mechanism called MFA (Multi-matrix Factorization Attention) that cuts the KV-cache overhead and compute cost of long contexts, and a system-level split called AFD (Attention-FFN Disaggregation) that decouples attention from feed-forward computation so the two can be scheduled on different hardware paths. The company even open-sourced StepMesh, a communication library for this setup, so the claimed throughput is not locked inside its own cloud.
Vision without drowning in tokens
Multimodal models usually pay a tax: every image floods the context with visual tokens. Step 3 uses a 5-billion-parameter vision encoder plus a two-layer 2D convolution that downsamples visual features to one-sixteenth of their original token count. The stated goal is to keep image understanding cheap enough to deploy, not just to demo.
On public benchmarks tracked by model registries, Step 3 posts numbers that put it in conversation with other open frontiers: roughly 82.9 on AIME 2025 (math), 73.0 on GPQA-Diamond (scientific reasoning), 67.1 on LiveCodeBench (coding) and 74.2 on MMMU (multimodal understanding).
Why open weights, why now
The strategic logic is straightforward. A model you can self-host becomes infrastructure other people build on — and StepFun is explicit that Step 3 was designed "for the inference era," where running cost per token, not raw parameter count, decides adoption. Because only 38 billion parameters fire per token and the design targets mainstream accelerators, the company says it can run on a rack of eight 48 GB-class GPUs, which is squarely within reach of many enterprise teams.
That places StepFun alongside DeepSeek and Alibaba's Qwen in the camp of Chinese labs shipping open weights at frontier scale — a different posture from labs that keep their best models behind APIs.
What readers can do now
- If you build with open models, benchmark Step 3 against Qwen and DeepSeek on your own vision-and-reasoning tasks; its MFA/AFD design is meant to win on cost-per-token, which matters more than leaderboard peaks in production.
- If you care about sovereignty and auditability, download the Apache-2.0 weights and inspect the architecture yourself — open weights, not just open APIs, are what let you verify and customize.
- If you follow the open-model race, watch StepFun's cadence: a video model, a speech model and now a flagship multimodal model all open in a single year signals a portfolio strategy, not a one-shot.
Honest limitations
Core facts (StepFun founded April 2023 in Shanghai by Jiang Daxin; Step 3 released 31 July 2025 as open weights, Apache 2.0; 321B total / 38B active parameters; 64K context; MFA and AFD architecture plus open-sourced StepMesh; 5B vision encoder with 1/16 visual-token downsampling; benchmark figures ~82.9 AIME2025 / 73.0 GPQA-Diamond / 67.1 LiveCodeBench / 74.2 MMMU; Step-Video-T2V and Step-Audio 2 mini open-sourced in 2025; Step-2 trillion-parameter MoE) are sourced from StepFun's official Hugging Face model card (stepfun-ai/step3) and platform disclosures, IT之家 (ithome.com) and the DataLearnerAI model registry. The architecture and parameter counts are vendor-published and reproducible from the released weights; benchmark scores are vendor or third-party registry figures, not the result of an independent head-to-head harness. Step 3 is a text-and-image multimodal model, not a text-to-video generator despite some registries mislabeling its output modality. No RMB figures appear in this article, so no currency conversion applies. Analysis is current to 29 September 2026.
