---
title: "StepFun's open-weight \"Flash\" model is built for agents, not chat"
date: 2026-10-04
category: Foundation Models
site: NeuroAI
canonical: https://neuroai.site/a/na-model-stepfun-step37-flash-agent
language: en
---

# StepFun's open-weight "Flash" model is built for agents, not chat

> StepFun open-sourced Step 3.7 Flash, a 196B MoE model tuned for tool-calling and long agent runs — a bet that the model, not just the wrapper, must change for production AI.

Chatbots answer questions. Agents do the work — and most models still fumble the doing. Tool calls break halfway. Long tasks drift off script. The demo works; production doesn't. A Shanghai startup is betting that the model itself, not just the wrapper, has to change.

## The model: Step 3.7 Flash

On May 29, 2026, StepFun (阶跃星辰) released and open-sourced Step 3.7 Flash under the Apache 2.0 license. It is a sparse Mixture-of-Experts (MoE, 混合专家) model: about 196B parameters plus a 1.8B vision encoder, with roughly 11B parameters active per token.

Key specs reported by the company:

- **256K context window**

- Up to **400 tokens per second** generation speed

- **Three reasoning levels** (low / medium / high) so developers trade speed and cost against capability

- Native multimodal input — images directly, and video on the hosted API

The design goal is explicit: a model cheap and reliable enough to call hundreds of times inside a long agent loop, where a slow frontier model would be uneconomic.

## Why "Flash" means agent-grade

StepFun frames Flash not as a lighter chatbot but as infrastructure for production agents (智能体). The optimizations target the failure modes that kill real deployments:

- **Stable tool calls.** The model is tuned to keep calling APIs, browsers, terminals, and Office tools correctly across many turns without drifting.

- **Proactive verification.** When information is uncertain, it can launch web and visual searches to cross-check before acting.

- **Structured output.** It turns screenshots, charts, and documents into structured results and executable tasks.

On agent benchmarks the company cites:

- **Toolathlon:** 49.5%

- **ClawEval-1.1** (daily autonomous tasks): 67.1%

- **GDPval** (44 professions): 45.8%

- **τ²-bench Telecom:** over 98% pass rate at all three reasoning levels

These are vendor-reported; we did not re-run them.

## Plays well with existing agent stacks

Step 3.7 Flash is built to drop into tools developers already use. StepFun says it is compatible with Claude Code, OpenClaw, Hermes Agent, KiloCode, RooCode, and OpenCode, and supports the MCP and Skills protocols. That means a team can swap it in for a closed model behind the same agent harness, on cloud or local hardware.

Its predecessor, Step 3.5 Flash (February 2026), set the stage: StepFun says it hit number one on OpenRouter's OpenClaw monthly call-volume chart within a month of release and passed 300,000 downloads on Hugging Face in the same window.

## The company behind it

StepFun was founded in 2023. CEO Jiang Daxin (姜大昕) is a veteran researcher; chairman Yin Qi (印奇) co-founded the computer-vision unicorn Megvii (旷视). Under Yin, the strategy shifted to "AI + terminals" — models pre-installed in OPPO and Honor devices with reported total installs above 42 million. The company is also preparing a Hong Kong listing.

For context on pricing, StepFun's earlier 321B Step 3 launched with promotional API rates around 1.5 yuan input and 4 yuan output per million tokens — roughly **US$0.21** and **US$0.56** — a hint at the cost tier Step 3.7 Flash's hosted API sits near.

## What comes next

StepFun previewed Step 5 on September 20, 2026: a 600B MoE with 27B active parameters, a 1M-token context, and vision input, with open weights promised for mid-October. The direction is clear — bigger windows, agent focus, open weights.

## How to think about Flash-class models

Step 3.7 Flash belongs to a broader "Flash" trend: fast, cheap, open-weight models designed to be called thousands of times inside agents rather than queried once in a chat box. Google's Gemini 3.5 Flash and Anthropic's Claude Haiku occupy the same niche on the closed side. The open variants matter because they let teams own the model and avoid per-call lock-in.

The trade is simple. You give up some peak reasoning for speed and cost, then compensate with longer, tool-rich loops. For tasks like "scan this inbox, file the receipts, and draft replies," that exchange usually wins — and it is why StepFun and its peers are aiming Flash models squarely at production agent infrastructure rather than the chat demo.

## Honest limitations

- All benchmark figures come from StepFun; we did not independently verify them, and agent benchmarks vary wildly by harness and prompt.

- "Production-grade" is a claim, not a certificate; real enterprise reliability depends on the surrounding system, not the model alone.

- The open-weight release handles images; video understanding is a hosted-API feature only, so self-hosted users get a narrower model.

- We did not test latency, cost, or tool-call stability ourselves.

## What readers can do now

- **Self-host it:** pull Step 3.7 Flash from Hugging Face and serve it with vLLM or Ollama to test agent workflows privately.

- **Swap it into your stack:** point an MCP-compatible coding agent (Claude Code-style) at a local Step 3.7 Flash endpoint and compare tool-call reliability against your current model.

- **Track the open-weights wave:** watch Step 5's October weight release if you need a 1M-context open agent model.

---

Published by NeuroAI (https://neuroai.site/) — https://neuroai.site/a/na-model-stepfun-step37-flash-agent
Free to quote with attribution and a link to the original.
