A four-year-old can watch you pour juice once and pour it themselves the next day. Most AI assistants still need to be told, step by step, what "pour" means.
That gap — between understanding language and acting in the world — is the line a generation of Chinese AI builders is now trying to cross. At the 2026 World Artificial Intelligence Conference (WAIC) in Shanghai, one of them put the shift into plain words.
The founder and the moment
Yin Qi (印奇) chairs StepFun (阶跃星辰), a Shanghai large-model (大模型) lab, and Qianli Technology (千里科技), an automotive-AI company. In a main-forum keynote at WAIC 2026, titled "When AI Agents Enter the Physical World," he argued that the industry has reached an inflection point — and that the next decade belongs to AI agents (智能体) that perceive, decide and act, not to chatbots that answer.
His framing came from a 15-year vantage point. "AI entrepreneurship has evolved from a niche track into a major global consensus," he told the audience, looking back on a career that began when AI was a backwater and now sits at the center of geopolitics and capital markets alike.
The threshold he says we just crossed
The core claim is concrete. According to Yin, model capabilities in 2026 crossed "a critical threshold" — AI has advanced, in his words, "from executing tasks lasting only a few seconds to working independently for tens of hours."
That is a meaningful distinction. A chatbot that answers one question is a parlour trick. A system that can hold a goal for hours, call tools, fix its own mistakes and finish a multi-step job is a different category of product. Whether the industry has truly arrived there is debatable; Yin's point is that the direction is no longer in doubt, only the speed.
He also named a new yardstick: after natural language, "programming is becoming a crucial benchmark for measuring AI capability." In other words, the field is beginning to judge models not by how well they talk, but by whether they can write and run the code that changes things.
From chatbot to "smallest productive unit"
Yin's sharper prediction is about form. He expects AI agents (智能体) to evolve "from chatbots into the smallest productive units capable of perception, decision-making, and task execution." Engineers, designers and researchers, he said, will each get a dedicated agent — so that "a single individual possesses the capabilities of an entire team."
He laid out three structural shifts the agent wave supposedly drives:
- A new system. An "Agentic OS" connects models with data, tools, interfaces and devices, and decides how far any agent can actually go.
- A new body. Computers, phones, cars and humanoid (人形机器人) robots become different "bodies" for the same agent, moving intelligence across terminals.
- A new network. An "A2A" (agent-to-agent) network where agents hold their own identity, ability and credit, find partners and complete transactions.
It is a clean story: one brain, many bodies, a machine economy underneath.
The part he did not skip
Yin did not pitch this as pure upside. Agents entering the real world, he warned, bring "order restructuring," not just capability leaps. The hard questions — who the agent acts for, who answers for what it does, how identity stays trustworthy and behavior stays traceable — are unresolved, and he argued the industry has to answer them together.
His closing line was the one worth keeping: "The future is not a world where machines replace humans, but a world where humans and intelligence co-evolve." And once agents reach the physical world, he said, "every person's possibilities will be amplified tenfold."
Why it matters beyond China
The agent thesis is not uniquely Chinese — OpenAI, Anthropic and Google are all pushing agentic products in the same window. What is distinctive about the Chinese version is how tightly it is coupled to hardware. Yin's "one brain, many bodies" only works if you can actually build the bodies cheaply and at scale, and China's manufacturing base in cars, phones and humanoid (人形机器人) platforms is exactly the lever Western software-first labs lack. That is why a large-model (大模型) founder spends his keynote talking about robots and vehicles rather than benchmark scores.
What this means if you build or buy
- Stop scoring models by benchmark leaderboards. If Yin is right, the unit that matters is the agent that finishes a job, not the model that tops a test.
- Watch the "body," not just the "brain." The interesting Chinese bets in 2026 are increasingly in cars, phones and humanoid (人形机器人) platforms that give agents somewhere to act.
- Put governance on the roadmap now. Yin's own caution — accountability, identity, traceability — is the part enterprises ignore at their own risk.
Whether agents truly leave the screen this year is still an open question. But the person setting the thesis is not a commentator; he runs one of China's better-funded large-model (大模型) labs, and he is betting his company's direction on it.
Honest limitations
This article is built from Yin Qi's WAIC 2026 keynote as reported by 21st Century Business Herald and reprinted by StepFun, plus an English recap (BigGo Finance). It is a single-founder viewpoint, not a consensus position, and the quotes are his claims about trajectory, not audited results. Phrases such as "tens of hours" of autonomous task time are his characterization of model capability, not a standardized measurement, and we did not independently verify them. WAIC is an industry conference with promotional incentives; treat the agent thesis as a strategic bet, not a forecast. No currency conversions are involved because the speech contained no prices.
