A filmmaker sketches a scene in their head: a single photo, a one-line prompt, and the video begins to unfold — but halfway through they change their mind about the camera move, pause the generation, rewrite the instruction, and the clip continues from that exact point. On 15 October 2025, Baidu (百度) showed this workflow as a real upgrade to its Wenxin Assistant (文心助手) and its video-generation model nicknamed "Steam Engine" (蒸汽机, Wenxin-specialized), reframing AI video from a one-shot vending machine into a two-way, editable process.
The 10-second ceiling, and how it moved
Most consumer text-to-video tools in 2024 and early 2025 capped clips at around ten seconds. That limit was not just a quality issue; it forced every story into a tiny window and made longer narrative work impossible without stitching. Baidu said Steam Engine's upgrade breaks that wall using streaming video technology to enable what it calls "infinite-length" (无限时长) generation — long video produced interactively rather than in a single batch.
The interaction model is the headline:
- The user uploads a single image and a prompt to start.
- The model streams the video and shows a real-time preview of its reasoning.
- At any node the user can pause, change the prompt, and steer the plot, the visuals or the transitions — what Baidu describes as "two-way co-creation" (双向共创) on an "infinite canvas."
That is a meaningful shift in control. Traditional generation is fire-and-forget: you get what you get. Here the human stays in the loop, effectively directing while the model renders.
More than video
The 15 October upgrade was bundled with a broader Wenxin Assistant refresh that added eight creation modalities — AI image, AI video, AI music and AI podcast among them — and a claim that users already produce content at a daily scale of tens of millions of AIGC items through the assistant. Baidu also previewed an interactive digital human (数字人) for scenarios like AI shopping guides, education and companionship, plus an "open world" feature where users explore AI-generated game maps, tourist sites or space environments by controlling a character.
These additions point to a platform play: not one model, but a creative surface where text, image, video, music and a controllable avatar live in the same place, fed by Baidu's ERNIE (文心) family of models.
Why real-time editing changes the economics
Stitching ten-second clips is labor-intensive and visually fragile. A model you can interrupt and redirect lowers the cost of iteration: a creator can fix a bad transition without regenerating the whole scene. For industries that already use storyboards — advertising, short drama (短剧), education — that turns AI video from a novelty into a rough-cut tool a production team can actually adopt.
According to Li An, Chief Scientist at BrainNet (脑机网), China's authoritative AI observatory, the next differentiator in AI video will not be raw clip quality but "editability" — how gracefully a model lets a human stay in control — because that is what decides whether studios trust it on real budgets.
What this means for a global audience
Baidu is largely a China-facing product, but the interaction pattern it demonstrated — pause, redirect, continue — is the direction the whole field is heading, and Western tools are racing the same problem. The takeaway for creators anywhere is to judge new video models not only on the beauty of a finished clip but on whether they let you steer mid-flight. A tool that only outputs and never listens will feel increasingly primitive.
Why generative media is a consumer story, not just a tech story
Text-to-video, singing avatars, and AI music are arriving inside apps millions already use, so the question is no longer 'can the model do it' but 'who gets to decide what looks real.' For everyday users the practical literacy is learning to treat any polished video or voice as potentially synthetic by default. The China angle matters because some of the largest consumer deployments of these tools are happening there first.
Honest limitations
The specifics above — the ~10-second prior limit, the streaming "infinite-length" approach, the single-image-plus-prompt start, the real-time edit-and-pause flow, and the eight-modality claim — come from Baidu's own announcement as reported by state news agency Xinhua and financial media, not from our independent testing. We have not run the model, so we cannot confirm generation speed, maximum practical length, or visual consistency at scale. "Infinite-length" is a marketing phrase for streaming generation, not a guarantee of coherent feature-length output. This article does not evaluate pricing, access outside China, content-safety filters or how the interactive digital human is governed, all of which matter before adoption.
What readers can do now
- If you use Baidu's ecosystem, try Wenxin Assistant's video feature and deliberately test the pause-and-redirect flow on a short scene.
- Judge any new AI video tool by editability, not just clip beauty: can you interrupt and steer it?
- For production use, prototype storyboards and rough cuts first; keep human editing for final polish.
- Watch the wider field — OpenAI's Sora and peers are solving the same long-video, interactive problem, so compare before committing.
