A filmmaker wants a 60-second ad where the same spokes-character appears in five different scenes, holding the same product, with the same face and the same jacket. Until recently, every AI video tool eventually broke that promise: the character would subtly morph between shots, the logo would drift, the product would reshape itself. Consistency, not raw image quality, was the wall.
On October 21, 2025, the Chinese startup ShengShu Technology (生数科技) pushed that wall back. It shipped a feature called "reference-to-video" (参考生视频) inside its Vidu Q2 model, and it is the clearest attempt yet to solve the most boring, most important problem in commercial AI video.
The problem every AI video tool kept hitting
Text-to-video models are good at a single beautiful shot. They are bad at remembering. Ask for "a woman in a red coat walking through a market," generate it three times, and you get three different women, three different coats. For advertising, film and anything with a recurring character, that is fatal. The industry calls this subject consistency, and for years it was the main reason AI clips stayed in the "demo" bucket instead of the "ship it" bucket.
ShengShu's answer is to let the user upload the things that must stay fixed.
What ShengShu (生数科技) actually shipped
The "reference-to-video" (参考生视频) feature accepts up to seven reference images — faces, gestures, scenes or props — and fuses them with a text prompt. A "multiple-entity consistency" engine then keeps each element distinct and faithful to its source, even as the camera moves or the scene changes.
Concretely, a creator can:
- Pin a specific actor's face and a specific product into the same clip.
- Combine unrelated elements (a historical figure, a meeting room, a prop) into one coherent scene.
- Generate transition animations from a first and last frame.
- Export at up to 4K, with clips running roughly 2 to 8 seconds per generation.
The company also opened a Vidu Q2 MaaS API on the same day, so businesses can fold the feature into their own pipelines.
Why this matters more than "prettier video"
Most generative-video competition has been a beauty contest — who renders the nicest waterfall. Consistency is different: it is the feature that turns a toy into a tool. An e-commerce team can finally generate a 360-degree product spin with the logo locked in. An animator can keep a character's face stable across a sequence. A brand can localize one hero film into ten languages without reshooting.
ShengShu frames this as the shift "from AI creation to AI performance" — teaching the model to act a role rather than merely paint a frame. That is marketing language, but the underlying capability (stable identity across shots) is exactly what production teams have been waiting for.
How far the platform has come
Vidu is not a one-shot product. Its lineage is worth a line:
- The team says its U-ViT architecture was an early Diffusion–Transformer hybrid, and its DPM-Solver diffusion solver won an ICLR 2022 Outstanding Paper award.
- Vidu 1.5 (late 2024) introduced multi-entity consistency from reference images.
- Vidu 2.0 (early 2025) cut generation to under ten seconds per clip at roughly half the then-industry cost.
- Vidu Q1 added natively generated, synchronized 48 kHz audio.
- Vidu Q2 (Sept 25, 2025 model; reference-to-video on Oct 21, 2025) is the consistency leap.
ShengShu reports that, since Vidu's launch in April 2024, the platform has reached users in more than 200 countries, gained roughly 30 million users, and produced over 400 million videos. These are company-reported figures and should be read as the vendor's own claims, not audited counts.
Where it still falls short
Reference-to-video does not mean "finished film." Clips are still short (seconds, not minutes), and complex emotional audio can lag behind leading US models such as Google's Veo. Independent tests noted that Vidu Q2's multilingual lip-sync is strong but its emotional delivery is more subdued than Veo 3.1. And like every generative-video tool, it inherits the deeper questions about copyright, deepfakes and consent that no architecture alone resolves.
Honest limitations
- The headline adoption numbers (30M users, 400M videos, 200 countries) come from ShengShu's own press materials and are not independently audited.
- The "world's first" and "SOTA" claims about U-ViT and DPM-Solver are the company's framing; we did not re-verify the academic comparisons.
- The feature launched in China and via global API; real-world quality on non-Chinese faces, languages and products is best confirmed by testing, not by the launch demo.
- We did not benchmark Vidu Q2 against Veo or Sora ourselves; comparisons here come from third-party write-ups and the company's own side-by-side tests.
What readers can do now
- Test the consistency claim yourself — upload one face and one product image to Vidu and generate a 5-second clip; check whether the face and product survive a camera move.
- If you produce ads or shorts, prototype a localized hero film with reference-to-video (参考生视频) before commissioning a reshoot.
- Watch the benchmark, not the trailer — judge any AI-video tool by whether a character stays identical across ten shots, not by its best single frame.
