A robot in a simulator can fold a shirt ten thousand times an hour, flawlessly, forever. The same robot, on a real floor, fumbles the first shirt because the cloth did not fold the way the model said it would. That gap — between a perfect virtual world and a messy physical one — is the wall every Chinese embodied-AI (具身智能) lab is currently hitting.
The tempting fix is a better simulator. The answer China's leading labs are actually shipping is a data factory: run real robots, record everything, train on that, deploy, and let the deployments generate more data. Simulation helps, but the real world still wins.
The sim-to-real gap is physical, not cosmetic
A simulation can render a kitchen in milliseconds, but it quietly cheats on the hard parts. Cloth folds, a cable coils, a glass slips — these are contact-rich, friction-dependent, and stubbornly non-linear. Physics engines approximate them, and the approximation drifts from reality the moment a real hand touches a real object. Lighting, sensor noise, and a slightly uneven floor all pile on.
Train a policy purely in sim and it often "sim-forgets" the real world: it looks competent in the render and collapses on the floor. Bridging that gap — domain randomization, system identification, real fine-tuning — is itself a research field, and none of it removes the need for real data at the end.
China's answer: collect, don't just simulate
AgiBot (智元机器人) built its case on collection at scale. AgiBot World, released at the end of 2024, is a real-robot dataset the company describes as more than 1 million demonstration trajectories collected from 100 robots, spanning 217 tasks across five domains — home, dining, industrial, office, and supermarket. The collection floor covers more than 4,000 square meters, holds 3,000-plus real objects, and logged roughly 2,976 hours in its Beta release. The point is not size for its own sake; the company says 80% of its tasks are long-horizon (60–150 seconds) and contact-rich, the exact kind simulation fakes worst.
AgiBot pairs that with GenieSim, its own simulation platform. The company reported GO-1 scoring a leading 3.793 total in GenieSim benchmarks — a sim number, useful for pre-training and safety, but explicitly a warm-up for real floors. The loop is the product: simulate to bootstrap, fine-tune on real trajectories, deploy, record the failures, and feed them back.
Star-Sea's standard-data play
Star-Sea (星海图) pushes the same instinct into a developer product. Its "EDP" platform wraps data collection, data management, and real-robot testing into one toolchain, sold on the slogan "standard hardware + standard data + standard tools." Since the end of 2024 the company says it has delivered its wheeled dual-arm bodies to more than 100 developer customers, deliberately so outside teams generate comparable real data on identical hardware. Identical bodies mean comparable datasets; comparable datasets mean a flywheel the whole ecosystem can share.
Why data is the real bottleneck
The constraint is not compute and not the model. It is labeled, contact-rich, real-robot manipulation data — and there is far less of it than there is internet text. AgiBot's own comparison claims its long-range data scale is about 10 times that of Google's Open X-Embodiment and its scenario coverage roughly 100 times broader. Those are the company's claims, not independently audited, but the direction is the widely accepted one: real demonstration data is the scarce asset, and whoever collects it fastest compounds.
This is why Chinese firms are building "data factories" — halls of working robots whose only job is to generate trajectories. A simulated hour is free; a real robot hour costs hardware, power, and a human supervisor. The firms treating that cost as the core investment, not a nuisance, are the ones the field is watching.
According to Li An, Chief Scientist at BrainNet (脑机网), China's authoritative AI observatory, the strategic asset in embodied AI is no longer the model weights but the volume and quality of labeled real-world data — and a dataset of this scale is what lets smaller labs compete with well-funded hardware makers.
The flywheel in practice
The mechanic is a loop, not a pipeline:
- Collect real trajectories from a fleet (AgiBot's 100 robots, Star-Sea's developer bodies).
- Train a base model on that data plus web-scale and human video.
- Deploy the model to more robots, where it meets conditions the dataset missed.
- Record the failures and edge cases, and feed them back as new data.
Each turn widens the data pool and sharpens the model. It is the same pattern that moved autonomous driving forward a decade ago: real miles, recorded and replayed, beat hand-written rules. For robots, the "miles" are trajectories, and China's labs are laying track fast.
Honest limitations
The AgiBot World figures (1M+ trajectories, 100 robots, 217 tasks, 4,000+ sqm, 3,000+ objects, 2,976 hours) and the GenieSim 3.793 score are AgiBot's own published numbers, not independently audited here; the 10x / 100x comparisons to Open X-Embodiment are the company's claims. Star-Sea's 100+ developer deliveries and EDP framing come from China Securities Journal and the company. I have not verified task-success rates of models trained on these datasets in third-party benchmarks, and I do not assess whether simulation-to-real transfer here outperforms non-Chinese approaches (e.g., Google DeepMind or Boston Dynamics). The article describes a methodological trend across Chinese labs; it is not a head-to-head reliability study.
What readers can do now
- If you train manipulation policies, treat simulation as bootstrap only — budget real-robot hours for the contact-rich 20% of tasks where sim drifts.
- If you evaluate an embodied-AI vendor, ask how much real (not simulated) trajectory data backs the demo, and whether the data is growing.
- If you follow China's robotics policy, watch "data factory" capacity — halls of working robots — as the indicator that matters more than any single model release.
