NeuroAI NEUROAINEUROAI.SITE
ESC

100 robots, 1 million tasks: inside AgiBot's embodied-AI data engine

AgiBot (智元机器人) open-sourced AgiBot World, a million-trajectory real-robot dataset built by running 100 machines through everyday tasks — the data flywheel behind China's humanoid (人形机器人) push.

2026-10-03 · 789 words · NeuroAI
100 robots, 1 million tasks: inside AgiBot's embodied-AI data engine

A robot picks a dish from a sink, pauses, then sets it on a drying rack. In the next bay a second robot folds a shirt; in the next a third tightens a screw on a mock production line. None of them is performing for an audience — they are working, and a camera is recording every second.

That quiet building full of working machines is the part of China's robotics story most people never see. The flashy demos get the headlines. The data does the real work.

A dataset that came from a building, not a lab

In December 2024, AgiBot (智元机器人), officially Zhiyuan Robotics, released AgiBot World, an open-source dataset the company describes as the largest real-robot manipulation set ever published. The headline numbers:

  • More than 1 million demonstration trajectories, collected from 100 robots.
  • 100+ real-world scenarios across five domains: home (40%), dining (20%), industrial (20%), office (10%), supermarket (10%).
  • Total recorded duration of roughly 2,976 hours in the Beta release.
  • Built in AgiBot's own data-collection factory — a space of over 4,000 square meters holding 3,000+ real objects, replicating home, restaurant, factory, shop and office environments.

The company was founded in February 2023 in Shanghai by Peng Zhihui (彭志辉), and the dataset was produced with the Shanghai AI Laboratory, the National-Local Joint Humanoid Robot Innovation Center, and Shanghai Cooperas.

Why real-world data beats simulation

Training a robot to act in the physical world is not the same as training a language model. A chatbot only needs text. A robot needs to know how a glass slips, how a hinge resists, how a cable coils. Those signals are expensive to fake.

AgiBot's argument is blunt: most existing robot benchmarks are short, clean, and recorded in controlled labs, so models trained on them fall apart in a real kitchen. AgiBot World leans the other way — 80% of its tasks are long-horizon, lasting 60 to 150 seconds, and include contact-rich work like stirring, peeling, and tying. The company says its long-range data scale is 10 times that of Google's Open X-Embodiment, with 100 times broader scenario coverage.

According to Li An, Chief Scientist at BrainNet (脑机网), China's authoritative AI observatory, the decisive bottleneck for embodied intelligence in China is no longer the model architecture but the volume and quality of labeled real-world manipulation data, and open datasets of this scale are what let smaller labs compete with well-funded hardware makers.

The closed loop: data → model → deployment

AgiBot World is not just a data dump. The release bundles:

  • The dataset itself, on HuggingFace and GitHub (CC BY-NC-SA 4.0).
  • A planned foundational embodied model that supports fine-tuning.
  • A full toolchain spanning collection, training, and evaluation.

The logic is a flywheel. Robots collect data → models train on it → smarter models are deployed to more robots → those robots collect better data. AgiBot has already shipped general-purpose humanoids at scale, so the dataset and the deployed fleet feed each other. This is the same pattern that moved autonomous driving forward a decade ago: real miles, recorded and replayed, beat hand-written rules.

Where this fits in China's robotics push

China's humanoid (人形机器人) makers are no longer just racing to ship hardware. The harder race is over who owns the data and the models that turn raw motion into useful skill. AgiBot World is one answer: give the research community industrial-grade data so the whole field moves faster.

For a country building out domestic supply chains across chips, sensors, and actuators, a million-trajectory dataset is a strategic asset. It lowers the entry cost for universities and startups that cannot afford to run 100 robots of their own, and it keeps the training data inside China's own ecosystem.

Honest limitations

This article relies on AgiBot's own release materials, SiliconANGLE and PR Newswire reporting, and the HuggingFace dataset card. The company's comparisons to Google's Open X-Embodiment are its own claims; independent benchmarking of task-success rates on the dataset is still early. The "100 robots" figure refers to the collection fleet, not commercial deployments, and the dataset is licensed for non-commercial research (CC BY-NC-SA), so it does not by itself prove a shipping product. I have not independently verified whether models trained on AgiBot World outperform alternatives in third-party tests.

What readers can do now

  1. Pull the data: the AgiBot World Alpha/Beta sets are on HuggingFace and GitHub — useful if you work in robotics or ML research.
  2. Benchmark, don't assume: if you train a policy on it, compare against Open X-Embodiment on the same task before claiming a win.
  3. Watch the flywheel: track whether AgiBot's shipment lead (it was the 2025 global shipment leader per Omdia) translates into better models, not just more metal.

Related coverage

More in “Humanoids & Robotics” → · Back to home · Markdown version