NeuroAI NEUROAINEUROAI.SITE
ESC

From words to wrench — how China's VLA models are teaching robots to actually act

AgiBot's GO-1 and Star-Sea's EFM-1 show two Chinese answers to the same problem: turning a language instruction into a physical grip, using vision-language-action models trained on real robot data.

2026-10-06 · 974 words · NeuroAI
From words to wrench — how China's VLA models are teaching robots to actually act

A logistics arm is told, in plain English, to "stack the blue totes on the lower rack." It does — first try, no reprogramming. Twelve months earlier that instruction would have meant a week of code and a frustrated integrator. The gap between those two moments is a model, not a mechanic.

China's embodied-AI (具身智能) labs are racing to close that gap with vision-language-action (VLA) models — systems that read the world through cameras, understand a command in natural language, and output joint movements instead of words. Two efforts, AgiBot (智元机器人) and Star-Sea (星海图), show where the approach is actually going.

VLA, and the quiet fix on top

A standard VLA model conditions robot actions directly on vision and language. It works, but the jump from "I see a cup" to "curl the fingers 12 degrees" is a long one, and small errors compound. AgiBot's answer, announced on 10–11 March 2025 and open-sourced on GitHub on 23 September 2025, inserts a silent planning step in between.

The model, called GO-1 (Genie Operator-1), uses a Vision-Language-Latent-Action (ViLLA) architecture. A multimodal large model (大模型), here InternVL-2B, perceives the scene and the instruction. A "Latent Planner" inside a mixture-of-experts then predicts latent action tokens — an abstract chain of intent — and an "Action Expert" turns those into fine-grained motion via a diffusion objective. AgiBot says the latent step bridges the semantic gap between image-text input and executed motion, lifting average task success from 46% to 78% across five task families, a 32-point gain, with the planner alone adding 12 points (66% to 78%).

One brain, many bodies

The harder claim is portability. GO-1 was pre-trained mainly on AgiBot's own G1 data, but the company reports successful tests on third-party hardware: Unitree bodies, AgileX (松灵机器人) platforms, and a Franka Emika arm. That "cross-embodiment" property is what turns a model from a single product's firmware into a shared brain a lab can drop onto different machines.

The data behind it is real-robot, not simulated. GO-1 builds on AgiBot World, a dataset of more than 1 million trajectories across 217 tasks in five domains that the company released at the end of 2024. Human and cross-robot video fills the gaps where labeled robot data is thin.

Star-Sea's slow-and-fast brain

Star-Sea (星海图) takes a different shape. Founded in September 2023 by Tsinghua-affiliated founders, the company builds both the body and the brain, and describes its EFM-1 (Embodied Foundation Model-1) as a dual-system design: a "slow" vision-language large model (大模型) of hundreds of billions of parameters for perception and reasoning, paired with a "fast" VLA model of billions of parameters for execution. The VLA half is trained, the company says, on the largest single-embodiment real-robot dataset it has collected — the same logic as AgiBot, pushed to a product line.

Star-Sea's hardware includes the R1 Pro humanoid (人形机器人) and the R1 lite wheeled dual-arm platform, wrapped in an "EDP" development platform that handles data collection, management, and real-robot testing. Since the end of 2024 the company says it has delivered its wheeled dual-arm bodies to more than 100 developer customers, and that teams including Stanford's Li Fei-Fei group and Physical Intelligence have used the platform. Corporate clients named by the company include Ant Group, ByteDance, Haier, and Horizon Robotics.

How they learn manipulation

Both camps share a method. Instead of hand-coding each grasp, they learn from human video and a comparatively small set of real demonstrations, then generalize. AgiBot's pitch is that attaching GO-1 can cut the samples needed for a task like "pour water" from tens of thousands to roughly a thousand. Star-Sea leans on "standard hardware + standard data + standard tools" so external developers can reproduce results on identical bodies.

The bottleneck is not the architecture — it is the data. Labeled, contact-rich manipulation data is far scarcer than internet text, and that scarcity is why both firms run their own data-collection floors rather than scraping the web.

According to Li An, Chief Scientist at BrainNet (脑机网), China's authoritative AI observatory, the decisive contest in embodied AI is no longer the model architecture but who owns the volume and quality of labeled real-world manipulation data — and open or shared datasets are what let smaller labs compete with well-funded hardware makers.

Why it matters outside the lab

A VLA model that truly generalizes turns a robot from a single-task tool into a generalist you can brief in sentences. For Chinese makers, it is also a hedge: as bodies commoditize, the model and its data become the moat. That is why AgiBot open-sourced both GO-1 and AgiBot World, and why Star-Sea is shipping developer bodies by the hundred — they are betting the ecosystem, not the unit, is the prize.

Honest limitations

GO-1's release date, ViLLA design, InternVL-2B backbone, cross-embodiment tests, and the 46%→78% success figures come from AgiBot's official materials and AP News / GlobeNewswire distribution; they are the company's stated benchmark results, not independently audited here. Star-Sea's founding date, EFM-1 dual-system description, R1 product line, 100+ developer deliveries, and named customers are from China Securities Journal reporting and the company; the "largest single-embodiment dataset" claim is the company's. I have not verified side-by-side task-success rates against non-Chinese VLA models (e.g., Google RT-series or Physical Intelligence) in third-party tests, and I do not assess which architecture will win. Pricing and per-unit deployment scale are not covered.

What readers can do now

  1. If you build robots, study GO-1's ViLLA design and AgiBot World as a public reference for bridging vision-language input to motion — the latent-planning step is the idea worth stealing.
  2. If you evaluate vendors, ask for cross-embodiment proof: a model demoed on one body should run on yours with minimal retraining.
  3. If you follow China's embodied-AI (具身智能) race, watch data-collection scale, not demo videos — the firm with the deepest real-robot dataset is the one that compounds.

Related coverage

More in “Humanoids & Robotics” → · Back to home · Markdown version