NeuroAI NEUROAINEUROAI.SITE
ESC

China turns data labeling into an industry — with a 20% growth mandate

A 2024 opinion from the National Data Administration treats data labeling as core AI infrastructure, targeting 20%+ annual growth by 2027 and seeding seven base cities.

2026-10-06 · 1129 words · NeuroAI
China turns data labeling into an industry — with a 20% growth mandate

A medical-imaging startup has the algorithm. What it lacks is 50,000 chest scans, each one carefully marked by a human to show exactly where the lesion sits. That marking work — invisible, tedious, essential — used to be a cottage industry of freelancers. Now it has a national policy behind it.

In late 2024, four central bodies put data labeling on the industrial-policy map. The message was blunt: if you want good AI, you have to treat the labor of preparing its training data as a real, supported sector — not an afterthought.

The document, and who signed it

The document is the Implementation Opinion on Promoting the High-Quality Development of the Data-Labeling Industry — 《关于促进数据标注产业高质量发展的实施意见》 — document number 发改数据〔2024〕1822号. It was dated 26 December 2024 and published publicly on 13 January 2025.

It was issued by four ministries:

  • the National Development and Reform Commission (国家发展改革委, NDRC);
  • the National Data Administration (国家数据局, NDA);
  • the Ministry of Finance (财政部);
  • the Ministry of Human Resources and Social Security (人力资源社会保障部).

The presence of the National Data Administration as a lead signatory is itself notable. The NDA, created in 2023, is the body now steering how China's data is governed and monetized; putting data labeling under its wing signals that labeled data is treated as national infrastructure, not merely a vendor service.

What data labeling actually means here

The opinion defines the industry precisely: the processing of data through screening, cleaning, classification, annotation, tagging, and quality inspection. In plain terms, it is the work of turning raw, messy data into something a machine can learn from.

The opinion calls this the "core productive force" of high-quality AI datasets and the bridge between raw data, algorithms, and real applications. That framing matters because it elevates labeling from low-skill piecework to a strategic input — on par with compute and models.

Why Beijing is treating labels as industrial policy

The stated problem is a bottleneck. High-quality, well-labeled data is scarce, and the opinion notes that shortage directly constrains the training of large models (大模型). Without clean labels, a model trained on mountains of raw text or images learns the wrong things.

The policy logic runs in three steps:

  • Release public-data labeling demand — governments should compile public-data labeling catalogs, and labeling services should enter government procurement.
  • Mine enterprise demand — including a "state-owned enterprise data-efficiency enhancement action" to unlock corporate data for labeling, plus focused work on transport, healthcare, finance, science, manufacturing, and agriculture.
  • Build the ecosystem — cultivate leading firms, "gazelle" and "unicorn" labeling companies, open-source platforms, and third-party services.

The opinion also nods to autonomous driving, low-altitude economy, and digital trade as scenarios where business innovation should pull labeling demand forward.

The 2027 numbers

The headline target is quantitative and unambiguous: by 2027, the data-labeling industry should show significantly stronger specialization, intelligence, and innovation capacity, with its scale rising sharply and a compound annual growth rate above 20%.

It also commits to:

  • a batch of influential, technology-driven data-labeling enterprises;
  • a set of industry–research–application innovation载体 (platforms);
  • a group of distinctive, effective data-labeling bases;
  • a relatively complete industry ecosystem.

A 20%-plus CAGR is a demanding bar. It tells localities and investors that labeling is meant to grow faster than the broader digital sector — and that the state expects to measure it.

Seven cities, one experiment

Ahead of the opinion, in May 2024, the National Data Administration named seven cities to host data-labeling base pilots: Chengdu, Shenyang, Hefei, Changsha, Haikou, Baoding, and Datong. These bases are meant to test the six tasks the NDA outlined — technological innovation, industry empowerment, ecosystem cultivation, standards application, talent development, and data security.

The base model is deliberate. Rather than spreading support thinly, the state picks geographic clusters, seeds them with public and enterprise demand, and hopes a local supply chain of annotators, toolmakers, and quality auditors grows around them — much as earlier policies built chip or new-energy clusters.

From human taggers to smart tools

The opinion is not only about people with mice. It pushes technology breakthroughs — cross-modal semantic alignment, 4D labeling, and large-model labeling — and encourages intelligent labeling tools for review, quality assessment, and expert annotation based on chains of thought. It even calls for self-controlled, software–hardware-integrated labeling equipment.

This is a quiet but important shift: as models get better at pre-labeling, the human role moves up the value chain from marking boxes to verifying and curating. The opinion prepares the industry for that transition rather than defending the old manual model.

It also ties labeling to the broader subsidy toolkit. The opinion explicitly encourages localities to use data vouchers, algorithm vouchers, and computing vouchers (数据券、算法券、算力券) to lower labeling firms' costs — linking this policy to the computing-voucher schemes covered separately on this site.

The honest open questions

The opinion sets direction, not destiny. Several things remain unsettled: how "quality" will be standardized across modalities; how public-data labeling will reconcile openness with privacy; and whether a 20% CAGR survives contact with real procurement budgets. The document also promises national occupational standards for AI trainers and data labelers, building on the fact that "data labeler" was already folded into the "AI trainer" occupation in China's national occupational catalog back in February 2020.

Honest limitations

This article is built on the official NDRC notice (发改数据〔2024〕1822号, dated 26 December 2024, published 13 January 2025) and the National Data Administration's accompanying expert interpretations, cross-checked against Xinhua and CCTV reporting. Caveats:

  • The 20% CAGR and 2027 scale targets are policy goals, not measured outcomes; actual growth will depend on procurement and local implementation.
  • The seven base cities and the February 2020 occupational-classification fact come from National Data Administration and Ministry of Human Resources statements, not from an independent audit.
  • The opinion is a guidance document, not a binding statute; its force depends on subsequent local rules, procurement, and subsidy programs.
  • This piece does not estimate the industry's current size or dollar value, because the official document does not publish a baseline RMB figure, and no currency conversion is therefore applied.
  • It does not assess labor-condition or wage questions inside the labeling workforce, which the opinion addresses only in broad talent-development terms.

What readers can do now

  1. If you train models and operate in China, track the seven base cities (Chengdu, Shenyang, Hefei, Changsha, Haikou, Baoding, Datong) as potential sources of curated, sector-specific datasets, especially in healthcare, transport, and agriculture.
  2. If you run a data or AI team, watch for local "data voucher / algorithm voucher" programs the opinion encourages — they can lower the cost of buying labeling and corpus services.
  3. If you follow AI supply chains, treat labeled-data capacity as a strategic input on par with compute; the 20%-growth mandate signals where government-supported capacity will concentrate through 2027.

Related coverage

More in “Policy & Governance” → · Back to home · Markdown version