NeuroAI NEUROAINEUROAI.SITE
ESC

China is writing the rulebook for what AI models are allowed to learn

Two national standards taking effect on 1 November 2025 — GB/T 45652-2025 and GB/T 45674-2025 — set security rules for the training data and annotation behind generative AI (生成式人工智能), a different governance layer from the already-mandated output labels.

2026-10-02 · 780 words · NeuroAI
China is writing the rulebook for what AI models are allowed to learn

A model is only as trustworthy as what it learned from. China decided that "what goes in" deserves its own rulebook — not just "what comes out."

On 1 November 2025, two recommended national standards on generative-AI (生成式人工智能) data took effect. They target the least visible part of the AI stack: the pre-training and fine-tuning data, and the human annotation that shapes it.

The two standards

  • GB/T 45652-2025 — "Cybersecurity Technology — Security Specification for Generative Artificial Intelligence Pre-training and Fine-tuning Data" (《网络安全技术 生成式人工智能预训练和优化训练数据安全规范》). It sets security requirements for pre-training and optimized-training (fine-tuning) data and the activities that process them, and describes how to evaluate compliance.
  • GB/T 45674-2025 — "Cybersecurity Technology — Security Specification for Generative AI Data Annotation" (《网络安全技术 生成式人工智能数据标注安全规范》). It covers the security of annotation platforms and tools, annotation-rule design, personnel requirements, and verification of annotations.

Both were published on 25 April 2025 and took effect on 1 November 2025. They were issued by the Standardization Administration of China (国家标准化管理委员会, SAC) together with the State Administration for Market Regulation (国家市场监督管理总局, SAMR). The official national-standards platform (openstd.samr.gov.cn) lists both as "current" (现行).

Why this is a separate layer

China's AI governance is built in layers, and the distinction matters:

  • Output layer — GB 45438-2025, a mandatory standard, requires visible and hidden labels on AI-generated content. (That rule took effect on 1 September 2025 and is already covered elsewhere.)
  • Service layer — GB/T 45654-2025 sets baseline safety requirements for generative-AI services.
  • Input layer — GB/T 45652 and GB/T 45674, the subject here, govern the training data and the annotation that feed the model.

The key nuance: GB 45438 carries the force of law (it is a GB, not GB/T), while 45652 and 45674 are recommended (GB/T). They are not optional in spirit — they are the technical reference regulators and third-party assessors use when reviewing a model before it launches — but their legal weight is softer than a mandatory standard.

What the data standard actually demands

GB/T 45652-2025 applies to service providers handling pre-training and optimized-training data, and to third-party bodies assessing that data. In practice it pushes providers to:

  • Control the source and legality of training corpora.
  • Manage risks in how data is selected, cleaned and weighted.
  • Run security self-assessment on the data pipeline.
  • Keep records that auditors can inspect.

The companion GB/T 45674-2025 extends the same logic to annotation — the human-labeled examples that tune a model's behavior. It asks organizations to secure the annotation platform, write clear annotation rules, qualify the people doing the labeling, and verify the results.

Where the standards sit legally

Both standards are the technical tail of the Interim Measures for the Management of Generative AI Services (《生成式人工智能服务管理暂行办法》), in force since 15 August 2023. That measure requires security assessment and filing before a public generative-AI service launches; the data and annotation standards give assessors a concrete checklist.

They also feed the large-model filing (大模型备案) process: a filing dossier already expects corpus annotation rules and a keyword filter list. These standards turn those expectations into documented method.

What it does not do

The standards are about security and legality of data, not about making a model smarter or more capable. They do not specify architecture, benchmark targets, or compute. And because they are recommended rather than mandatory, the real enforcement pressure comes indirectly — through the filing and security-assessment gate a model must pass to reach the public.

Honest limitations

Core facts — the standard numbers, Chinese and English titles, publication date (25 April 2025), effective date (1 November 2025), and issuing bodies (SAMR + SAC) — are confirmed on the official national-standards platform (openstd.samr.gov.cn) and corroborated by multiple provincial market-regulation bulletins. The GB (mandatory) vs GB/T (recommended) distinction is taken from the standard codes themselves. The linkage to the 15 August 2023 Interim Measures and to the large-model filing process is drawn from the measures' text and from implementation guidance quoted by government and legal sources. This article describes the standards' scope; it does not reproduce their full technical clauses, and we have not independently audited any company's compliance. No RMB amounts appear, so no currency conversion applies. Analysis is current to 2 October 2026.

What readers can do now

  • If you train or fine-tune models in China, treat GB/T 45652 and 45674 as your pre-launch checklist — corpus legality and annotation controls are now reviewable points, not internal chores.
  • If you buy AI services, ask vendors how they source and annotate training data; the standards give you the vocabulary to demand answers.
  • If you follow governance, read China's AI rules as three layers — input (data), service (safety), output (labeling) — rather than one blanket "AI law."

Related coverage

More in “Policy & Governance” → · Back to home · Markdown version