---
title: "China's New Training-Data Rules Put Copyright and Quality on the Same Checklist"
date: 2026-10-03
category: Policy & Governance
site: NeuroAI
canonical: https://neuroai.site/a/na-policy-training-data-copyright
language: en
---

# China's New Training-Data Rules Put Copyright and Quality on the Same Checklist

> Three national standards effective November 1, 2025 spell out how AI training data must be sourced, labeled, and verified — with copyright sitting at the center.

A novelist in Guangzhou finds her published books inside a model's training set without permission. A startup in Shenzhen scrambles to prove its dataset is clean before a client audit. A standards engineer in Beijing writes the checklist that both sides will now be measured against.

China's answer to the training-data problem is not a single new law but a stack of standards, and the core of it went live on November 1, 2025. Three recommended national standards (GB/T) now define what "clean" means for the training data (训练数据) that feeds a generative model (生成式人工智能) — covering where the data came from, how it was annotated, and whether anyone's rights were stepped on along the way.

## The rules already on the books

Before the new standards, the obligation to use lawful training data sat in the **Interim Measures for the Management of Generative AI Services (生成式人工智能服务管理暂行办法)**, in force since August 2023. Those measures require providers to use data from lawful sources, respect intellectual property, and obtain consent where personal information is involved. They set the principle. What arrived in November 2025 was the operating manual.

The three standards, all published on April 25, 2025 and effective November 1, 2025:

- **GB/T 45654-2025** — Cybersecurity Technology — Basic Security Requirements for Generative AI Services.

- **GB/T 45652-2025** — Cybersecurity Technology — Security Specification for Pre-training and Optimized-training Data of Generative AI.

- **GB/T 45674-2025** — Cybersecurity Technology — Security Specification for Data Annotation of Generative AI.

All three are recommended (GB/T), not mandatory, which gives firms room to adopt them as the recognized benchmark rather than a hard ceiling.

## Three duties on the data

The training-data standard distills the obligation into three pillars:

- **Lawful sourcing (数据来源合法)** — data should come through user authorization, copyright verification, or proper procurement, lowering the risk of poisoned or misappropriated material and making accountability traceable.

- **Quality assurance (数据质量保障)** — content must be filtered for false, illegal, or harmful information, and checked for accuracy, diversity, and misleading content. Annotation quality must be verifiable.

- **Security management (数据安全管理)** — encryption, access control, and detection across collection, storage, transmission, and use, plus regular updates and adversarial training.

Copyright lives inside the first pillar. The interim measures already bar training on material that infringes another party's intellectual property; the standards turn that bar into a procurement and verification step rather than a courtroom afterthought.

## The copyright pressure that pushed this

The driver is not abstract. Chinese legal media has reported cases where a leading AI-art platform was sued by dozens of institutions for training on tens of thousands of copyrighted works without authorization, with the dispute dragging on for roughly ten months because courts lacked clear rules on AI-output copyright, fair-use boundaries, and infringement standards. That kind of case — creator versus model, with no shared rulebook — is exactly what a sourcing standard is meant to pre-empt.

## What the annotation standard actually demands

**GB/T 45674-2025** is the most operational of the three. It governs the data-annotation (数据标注) process across four dimensions: the annotation platform's security, the annotation rules, the personnel, and verification. It splits labeling into two types — functional annotation, which shapes what the model learns, and safety annotation, which filters harmful content — and requires safety annotations to be checked item by item (逐条核验). Annotation staff must be managed and trained, and the organization must keep records it can show a client or third-party auditor.

An earlier standard, **GB/T 42755-2023, Artificial Intelligence — Data Annotation Procedure for Machine Learning**, already mapped the end-to-end annotation project workflow. The 2025 standard tightens it specifically for generative models, where a single bad label can propagate across billions of parameters.

## Why "recommended" still carries weight

Because the standards are recommended, a firm is not automatically fined for not following them. But in practice they function as the default evidence of diligence. When a regulator or a business partner asks "is your training data compliant?", the fastest answer is "we meet GB/T 45652 and 45674." The National Data Administration has also pushed a high-quality-dataset initiative, signaling that clean, traceable, well-annotated data is treated as national infrastructure, not a private concern.

For foreign observers, the structure is familiar in shape: a principled law, then technical standards that make the principle auditable. The difference is speed and density — China issued the interim measures, then the labeling rule, then the training-data standards in a tight sequence, building the checklist bottom-up.

## What readers can do now

- If you train or procure models in China, require a data-source ledger (authorization, copyright check, or purchase record) as a contractual deliverable — the standards make this the expected proof.

- If you are a creator, document your works and registrations; the lawful-sourcing duty is your leverage if a model ingested them without permission.

- If you evaluate AI vendors, ask which of GB/T 45652, 45654, and 45674 they certify against, and request the annotation verification records.

## Honest limitations

This article draws on the published interim measures, the three GB/T standards and their implementation dates on the national standards platform, and Chinese legal-media reporting on copyright disputes. The standards are recommended, not mandatory, so "compliance" is benchmark adherence rather than a legal requirement in every case. Descriptions of the annotation standard's internal structure summarize the published specification; firms should read the full texts for binding detail. Figures on dispute duration are reported examples, not comprehensive statistics, and the National Data Administration's dataset initiative is characterized from general reporting rather than a single cited issuance.

---

Published by NeuroAI (https://neuroai.site/) — https://neuroai.site/a/na-policy-training-data-copyright
Free to quote with attribution and a link to the original.
