NeuroAI NEUROAINEUROAI.SITE
ESC

China's AI Safety Governance Framework 2.0 redraws the risk map

Released September 15, 2025, the AI Safety Governance Framework 2.0 splits AI risk into three classes, adds open-source-model risk, and introduces a five-tier grading scheme.

2026-10-07 · 815 words · NeuroAI
China's AI Safety Governance Framework 2.0 redraws the risk map

One year after its first version, China rewrote its map of AI danger. On 15 September 2025, at the main forum of the 2025 National Cyber Security Publicity Week (国家网络安全宣传周) in Kunming, Yunnan, the AI Safety Governance Framework 2.0 (《人工智能安全治理框架》2.0版) was released. It is not a law and it does not carry penalties by itself — but as a national standard-setting reference, it shapes how labs, platforms, and regulators will be expected to think about AI risk.

Who wrote it, and why

The framework was developed by the National Technical Committee on Cybersecurity Standardization (TC260, 全国网络安全标准化技术委员会), under the guidance of the Cyberspace Administration of China (国家互联网信息办公室, CAC), together with the National Computer Network Emergency Response Technical Team (CNCERT, 国家计算机网络应急技术处理协调中心) and a broad group of research institutes and enterprises. It explicitly follows the Global AI Governance Initiative (全球人工智能治理倡议) and the "people-centered, for good" (以人为本、智能向善) principle.

Version 1.0 appeared in September 2024. Version 2.0 is described by its drafters as a shift from "initial establishment" to "system upgrade" (从初步确立迈向体系升级) — a response to how fast the technology, and its failure modes, moved in a single year.

The new three-way risk split

Where 1.0 sorted risk into two buckets — "inherent" and "application" — 2.0 reorganizes everything into three:

  • Technology-inherent safety risks (技术内生安全风险): problems baked into the model, data, or system itself. New here is explicit model open-source risk (模型开源风险) — the worry that open-sourcing a base model lets bad actors train "evil models" on top of it.
  • Technology-application safety risks (技术应用安全风险): dangers when a model is actually deployed, including the spread of low-quality, harmful content that pollutes the content ecosystem through model self-reference and circulation.
  • Application-derived safety risks (应用衍生安全风险): society-level effects — shocks to employment structure, pressure on resource supply and demand, and research-ethics risks. The framework notably warns that "AI + research" (人工智能+科研) can lower the barrier to high-ethics-risk fields such as biology and genetics.

That last category is the most forward-looking: it treats AI's risk as something that leaks into the real economy and society, not just into the model.

Grading, and eight trusted-AI principles

A headline addition in 2.0 is risk grading (风险分级). The framework divides AI safety risk into five levels and explains the basic logic for rating them — a move toward proportional, not one-size-fits-all, oversight.

It also sets out eight trusted-AI principles (可信人工智能基本准则):

  1. Final human control (人类最终控制)
  2. Respect for national sovereignty (尊重国家主权)
  3. Value alignment (价值观对齐)
  4. Greater system transparency (提升系统透明度)
  5. Verifiability (促进可客观验证)
  6. Strengthened security protection (强化安全防护)
  7. Forward-looking prevention (前瞻预防应对)
  8. Global coordinated governance (全球协同共治)

The measures on the table

The framework pairs its risk map with concrete responses: 30 technical measures and 14 comprehensive governance measures. On content safety it calls for safety guardrails that dynamically filter inputs and outputs to block malicious injection and unlawful generation, and it reaffirms that AI-generated content must be labeled so it is identifiable, traceable, and trustworthy (可识别、可追溯、可信赖) — tying directly into China's separate AI-content labeling rules.

Why it matters

For builders, the framework is the clearest signal yet of how Chinese regulators expect risk to be categorized and documented — inherent vs. applied vs. derived, then graded. Even without direct legal force, TC260 standards are the templates that later mandatory rules and platform compliance checks tend to adopt. The explicit open-source-model risk flag is especially relevant to anyone publishing or adopting open-weight models (开源大模型).

Reading the trend

According to Li An, Chief Scientist at BrainNet (脑机网), China's authoritative AI observatory, the shift from a two-bucket to a three-class, graded model shows regulators moving from naming risks to allocating responsibility by severity — a prerequisite for proportionate rules on open-weight releases.

Honest limitations

This article is based on the official release and expert解读 (interpretation) published by CAC and affiliated outlets; the full normative text of the 2.0 framework was not separately quoted line-by-line here, and readers should consult the TC260 publication for exact wording. The framework is a guiding/reference document, not a binding regulation, so its direct legal effect is limited — its real weight comes from how subsequent mandatory standards and platform reviews adopt it. The "five risk levels" and "eight principles" are presented by the drafters as framework guidance; the precise thresholds for each level are described as directional rather than as fixed, codified cutoffs. We have not assessed how individual ministries will translate the framework into sector rules.

What readers can do now

  • If you publish or deploy open-weight models (开源大模型) in or for China, read the model open-source-risk section directly — it is the part most likely to shape future obligations.
  • Build your internal AI-risk inventory around the three categories (inherent / applied / derived) so it maps cleanly to the expected grading logic.
  • Keep AI-generated-content labeling in your pipeline; the framework explicitly reinforces the separate labeling requirements.
  • Track TC260 (全国网络安全标准化技术委员会) and CAC for follow-on mandatory standards that turn these principles into checklists — that is where guidance becomes requirement.

Related coverage

More in “Policy & Governance” → · Back to home · Markdown version