"Is our model good?" is the question every lab dodges with a leaderboard. China answered it with a standard: a published method for testing large models (大模型), effective the day it was released.
It is drier than a benchmark duel — and far more durable.
The standard, in plain terms
The document is GB/T 45288.2-2025, full title Artificial intelligence — Large-scale model — Part 2: Testing and evaluation for metrics and methods (《人工智能 大模型 第2部分:评测指标与方法》).
Verified on the national standards platform (samr.gov.cn):
- Issuer: SAMR and SAC.
- Published: 28 February 2025.
- Effective: 28 February 2025 (same day).
- Status: current (现行).
- Type: recommended (GB/T).
It belongs to a series. GB/T 45288.1-2025 (Part 1: General requirements) shipped the same day. Part 2 is the measurement half; Part 1 is the baseline the model must meet.
What it actually measures
The standard organizes evaluation into two blocks:
- Evaluation indicators (评测指标): understanding ability, and generation ability.
- Evaluation methods (评测方法): the dataset used, the environment, the tools, and how the test is run.
In other words, it is not a single score. It is a recipe: define what "understands" and "generates" mean, then specify how you prove it — data, setup, tooling, procedure. The appendix even lays out objective and subjective calculation methods.
Who wrote it
The drafting coalition is the clearest signal of its intent. The list includes the China Electronics Standardization Institute (CESI), the Institute of Automation at the Chinese Academy of Sciences, and companies across the stack: Baidu, Alibaba Cloud, Huawei Cloud, Tencent, iFlytek, Zhipu, Xiaomi, Ant, SenseTime, Tsinghua and others.
That breadth matters. A standard only works if the labs being measured helped write it. China's largest model builders co-authored the ruler they will be measured by.
Why evaluation is a policy lever
Training a model is a research act. Measuring it is a governance act. Once a country defines how a model is evaluated, it can:
- Compare domestic models on a common basis instead of each lab's favorite benchmark.
- Feed procurement and certification — government and state-owned buyers need a defensible "good enough."
- Tie evaluation to the filing and security regime (see GB/T 45654-2025) so that "safe" and "capable" are assessed, not asserted.
This is the quiet infrastructure of China's AI governance: not one ban, but a stack of measurement rules.
How it differs from the security baseline
Two standards, two jobs:
- GB/T 45654-2025 (security baseline): does the model behave safely before launch?
- GB/T 45288.2-2025 (evaluation): how capable is the model, and how do we prove it?
One is about risk; the other about performance. Together they let a regulator ask both "is it safe?" and "does it work?" in standardized language.
What a buyer should take away
For anyone procuring a Chinese model, the standard is a negotiation tool. Instead of accepting a vendor's favorite leaderboard, you can ask which understanding and generation indicators they test against, and how. A serious answer names the dataset and the method; a vague one reveals the gap. The standard also gives state-owned and enterprise buyers a defensible basis to say "this model meets the national evaluation method" — useful exactly when procurement needs a paper trail that survives an audit.
The limits of any yardstick
No evaluation standard captures everything users care about: tone, safety in the wild, cost at scale, or whether a model quietly drifts between versions. GB/T 45288.2-2025 measures what it measures and leaves the rest to judgment. That is true of every benchmark, including the ones Western labs publish. The value is comparability, not completeness — a common ruler beats ten private ones, even if the ruler itself is imperfect. For buyers, the practical move is to use the standard as a floor, then layer their own red-team tests on top for the risks this text does not reach.
Honest limitations
This article reflects the standard's scope and dates from the official national-standards record and its drafting-unit list; it does not reproduce the full technical text. Because the standard is recommended (GB/T) and method-oriented, it sets a procedure, not a pass/fail bar — adoption depends on buyers and regulators choosing to require it. We cannot quantify how many labs have conformed in practice. It also targets general large models and does not replace domain-specific evaluation for medical, scientific or safety-critical use.
What readers can do now
- If you evaluate Chinese models, map your internal benchmark onto GB/T 45288.2-2025's understanding/generation split — it is becoming the common language.
- Pair this with the security baseline (GB/T 45654-2025) when assessing a vendor: capability and safety are now separately standardized.
- Watch for Part 3+ of GB/T 45288; evaluation standards tend to expand into domains like multimodal and embodied models. Today's text covers text large models, and the frontier is already moving past that, so the standard is a starting line rather than a finish.
