NeuroAI NEUROAINEUROAI.SITE
ESC

A 293-Billion-Parameter Model Trained Entirely on Domestic Chinese Chips

iFlytek's Spark X2.5 was trained and served end-to-end on Chinese accelerators — and shipped with open on-device models small enough to run on a laptop. The constraint moved from silicon to electricity.

2026-09-07 · 702 words · NeuroAI
A 293-Billion-Parameter Model Trained Entirely on Domestic Chinese Chips

Most people have had this experience: an AI tool goes down mid-task, or a subscription you just renewed quietly changes its limits. It is easy to assume the platform is being difficult.

The harder explanation is usually this: someone else owns the power station, and someone else owns the furnace.

On 7 September 2026, iFlytek released Spark X2.5 and moved that problem a long way forward.

Key takeaways

  • Architecture: Spark X2.5 uses a 293B-A30B mixture-of-experts design — 293 billion total parameters, with about 30 billion active on any given query.
  • The headline is not the parameter count: its full training and inference pipeline runs on entirely domestic Chinese compute.
  • On-device models shipped alongside it: 4B and 1.7B variants, open-sourced, supporting up to 1 million tokens of context, deployable locally.
  • The 1.7B model is small enough to run on an ordinary laptop or a high-end phone — meaning documents never have to leave your machine.
  • The next cost battlefield is electricity, not chips. When compute stops being scarce, power and cooling become the variable that sets price.

Why "all-domestic" is the actual dividing line

For years, Chinese model labs have operated with supply uncertainty hanging over high-end accelerators: how many can be bought this quarter, and whether they can be bought next quarter, was never fully in their control.

That uncertainty has a direct consequence for users. A company that cannot guarantee its own compute cannot commit to long-term pricing — so quotas change, and peak-hour queues appear.

Running the full pipeline on domestic hardware is, in effect, building the power station inside your own fence. The transmission chain to an end user is short: stable supply makes long-term price competition possible; price competition stops your tools raising prices every few months; lower cost makes genuinely offline, free personal tools viable.

The open on-device releases may matter more

The flagship announcement is the one that gets coverage, but the 4B and 1.7B open models are the part with immediate consequences.

A 1.7B model with a one-million-token context window can ingest seven or eight full-length novels at once — or every meeting note, contract and product document your company produced in the last six months — and do it without uploading any of it.

For anyone handling financial records, legal material or confidential drafts, that is more practically valuable than any benchmark.

What gets more expensive when models get cheaper

There is a counter-intuitive consequence worth stating plainly. Once compute is no longer the bottleneck, the scarce input changes — and it becomes electricity.

A large model is a machine for turning power into intelligence. Training runs consume the energy of a small town; every query you send spins up chips that draw current, generate heat and require cooling.

So the competitive frontier shifts from who has the most accelerators to who has cheap green power and data centres sited where electricity is abundant. The likely end state for users is that AI pricing starts to vary by region much as electricity tariffs do.

Three cooling observations

Running is not catching up. Completing a full domestic pipeline answers "can it be done," not "how close is it to the frontier." Process-node gaps and large-cluster interconnect efficiency still need time.

Open is not the same as good. A 1.7B model will summarise and pre-screen long documents competently. It will not replace a frontier cloud model for complex reasoning. Do not mistake "runs locally" for "performs equivalently."

Ecosystem is the real moat. Half of a model's usefulness is capability; the other half is how many people have built tooling, written tutorials and hit the edge cases already. On that measure, Chinese open models are still accumulating.

What to actually do

Test your most expensive workflow first — long-document organisation, meeting minutes, contract pre-screening — on a local model with non-sensitive material. One afternoon of that beats reading ten reviews.

Do not pay a premium for the word "domestic." Another round of price cuts is likely within six months. Buy the tool that solves your problem.

And keep one habit regardless of vendor: sensitive material does not go to the cloud.

Specifications and release details as announced by iFlytek on 7 September 2026.

More in “Chips & Compute” → · Back to home · Markdown version