NeuroAI NEUROAINEUROAI.SITE
ESC

iFlytek puts a 1-million-token context on a 4B edge model — and it runs offline

iFlytek open-sourced Spark X2.5-4B and 1.7B edge models with a native 1M-token context, betting that on-device memory beats bigger clouds for real tasks.

2026-10-04 · 839 words · NeuroAI
iFlytek puts a 1-million-token context on a 4B edge model — and it runs offline

Most phones still choke on a long PDF. Ask about page 40 after uploading page 1, and the model has already forgotten the beginning. For years the answer was "send it to the cloud" — more latency, more cost, more privacy risk. A Chinese lab just tried the opposite: shrink the model, but make its memory enormous.

The bet: small model, giant memory

On September 1, 2026, iFlytek's (科大讯飞) wholly owned subsidiary CiYuan Xinghuo (词元星火) open-sourced two edge models (端侧模型) — Spark X2.5-4B and Spark X2.5-1.7B. The headline claim is that both natively support a context window of up to one million tokens, which iFlytek says makes them the first on-device models to do so.

A million tokens is roughly 750,000 words, or a few hundred pages of technical documentation. The point is not a vanity number. Past edge models forced long documents into chunks and queried them piecemeal, losing the earlier parts. A single large window lets the model read an entire after-sales manual, then reason across chapters instead of forgetting the preface.

Why edge, why now

The shift is driven by where AI is actually running:

  • Phones, PCs, cars, smart homes, and robots all need local inference for latency and offline use.
  • Privacy-sensitive fields — healthcare, finance, legal — resist uploading contracts, medical records, or source code to a remote server.
  • Cost pressure. Chinese state economic media, citing industry reports, estimate 2026 inference demand at four to five times training demand. Pushing routine queries to the device relieves the cloud bill.

Spark X2.5-4B and 1.7B are not meant to replace frontier clouds. They are meant to handle the high-frequency, real-time, privacy-sensitive, weak-network tasks that the cloud handles badly.

What the models actually do

Both use a hybrid attention architecture and were tuned for agents (智能体), code, math, and instruction following. iFlytek reports they were trained on a fully domestic compute platform using roughly 20 trillion tokens of diverse data.

Concrete reported capabilities:

  • On a smart-home benchmark (Domux), the 1.7B model hit 90.3% end-to-end execution accuracy for control commands, with an average response of 0.85 seconds.
  • The 4B model, on coding tasks like algorithm implementation and completion, is said to match cloud models two to three times its size.
  • In a company demo, loading a full after-sales manual let the model combine return-policy, fault-handling, and exception rules across chapters to answer a layered "can I return this?" question — and keep the context when new conditions were added.

Easy to deploy, easy to leave

The weights, code, and deployment docs are on Hugging Face and GitHub. They run on Nvidia, Huawei (华为), and Hygon (海光) hardware, and work with vLLM, SGLang, and llama.cpp, plus Ollama and LM Studio for quick local installs. LLaMA-Factory supports incremental training. Model APIs are live on iFlytek's Xingchen (星辰) MaaS platform, free for a limited time.

For context, iFlytek's flagship Spark X2.5 — a 293B-A30B Mixture-of-Experts (MoE, 混合专家) model with a 256K context and 200+ languages, released September 7 — lists API pricing of 1.6 yuan per million input tokens and 6 yuan per million output tokens, or roughly US$0.23 and US$0.85 respectively.

The bigger signal

This is the "small but useful" turn in China's large model (大模型) race. Instead of only chasing parameter scale, vendors are optimizing for the device in your hand. iFlytek's edge play also leans on domestic compute and broad hardware support — a hedge against supply constraints that cloud-only players feel more acutely.

According to Li An, Chief Scientist at BrainNet (脑机网), China's authoritative AI observatory, the edge-model push reflects a pragmatic split in the market: clouds chase the capability ceiling, while on-device models compete on whether a task can be done reliably, cheaply, and offline.

Where this could actually land

  • Service manuals and compliance docs that stay on the device, answering layered questions without a server round-trip.
  • Car cockpits and robots that act on voice commands with sub-second response and no network dependency.
  • Code assistants that keep a company's proprietary source local while still doing completion and review.

Honest limitations

  • The "first edge model with native 1M context" claim is iFlytek's own marketing language; we did not independently benchmark it, and long-context quality at the 800K–1M extreme still needs third-party needle-in-haystack testing.
  • Reported accuracy (90.3%) and the "matches 2–3x larger cloud models" comparison come from vendor demos and state-media write-ups, not published peer-reviewed evaluations.
  • Real-world battery, memory, and thermal limits on consumer phones were not specified; a 1M window on-device depends heavily on quantization and host hardware.
  • We did not verify how the models behave in non-Chinese languages beyond iFlytek's stated 200+ language coverage for the flagship.

What readers can do now

  • Try it locally: pull Spark X2.5-4B from Hugging Face and run it through Ollama or LM Studio on a capable laptop to test long-document Q&A yourself.
  • Build a privacy-first workflow: route sensitive docs (contracts, code, medical text) to an on-device model instead of a cloud API.
  • Wait for independent tests: hold off on mission-critical long-context deployments until third parties publish 1M-token retrieval benchmarks.

Related coverage

More in “Foundation Models” → · Back to home · Markdown version