NeuroAI NEUROAINEUROAI.SITE
ESC

Cambricon's 700-TFLOPS Siyuan 690 becomes the compute base for China's top open models

Cambricon (寒武纪) mass-produced its 700-TFLOPS Siyuan 690 (思元690) chip in 2026 and now ships Day-0 support for DeepSeek and Zhipu models, anchoring China's domestic AI stack.

2026-10-05 · 808 words · NeuroAI
Cambricon's 700-TFLOPS Siyuan 690 becomes the compute base for China's top open models

When DeepSeek shipped a new flagship model in April 2026, it did not ask its chip supplier to catch up later. The adaptation code was ready the same day. The silicon behind that speed was not Nvidia — it was Cambricon's Siyuan 690 (思元690), a Chinese accelerator that has quietly become the floor under the country's most-used open models.

For years, "domestic substitution" meant a slower chip you bought because policy told you to. That complaint is losing its edge. The story of 2026 is less about one chip beating another and more about an entire Chinese compute stack — chips, frameworks, and models — learning to launch in lockstep.

The chip that changed the math

Cambricon (寒武纪), the Shanghai-listed AI chip designer, mass-produced its flagship cloud accelerator, the Siyuan 690, in early 2026. By the specifications Forbes China and 21st Century Business Herald reported, it delivers more than 700 TFLOPS in FP16 compute, carries 196 GB of HBM3 memory, and reaches interconnect bandwidth above 890 Gbps. Those numbers place it in contention with high-end foreign accelerators rather than a generation behind.

The business result was stark. Cambricon posted its first profitable full year in 2025 — revenue of about 64.97 billion yuan (≈ US$915M / HK$7.15B) and net profit near 20.59 billion yuan (≈ US$2.9B / HK$22.6B) — and on 30 June 2026 its market value crossed 1 trillion yuan, the first STAR Market company to do so. (Read those profit and revenue scales cautiously: they are company-disclosed and the valuation was a single-day peak, not a steady state.)

Day-0 adaptation: the real moat

Raw specs alone do not win developers. What separates Cambricon's 2026 from its earlier years is timing. When DeepSeek-V4 launched in April 2026, Cambricon completed Day-0 adaptation and open-sourced the code the same day. When Zhipu open-sourced its flagship GLM-5.2, Cambricon was among the eight domestic compute platforms ready at launch.

"Day-0" means a model author can tell users "it runs on Cambricon out of the box" on release day, instead of "support is coming." That removes the biggest friction in switching off Nvidia: the wait, the custom kernels, the uncertainty. For model labs racing on release cadence, a chip that is ready on hour one is far more valuable than one that is 10% faster but arrives a month late.

From chip to cloud deployment

Cambricon's gains are not only in model labs. Chinese internet and cloud operators — Alibaba, Tencent, Baidu, ByteDance — have sharply increased domestic accelerator purchases through 2026, deploying them across model training, inference, and content workloads, according to 21st Century Business Herald. ByteDance alone has deployed well over 100,000 Siyuan 590 and 690 units for recommendation, AIGC, and moderation work.

This is the chip-maker × cloud-maker collaboration the market has been waiting for: the silicon vendor optimizes for the exact models the clouds run, and the clouds commit volume orders that fund the next chip. It is a tighter loop than the traditional "buy GPUs, hope they fit" relationship with foreign suppliers.

The strategic "so what"

For global readers, the point is not "China caught Nvidia." It is that a self-contained AI supply chain — domestic chips, domestic frameworks like CANN and NeuWare, domestic models — is now coherent enough to ship products without foreign links at every step. When export controls tighten one input, the rest of the stack absorbs the shock instead of stalling.

That resilience has a cost side too. Cambricon's gross margin, near 54% in early 2026, still trails Nvidia's 75%+, and software maturity (PyTorch migration success, operator coverage) remains the weak link. But the direction — from "usable" to "good enough at scale" — is what is moving budgets.

What readers can do now

  • Track Day-0 counts, not just TFLOPS. When a new Chinese model launches, check how many domestic chip platforms claim same-day support; that number predicts real deployment speed better than a benchmark.
  • Watch cloud procurement signals. Alibaba Cloud, Tencent Cloud, and Baidu AI Cloud adding Cambricon to production workloads is a stronger adoption signal than any demo.
  • Compare total cost of ownership, including software. A cheaper accelerator can cost more if your team rewrites kernels for months; evaluate the published migration tooling before assuming savings.

Honest limitations

Figures here come from Forbes China and 21st Century Business Herald (via Southern+/Nandu reporting) and company disclosures; they are not independently audited, and the trillion-yuan market-cap figure was a single-day peak that later pulled back. Specific deployment volumes at individual cloud firms are reported, not confirmed by those firms' own filings, so they should be read as directional. This article focuses on the compute-base and ecosystem angle; it does not benchmark Siyuan 690 against Nvidia H200/B200 on neutral workloads, nor assess export-control exposure or software-ecosystem gaps in depth. Currency conversions use ≈ ¥7.1 = US$1 and ≈ ¥1 = HK$1.1 as approximations.

Related coverage

More in “Companies & Stack” → · Back to home · Markdown version