For years, Cambricon was the poster child of China's AI-chip ambition and its disappointments at the same time. A lab with brilliant papers and a stock that kept falling. Then the supply of the chips everyone actually wanted got cut off — and the part Cambricon had been quietly building suddenly looked like the most useful thing in the room.
The story of the MLU590 is less about a single spec sheet and more about timing, survival, and the slow grind of building a domestic alternative to Nvidia.
From academic project to "AI chip first stock"
Cambricon was founded in 2016 by Chen Tianshi (陈天石) and his brother Chen Yunji (陈云霁), spun out of the Chinese Academy of Sciences' Institute of Computing Technology. The brothers had pioneered the idea of a "deep-learning processor" years earlier; the company turned it into products.
It listed on Shanghai's STAR Market (科创板) in 2020, billed as the first AI-chip stock of its kind. But the early years were brutal: the company lost money every year from founding through 2024, and its shares fell far below their debut. A pivotal blow came when Huawei, once a major customer for its IP, moved to its own in-house architecture in 2019.
What changed the trajectory was exogenous. As the United States tightened export controls on high-end AI accelerators, Chinese buyers that had planned their roadmaps around Nvidia's A100 found themselves needing a substitute — fast.
The chip that changed the math
That substitute was the MLU590 (思元590), Cambricon's fifth-generation data-center accelerator. It entered volume production in 2024 and landed squarely in the window when the A100 was no longer freely available.
Here is the uncomfortable truth about the MLU590's specifications: Cambricon has never published an official datasheet. Almost every number in circulation comes from media and analyst compilations, and they disagree. What is broadly reported and consistent:
- Built on SMIC's N+2 7nm-class process (deep-ultraviolet lithography, not EUV).
- Equipped with on the order of 80 GB of high-bandwidth memory.
- Peak FP16 throughput in the roughly 256–345 TFLOPS range, depending on the source.
- Performance estimated at about 80% of an Nvidia A100 on relevant workloads.
The honest read is that the MLU590 is not an H100-killer. It is an "A100-class, domestically producible" chip — and for Chinese buyers cut off from the real thing, the relevant comparison was never the 590 against an H100. It was the 590 against nothing.
The software half of the fight
Cambricon understood early that a chip is only as good as the stack around it. Its answer to Nvidia's CUDA is NeuWare, a unified software platform covering compiler, runtime, model-optimization tools and SDKs, first introduced back in 2017. On top of it sits BANG C, a proprietary programming language for the MLU, analogous to CUDA C.
This is where the real moat — and the real weakness — lives. A hyperscaler can swap a card more easily than it can rewrite a training pipeline. Cambricon's pitch is that NeuWare is mature enough to be a genuine alternative, but the ecosystem is still thinner than CUDA's, and that gap, not the silicon, is what makes buyers hesitate.
What comes next: the MLU690
The part drawing the most attention now is the MLU690, Cambricon's next-generation accelerator. Reporting through 2025 described it as designed to approach, or rival, Nvidia's H100 — a design goal attributed to sources and analysts, not a measured result. As of late 2025 it was still in the testing phase, with mass production possibly slipping toward the second half of 2026.
That "if" is large. Even an H100-class domestic chip would still trail Nvidia's current generation on throughput, memory bandwidth and software maturity. And Cambricon's 2026 output is constrained by two scarce inputs downstream of export controls: advanced foundry capacity at SMIC and a stable supply of HBM memory.
The bottom line on the business
The strategic shift showed up in the numbers. Cambricon's cloud-product-line revenue reportedly surged 1,188% in 2024 as the MLU590 ramped. In 2025 the company posted its first annual profit since listing — roughly RMB 2.06 billion (≈ US$286M / HK$2.25B) in net profit, according to its reported earnings, after years of accumulated losses. That is a genuine turning point, not a press-release mirage.
What readers can do now
- If you procure AI compute in China, evaluate the MLU590 on a real workload, not a spec comparison. Its strength is inference and mid-size training clusters; treat unsupported precision formats and cluster-scale limits as things to test, not assume.
- If you are a software builder, budget time for the NeuWare/BANG C learning curve if you consider Cambricon hardware. The porting cost, not the chip price, is usually the hidden line item.
- If you follow the sector, watch the MLU690's production timeline and HBM supply more than any single benchmark. Availability, not peak FLOPS, is the constraint that decides whether Cambricon stays profitable.
Honest limitations
Core facts (Cambricon founded 2016 by Chen Tianshi and Chen Yunji, spun from the Chinese Academy of Sciences; STAR Market listing in 2020; MLU590 volume production in 2024; first annual profit in 2025; MLU690 as next-gen, H100-class design goal, testing phase with possible H2 2026 mass production; NeuWare and BANG C software stack; US entity-list addition in 2022; SMIC N+2 process) are drawn from IEEE Spectrum's reporting on China's AI-chip race, corroborated by Chinese financial outlets (中新经纬, 36Kr, 新浪财经) citing the company's earnings, and the company's own product materials. The MLU590's exact specifications (process node, HBM capacity, FP16 throughput, A100-comparison percentage) are NOT officially published by Cambricon; figures cited are analyst/media estimates and should be read as approximate. The 2025 net-profit figure of ~RMB 2.06 billion is reported from earnings and converted at roughly 7.2 RMB/USD and 1.09 HKD/RMB. This article is descriptive, not investment advice, and contains no stock recommendation. Facts are current to 28 September 2026.
