In 2014, a paper with an unusual name — "DianNao," Mandarin for "computer" — reached the top architecture conference ASPLOS and took best paper, the first time an Asian lab had. Its lead author was 31 years old and, by his own description at the time, working in a fog with no map.
That author is Chen Yunji (陈云霁), a researcher at the Chinese Academy of Sciences' Institute of Computing Technology (中科院计算所). More than a decade before the current AI gold rush, he was already asking a strange question: what would a chip look like if it were built for neurons instead of spreadsheets?
The bet nobody wanted
In 2008, when artificial intelligence was still a dirty word in many chip circles, Chen Yunji and his brother Chen Tianshi (陈天石) began mixing AI with processor design. Grant panels were unimpressed; the brothers did the work on the side. Chen Yunji had come up through the Loongson (龙芯) team under Hu Weiwu (胡伟武), becoming at 25 a chief architect of the eight-core Loongson 3. He knew processors. He wanted to know what a processor built for neural networks might look like.
What a deep-learning processor is
A normal chip is a generalist. A deep-learning processor (深度学习处理器) is a specialist: it trades flexibility for the ability to crunch the matrix maths that neural networks live on, using far less power. The promise is simple — run image recognition, speech, and prediction on a phone or a server without melting the battery.
How DianNao actually worked
DianNao was not a vague idea; it had a concrete shape. The design used a grid of small "processing elements," each handling a slice of a neural network's maths, fed by carefully managed on-chip memory so the chip spent its energy computing rather than shuttling data back and forth to main memory — the usual power hog in conventional processors. That architectural choice, more than raw clock speed, is what delivered the efficiency the team measured.
From DianNao to Cambricon
Working with France's Inria, Chen Yunji's group designed DianNao (电脑), the first deep-learning processor architecture, published at ASPLOS in 2014 and awarded best paper — a first for an Asian institution. They followed with DaDianNao (大电脑) for large networks and PuDianNao (普电脑) and ShiDianNao (视电脑) for other learning tasks.
In 2016 the team released the Cambricon instruction set — the "language" a program uses to talk to the chip — at the ISCA conference, where it drew the meeting's top peer-review score. That same year, Chen Tianshi founded the company Cambricon (寒武纪) to carry the work to market, while Chen Yunji stayed at the institute to keep researching.
Why the instruction set is the real prize
Chips are only as useful as the software written for them. The world runs on two old instruction sets, x86 and ARM, both foreign-born. Chen Yunji's bet was that smart chips needed a new instruction set (指令集) of their own — one where a single command could drive a whole group of neurons' worth of computation. Whoever owns the instruction set owns the ecosystem, and China had rarely owned one.
The research claimed energy efficiency roughly a hundred times that of traditional chips for the same AI workload — a benchmark result the team measured, and one that helped turn a quiet academic project into national strategy.
Why the timing mattered
Chen Yunji started this work during what insiders called an AI winter, when neural networks were unfashionable and funding was scarce. That timing turned out to be an advantage: by the time deep learning exploded after 2012, his group already held years of lead-time and a stack of best-paper results. The lesson he often draws is unglamorous — steady, patient groundwork beats chasing the trend, and a researcher's job is to work the quiet field others skip.
The man behind the quiet revolution
Chen Yunji entered the University of Science and Technology of China's gifted-students program at 14, earned his PhD at 24, and became a full researcher at 29. In 2015 he appeared on MIT Technology Review's "Innovators Under 35" list. By 2024 he was deputy director of the Institute of Computing Technology and a recipient of the national May 1 Labour Medal. International peers, including the journal Science, have called him a pioneer of the field.
Honest limitations
Chen Yunji is a researcher; Cambricon the company is a separate entity with its own commercial and financial pressures, and its story is not this profile's to tell. The "hundredfold" efficiency figure is a research benchmark under specific workloads, not a guarantee in every product. An instruction set is only as strong as the software built on top of it — adoption, not invention, is the hard part — and established ecosystems such as Nvidia's CUDA remain the bar to beat. This profile covers the scientist's research contribution, not the company's market fortunes.
What readers can do now
- If you study computer science, the frontier is at the seam between architecture and machine learning; the next DianNao will likely come from there.
- If you invest or strategise, learn to separate a researcher's breakthrough from a company's quarterly results — they are different bets.
- If you follow tech policy, watch the open instruction-set (指令集) contests; whoever sets the standard shapes a decade of hardware.
