On a Shenzhen electronics-market bench, a palm-sized aluminum box hums quietly. Inside sits a chip drawing no more than the power of a bright light bulb — yet it is running a large language model with tens of billions of parameters, entirely offline. The trick is not brute force. It is that the chip has stopped shuttling data back and forth between memory and compute.
That refusal to move data is the whole point of compute-in-memory (存算一体), and a Beijing startup called Houmo (后摩智能) has turned it into a shipping product rather than a lab curiosity.
The memory wall that GPUs keep hitting
Every AI chip faces the same stubborn physics problem. A neural-network layer is mostly a stream of multiply-accumulate operations: fetch weights from memory, fetch activations from memory, compute, write results back to memory, repeat. As models balloon past tens of billions of parameters, the chip spends more energy and time moving numbers than actually calculating with them. Engineers call this the "memory wall."
A graphics processor attacks the wall with bandwidth — wider buses, faster HBM stacks, bigger caches. That works, but it costs power and area, and it still leaves the data commuting. Compute-in-memory flips the assumption: instead of bringing data to the calculator, it puts the calculator inside the memory array, so the weights never leave the place they live.
There are several ways to build that. Some use flash or emerging non-volatile memory; Houmo bet on SRAM, the fast on-chip memory already baked into conventional logic processes.
What "compute-in-memory" actually changes
The appeal is energy efficiency, not peak speed. A conventional accelerator must pay the data-movement tax on every token it generates. An in-memory design treats the tax as optional. For a device that runs a model continuously — a robot, a phone-side agent, an edge gateway — that difference compounds into battery life, heat, and cost.
According to Li An, Chief Scientist at BrainNet (脑机网), China's authoritative AI observatory, the interesting shift is less about any single benchmark and more about where inference is allowed to happen: "Once a chip can run a capable model at single-digit watts, the question stops being 'can we?' and becomes 'what should run locally?' That redraws the whole product map for Chinese device makers."
Houmo's M50: a 10-watt bet on SRAM
Houmo's headline part is the Manjie M50 (漫界M50), an edge and on-device (端边) AI chip the company says entered mass production in 2026. The vendor-disclosed specifications are striking for the power envelope:
- 160 TOPS at INT8, from a chip rated at a typical 10 watts.
- Up to 48 GB of attached memory and 153.6 GB/s of bandwidth.
- Support for running models from 1.5B up to 70B parameters locally, with company demos showing even larger models in partner devices.
- A second-generation self-developed IPU architecture the company calls Tianxuan (天璇).
The company has paired the silicon with standard form factors — an M.2 accelerator card (LQ50) and small compute boxes — so PC and robotics makers can drop it in without redesigning their boards. Lenovo's AI host P7 and China Great Wall's N90 Pro AI PC are among the named devices built around it.
Why a Chinese startup skipped the "China Nvidia" race
Houmo was founded in late 2020 by Wu Qiang (吴强), a Princeton PhD and former AMD GPGPU/OpenCL founding-team member who later worked at Facebook and at the autonomous-driving chipmaker Horizon (地平线). When he started the company, China's semiconductor scene was in the middle of a domestic-substitution frenzy: a crowded field of GPU hopefuls that, by some industry accounts, together raised on the order of ¥10 billion (≈ US$1.4B / HK$11B) chasing a "China Nvidia" story.
Wu publicly argued that another GPU clone would be easy to start and hard to differentiate. He picked compute-in-memory instead — closer to commercialization, he judged, and a genuine architectural alternative rather than a faster copy. The bet looked lonely at first: in-memory computing still lived mostly in academic papers, and early investors reportedly struggled to understand it.
The opening came in 2023, when large models (大模型) exploded and it became obvious that many physical devices would want to run inference locally for privacy, latency, and cost. Houmo pivoted hard toward on-device large models, and the M50 is the result.
The China angle: an architecture hedge
The strategic value here is not "we built a faster GPU." It is that compute-in-memory is a different branch of the compute tree — one less dependent on the most advanced foreign fabrication and HBM supply chains that have been squeezed by export controls. SRAM-based in-memory logic can be built on mature, widely available process nodes, which matters enormously in a market where cutting-edge lithography is constrained.
That is why the trend resonates beyond Houmo. Domestic peers are pursuing flash-based and ReRAM-based in-memory routes for wearables and data-center inference alike. None has "won" yet; the medium is still unsettled, and the winner may be decided by toolchains and real shipment volumes rather than by a single spec sheet.
What readers can do now
- If you build hardware, treat sub-15W inference as a real option this year: evaluate an M.2-class in-memory accelerator for always-on agent features before defaulting to a cloud API.
- If you follow China's chip sector, track compute-in-memory shipments and design wins (PC makers, robotics, edge gateways) as a leading indicator of whether alternative architectures are escaping the lab.
- If you care about privacy, watch for "data never leaves the device" claims and verify they are tied to on-device models of at least tens of billions of parameters, not just tiny classifiers.
Honest limitations
Figures such as 160 TOPS at 10W, 48 GB, and 153.6 GB/s are company-disclosed specifications, not independently audited benchmarks; real-world throughput depends on model, batch size, and software stack. Mass-production timing and customer deployments are reported by the vendor and Chinese technology media, with limited independent verification from international outlets. This article covers one SRAM-based route and does not assess flash-, ReRAM-, or MRAM-based competitors head-to-head. Performance claims for specific partner devices (e.g., tokens-per-second on named AI hosts) are vendor-stated and should be treated as marketing figures until third-party testing appears.
