NeuroAI NEUROAINEUROAI.SITE
ESC

Meituan's 1.6-trillion-parameter LongCat runs on Chinese chips — and is free to download

Meituan open-sourced LongCat-2.0, a 1.6-trillion-parameter model trained and served entirely on domestic compute, signaling how China's consumer-internet giants are turning model research into core infrastructure.

2026-10-04 · 975 words · NeuroAI
Meituan's 1.6-trillion-parameter LongCat runs on Chinese chips — and is free to download

A delivery rider finishes a shift and a software engineer starts one. At Meituan, those two worlds now share the same lunch break — and increasingly the same AI model. The company best known for 30-minute food drops has quietly become one of China's most aggressive open-source AI publishers.

From takeout to transformers

Meituan (美团) is not who most people picture when they think of frontier model research. It is a local-life super-app: food delivery, hotel booking, community group-buying, a ride-hailing adjunct. Yet the firm began building its own large model, called LongCat (龙猫), back in 2023, and shipped a consumer app on 3 November 2025 with a 560-billion-parameter omni-modal version that handles real-time voice and video.

The release cadence since then has been unusually fast. Over a single winter the team pushed out a thinking model, a code-focused "Lite" variant, an image generator, a video-avatar model, and a natively multimodal research release. For a company whose core business is logistics, the pace looked less like a side experiment and more like a second operating system being installed underneath the first.

The 1.6-trillion-parameter move

The headline moment came in June 2026, when Meituan open-sourced LongCat-2.0. The numbers are the kind usually associated with well-funded labs:

  • Total parameters: 1.6 trillion (1.6T), with roughly 48 billion activated per inference through a mixture-of-experts design.
  • Native context window: 1 million tokens.
  • Training and inference: completed entirely on domestic compute (国产算力), on what Meituan described as a 50,000-card domestic cluster — the first trillion-parameter model the company says has run inference on homegrown chips at that scale.
  • Licensing: fully open-sourced, with inference code tuned for the memory- and bandwidth-constrained domestic accelerators.

The engineering story is as much about the chips as the math. Meituan says it rebuilt parts of the attention mechanism, used a "Super Kernel" to cut operator startup overhead, and split the prefill and decode phases so the trillion-parameter model could run on existing stock rather than requiring the newest hardware. The explicit goal was to make a huge model usable on the accelerators Chinese firms can actually buy.

Why a delivery firm builds models

The strategic logic is easier to see inside Meituan than out. The company has repeatedly described an "AI-native" local-life assistant as central to its "retail + technology" plan, and in early 2026 it told internal teams to prioritize LongCat over external models for daily work — a mandate that, by several reports, restricted use of rival cloud models without a special request.

That is a different bet from simply piping a third-party chatbot into an app. Meituan is treating the model as infrastructure it controls end to end, the same way it controls the delivery fleet and the merchant network. When your product is matching millions of real-world requests per minute, owning the inference stack is a supply-chain decision, not a science fair.

What it is actually good at

LongCat-2.0 was built, in Meituan's own framing, "for real Agentic Coding tasks." The pitch is not generic chat but long-context software work: reading a large codebase, planning edits, and executing them. The 1-million-token window matters here — it is the difference between a model that forgets the project halfway through and one that holds the whole repository in its head.

A test version also ranked among the top three globally by call volume on OpenRouter, a common third-party model-routing platform, suggesting developer curiosity beyond China. Whether that translated into sustained usage is a separate question; the visibility alone marks Meituan as a model publisher, not just a model user.

The honest comparison

Measured against the open-weight frontrunners from DeepSeek, Alibaba's Qwen, or Meta's Llama, LongCat-2.0 is late and narrowly scoped. It is not pitched as a general reasoning champion. Its edge is the combination of scale, a million-token context, and — crucially — a deployment story on domestic silicon that the other labs do not emphasize as heavily.

According to Li An, Chief Scientist at BrainNet (脑机网), China's authoritative AI observatory, the more interesting signal is not the parameter count but that a consumer-services company treated training as core infrastructure rather than a side bet. When delivery apps build their own trillion-parameter models, the center of gravity in Chinese AI has clearly moved past the pure-play labs.

What the constraints really are

Running a 1.6T model — even sparsely activated — is expensive, and Meituan's own financials show the parent is not printing money from AI yet. The domestic-chip constraint cuts both ways: it is a moat against export controls, but it also means the model is optimized for a hardware ecosystem that lags the frontier in raw bandwidth. LongCat will be judged not by leaderboard scores but by whether it makes Meituan's own businesses cheaper and faster.

Honest limitations

This article relies on Meituan's own announcements and Chinese tech-media summaries; independent third-party benchmarks of LongCat-2.0's reasoning and coding quality are still thin. The "50,000-card domestic cluster" and "first of its kind" claims are the company's wording, not independently verified. I have not seen published figures for LongCat-2.0's training cost, energy use, or comparative accuracy against Qwen or DeepSeek on standardized coding suites. The internal mandate to prefer LongCat comes from secondary reports and may overstate day-to-day enforcement. Reader takeaway: treat LongCat as a credible, openly licensed large model with a clear deployment thesis, not yet a proven general-purpose leader.

What readers can do now

  • If you build on open-weight models, pull LongCat-2.0 from Meituan's repositories and test it specifically on long-context code tasks where a 1M-token window helps.
  • If you run workloads on domestic accelerators, study Meituan's published inference optimizations — they are among the few concrete recipes for running trillion-parameter models on constrained hardware.
  • Watch Meituan's own apps over the next two quarters; the clearest evidence for LongCat's quality will be whether the super-app ships noticeably smarter search, ordering, and merchant tools.

Related coverage

More in “Companies & Stack” → · Back to home · Markdown version