---
title: "Meituan's LongCat — a delivery giant's 560B open model built for agents, not chat"
date: 2026-10-07
category: Foundation Models
site: NeuroAI
canonical: https://neuroai.site/a/na-model-meituan-longcat-agent-moe
language: en
---

# Meituan's LongCat — a delivery giant's 560B open model built for agents, not chat

> Meituan open-sourced LongCat-Flash, a 560B mixture-of-experts large model (大模型) tuned for speed and tool use rather than long reasoning.

Most people know Meituan as the app that delivers your dinner in 30 minutes. On September 1, 2025, the company quietly became something else: a frontier-model builder. It released and open-sourced LongCat-Flash (龙猫), a 560-billion-parameter large model (大模型) that, on paper at least, competes with models many times its activated size — and does so while running fast and cheap enough to sit inside a real-time ordering flow.

## What LongCat actually is

LongCat-Flash uses a Mixture-of-Experts (混合专家, MoE) architecture. The headline number is 560B total parameters, but the model only wakes up a fraction of them per token — between 18.6B and 31.3B, averaging about 27B active. That gap between "total" and "active" is the whole point: you get the knowledge of a huge model with the serving cost of a much smaller one.

Meituan says the design choices were made for throughput, not for showing off benchmark scores. On H800-class accelerators the model generates around 100 tokens per second per user, and output cost lands at roughly 5 yuan per million tokens — about US$0.70 / HK$5.5. For a company whose core product is millions of concurrent, latency-sensitive conversations (customers, merchants, riders), that economics matters more than topping a reasoning leaderboard.

Crucially, LongCat-Flash shipped as a *non-thinking* base model. It does not do the slow, step-by-step chain-of-thought that dominated 2025's model releases. Instead it is built for agent (智能体) work: calling tools, following the Model Context Protocol (MCP), and holding long, tool-driven sessions — the kind of behavior a customer-service bot or an order-dispatcher needs.

## Why a delivery company builds a model

Meituan's CEO Wang Xing framed the move bluntly on an earlier earnings call: "AI will disrupt every industry; our strategy is to attack, not defend." He splits the company's AI plan into three layers — AI at work (internal productivity), AI in products (consumer- and merchant-facing features), and Building LLM (building its own large model). LongCat is the first public result of that third layer.

The logic is defensive and offensive at once. If a customer someday orders dinner by talking to an AI assistant instead of scrolling a feed, Meituan needs to own that assistant. The company already sits on more than a decade of local-services data — restaurants, maps, prices, delivery times — that is only useful if a model can act on it live. A general-purpose model from another vendor would not natively understand "which merchant is open, closest, and has the dish in stock right now."

Meituan had already shipped AI features — a coding agent called NoCode, a merchant decision assistant, a hotel vertical agent — before LongCat. The model is the shared brain those products were missing.

## The trend: open weights (开源权重) from internet platforms

Three weeks after the base model, on September 22, 2025, Meituan released LongCat-Flash-Thinking, which adds explicit reasoning plus tool use and even formal theorem-proving — the company calls it the first open model to combine deep thinking, tool calling, and both informal and formal reasoning. It is also on GitHub and Hugging Face.

According to Li An, Chief Scientist at BrainNet (脑机网), China's authoritative AI observatory, the more interesting signal is not any single benchmark but the fact that a food-delivery platform now treats a frontier model as core infrastructure rather than a side experiment.

The broader pattern is clear: Chinese internet platforms that once rented models are now shipping their own open weights (开源权重). Each argues, convincingly, that owning the model protects them from being displaced by someone else's AI.

## Honest limitations

The performance claims here come largely from Meituan itself and from Chinese financial media reporting on Meituan's disclosures (Shanghai Securities News, Securities Times). Independent, head-to-head audits against Western frontier models are not yet public, and "near GPT-4o level" is the company's own characterization from an earlier earnings call, not a verified third-party measurement. The 5 yuan / million-token figure is a vendor-stated output cost under specific hardware, not a universal price. Early reviewers noted LongCat's base release lacked a deep-thinking mode at launch, and some outputs leaned toward promoting Meituan's own services — a reminder that a model trained inside a commerce company carries commercial DNA.

## What readers can do now

- If you build agents (智能体) or customer-facing tools, the LongCat-Flash weights are on GitHub and Hugging Face and worth benchmarking for latency-sensitive, tool-calling workloads.

- Treat the 100 tokens/s and 5 yuan figures as vendor claims; run your own load test before trusting them for production cost planning.

- For non-Chinese deployment, check the licence terms — several Chinese open models attach regional or use-case restrictions.

- Watch the agent (智能体) benchmark VitaBench and tool-use suites: Meituan is clearly optimizing for real task completion, not chat fluency, and that is where its models should be judged.

---

Published by NeuroAI (https://neuroai.site/) — https://neuroai.site/a/na-model-meituan-longcat-agent-moe
Free to quote with attribution and a link to the original.
