NeuroAI NEUROAINEUROAI.SITE
ESC

Tencent's Hunyuan-A13B — an open-weight reasoning model with a 256K context and a licence that bans EU use

Tencent released the weights of Hunyuan-A13B, an 80-billion-parameter mixture-of-experts model that activates only 13 billion per token. The catch is the licence — it excludes the EU, UK and South Korea and forbids distillation into other models.

2026-10-02 · 836 words · NeuroAI
Tencent's Hunyuan-A13B — an open-weight reasoning model with a 256K context and a licence that bans EU use

A model that fires only 13 billion of its 80 billion parameters per token sounds like a party trick. For Tencent, it is a deliberate bet that efficient open-weight reasoning, not raw size, is where Chinese large models (大模型) will compete.

On 27 June 2025, Tencent uploaded Hunyuan-A13B to Hugging Face. It is a mixture-of-experts (MoE) text model: roughly 80 billion total parameters, but only about 13 billion are active for any given token. That keeps the cost of serving it close to a much smaller dense model while still allowing serious reasoning on math, science and coding.

What "open weight" actually means here

Hunyuan-A13B is genuinely downloadable. Weights and Tencent's own GGUF quantizations sit on Hugging Face, and the code, technical report and licence live in the project's GitHub repository. Teams that do not want to run it themselves can call it through Tencent Cloud's API.

But "open weight" is not "open source" in the OSI sense. Tencent ships it under the Tencent Hunyuan Community License Agreement — a community licence with hard carve-outs:

  • Geography: the grant runs worldwide except the territory of the European Union, the United Kingdom and South Korea.
  • Commercial size: any licensee with more than 100 million monthly active users in the month before release needs a separate licence, granted at Tencent's sole discretion.
  • No distillation: the licence forbids using the model or its outputs "to improve any other AI model" other than Hunyuan itself.

For a self-hosted research or product team outside those three regions and under the user threshold, the model is usable. For a European startup or a giant consumer app, it is effectively off-limits without a separate deal.

The dual-mode reasoning pitch

Tencent markets Hunyuan-A13B as a "hybrid-reasoning" model. The idea is a fast response mode for routine queries and a slower, step-by-step reasoning mode for harder math, science and coding problems — letting a caller trade latency for depth on each request. This is the part Chinese labs are increasingly differentiating on, and it is the headline reason the release drew attention next to DeepSeek and Qwen.

It also ships a native long-context capability. Tencent's technical report describes a long-context training stage that scales the window from 32K to 256K tokens. In practice, the shipped configuration caps the context at 32,768 tokens by default; reaching 262,144 requires editing the configuration or passing a server flag. So "256K" is a real trained capability, not the out-of-the-box default — a distinction that matters for anyone sizing GPU memory.

Why efficiency, not size

The MoE design is the whole point. By activating only 13 billion parameters per token, Hunyuan-A13B delivers reasoning that tracks the active count, not the full 80 billion. That positions it against other efficient open releases — DeepSeek's smaller models and Alibaba's Qwen among them — competing on active-parameter efficiency rather than headline parameter count. The release is also part of a broader 2025 open-weight push from the Hunyuan team, which published smaller dense models such as Hunyuan-7B under the same community licence.

How to read the licence carve-outs

The territorial exclusion is worth a pause, because it reverses a usual pattern. Chinese open-weight releases are typically downloadable everywhere and restricted, if at all, by user scale. Carving out the EU, UK and South Korea suggests Tencent is managing regulatory exposure in jurisdictions with active AI-rulebook activity rather than simply limiting competition — a compliance-first posture that other Chinese labs may copy as AI legislation matures worldwide. For builders, the practical rule is simple: the licence, not the model card, defines where a deployment is legal, and checking it takes less time than a takedown notice costs.

Honest limitations

Quantitative claims (release date 27 June 2025; ~80B total / ~13B active parameters; 256K trained context with a 32,768-token shipped default; Hugging Face and Tencent Cloud availability; the licence's EU/UK/South Korea carve-out, 100-million-MAU commercial threshold, and no-distillation clause) come from Tencent's Hugging Face model card, the project GitHub repository and the licence text itself, cross-checked against third-party model indices. Benchmark scores are not independently reproduced here and are omitted rather than restated. Independent reviewers have questioned whether the shipped weights' fast/slow "thinking" toggle behaves exactly as documented; this article treats dual-mode reasoning as a designed feature per Tencent and does not claim verified benchmark superiority. The licence interpretation is read from the published licence text, not legal advice. No investment advice is given.

What readers can do now

  • If you run inference infrastructure, test Hunyuan-A13B as a cost-efficient reasoning option, but check the licence territory and the 100-million-MAU clause before any product launch — and budget the config edit if you need the full 256K context.
  • If you build with open models in the EU, UK or South Korea, treat Hunyuan-A13B as unavailable and compare Qwen, DeepSeek or similarly licensed alternatives instead.
  • If you track China's open-model strategy, read the licence carve-outs as a signal: Tencent is willing to open weights, but on terms that protect its own ecosystem and limit downstream model-training competition.

Related coverage

More in “Foundation Models” → · Back to home · Markdown version