---
title: "Baidu's ERNIE 5.0 Packs 2.4 Trillion Parameters Into One Native Multimodal Model"
date: 2026-10-03
category: Foundation Models
site: NeuroAI
canonical: https://neuroai.site/a/na-model-baidu-ernie5
language: en
---

# Baidu's ERNIE 5.0 Packs 2.4 Trillion Parameters Into One Native Multimodal Model

> Released 22 January 2026, Baidu's ERNIE 5.0 (文心 5.0) is a 2.4-trillion-parameter native omni-modal large model (大模型) that fuses text, image, audio, and video in one autoregressive backbone, activating under 3% of parameters per token.

Most multimodal models bolt a vision encoder onto a language brain after the fact. Baidu bet the opposite — build one model that learns text, image, audio, and video from the same scratch. Whether that bet pays off depends on what you ask it to do.

## What actually shipped

On **22 January 2026**, Baidu released **ERNIE 5.0 (文心 5.0)** in formal, production form. Its headline number is **2.4 trillion parameters**, trained from scratch as a single autoregressive large model (大模型) that handles text, image, audio, and video in one framework — what Baidu calls "native omni-modal" (原生全模态). Instead of late-fusion (gluing a vision tower onto a language model), ERNIE 5.0 fuses the modalities during pretraining, so cross-modal features interact deeply from the start.

Under the hood is an ultra-sparse mixture-of-experts backbone with modality-agnostic routing. Baidu says **fewer than 3% of parameters activate per token** — a 2.4-trillion-parameter model that, at inference, fires only a sliver of itself. That is how it stays usable despite the size.

## Why "native" is the argument

The late-fusion approach dominates the industry because it is easy: take a strong text model, strap on a vision encoder, declare victory. Baidu argues that path caps multimodal quality. ERNIE 5.0's unified token space and "Next-Group-of-Tokens Prediction" objective are meant to let, say, a video frame and a sentence share the same reasoning path.

In a live demo, Baidu fed the model a tutorial video of someone rebuilding a food-delivery app and it decomposed the steps and emitted runnable front-end code. In a creative task it mimicked a *Dream of the Red Chamber* character's voice to draft a business plan. Whether those parlor tricks survive enterprise load is the open question.

## What the benchmarks claim

On **15 January 2026**, the preview build **ERNIE-5.0-0110** scored **1,460 on the LMArena text leaderboard**, which Baidu says placed it first among Chinese models and eighth globally — ahead of GPT-5.1-High and Gemini-2.5-Pro on that snapshot. LMArena re-cuts constantly, so treat the ranking as a dated datapoint. Baidu also claims 40-plus benchmark wins across language and multimodal understanding versus Gemini-2.5-Pro and GPT-5-High.

ERNIE 5.0's agentic side comes from end-to-end reinforcement learning coupling reasoning with action and tool use, plus a "mentor" program of 835 domain experts who calibrate the model.

## The reasoning sibling: ERNIE X1

Baidu's reasoning track is **ERNIE X1**, a "deep-thinking" model that extends internal chains of thought before answering and can call tools (search, code interpreter) mid-reason. Baidu launched X1 claiming performance comparable to DeepSeek R1 at roughly half the price — a cost-reduction pitch that predates and frames the 5.0 era. X1.1 followed in September 2025 with a 64K context and tighter factuality.

## From model to product

ERNIE 5.0 is reachable through Baidu's Qianfan platform and the Wenxiaozhu (文心一言) app; Baidu said its assistant's monthly active users had passed **200 million** before launch. A separate **ERNIE-Image** model (8B DiT, open weights) shipped 15 April 2026, and **ERNIE 5.1** followed on 9 May 2026 at roughly a third of 5.0's parameters and about 6% of its pretraining compute — Baidu's efficiency play.

Qianfan list pricing for ERNIE 5.0 runs about **US$0.89 per million input tokens and US$3.54 per million output**; ERNIE 5.1 is cheaper at **US$0.59 / US$2.65**.

## What Baidu's flagship says about the model race

ERNIE 5.0 is less a single product than a statement of where Baidu is betting. A native omni-modal flagship plus a cheaper 5.1 sibling is the same two-tier play every major Chinese lab now runs: one model for the capability demo, one for the cost-sensitive customer. Baidu's edge is distribution — its search, cloud, and enterprise accounts give ERNIE a built-in audience that pure labs lack. The open-weight ERNIE-Image and the mentor program show it still wants developer mindshare, not just enterprise contracts. For readers, the useful read is to ignore the leaderboard snapshot and watch whether ERNIE's omni-modal claim survives contact with real mixed-media workloads, where late fusion usually breaks.

## Honest limitations

The 22 January 2026 release, 2.4T parameters, native omni-modal architecture, <3% activation, the 15 January LMArena score of 1,460 (No.1 China / No.8 global), the 200M MAU, the 835-expert mentor program, ERNIE-Image (15 April 2026, 8B, open weights), and ERNIE 5.1 (9 May 2026) come from Baidu's official ERNIE blog, stcn, cnstock, and Baidu Baike. The **"40-plus benchmark wins vs Gemini-2.5-Pro / GPT-5-High"** claim is Baidu's and unverified independently. The **LMArena 1,460 ranking** is a leaderboard snapshot that shifts constantly. ERNIE X1's **"comparable to DeepSeek R1 at ~half price"** is Baidu's launch positioning, not an audited comparison. The **US$0.89 / US$3.54** and **US$0.59 / US$2.65** prices are from model aggregators (Qianfan list) and vary by tier/region. No renminbi figure required conversion here; all cited pricing is in USD per published aggregator data.

## What readers can do now

- **If you need one model for text+image+audio+video:** pilot ERNIE 5.0 on Qianfan and test the native-omni claim on mixed-modality tasks your late-fusion stack fumbles.

- **If cost is the constraint:** start on ERNIE 5.1 (1/3 params, 6% pretrain cost) and reserve 5.0 for jobs that need the full omni-modal range.

- **If you benchmark:** re-run LMArena-equivalent tasks on today's cut before trusting the January 1,460 snapshot; leaderboards move weekly.

---

Published by NeuroAI (https://neuroai.site/) — https://neuroai.site/a/na-model-baidu-ernie5
Free to quote with attribution and a link to the original.
