Most multimodal models bolt a vision encoder onto a language brain after the fact. Baidu bet the opposite — build one model that learns text, image, audio, and video from the same scratch. Whether that bet pays off depends on what you ask it to do.
What actually shipped
On 22 January 2026, Baidu released ERNIE 5.0 (文心 5.0) in formal, production form. Its headline number is 2.4 trillion parameters, trained from scratch as a single autoregressive large model (大模型) that handles text, image, audio, and video in one framework — what Baidu calls "native omni-modal" (原生全模态). Instead of late-fusion (gluing a vision tower onto a language model), ERNIE 5.0 fuses the modalities during pretraining, so cross-modal features interact deeply from the start.
Under the hood is an ultra-sparse mixture-of-experts backbone with modality-agnostic routing. Baidu says fewer than 3% of parameters activate per token — a 2.4-trillion-parameter model that, at inference, fires only a sliver of itself. That is how it stays usable despite the size.
Why "native" is the argument
The late-fusion approach dominates the industry because it is easy: take a strong text model, strap on a vision encoder, declare victory. Baidu argues that path caps multimodal quality. ERNIE 5.0's unified token space and "Next-Group-of-Tokens Prediction" objective are meant to let, say, a video frame and a sentence share the same reasoning path.
In a live demo, Baidu fed the model a tutorial video of someone rebuilding a food-delivery app and it decomposed the steps and emitted runnable front-end code. In a creative task it mimicked a Dream of the Red Chamber character's voice to draft a business plan. Whether those parlor tricks survive enterprise load is the open question.
What the benchmarks claim
On 15 January 2026, the preview build ERNIE-5.0-0110 scored 1,460 on the LMArena text leaderboard, which Baidu says placed it first among Chinese models and eighth globally — ahead of GPT-5.1-High and Gemini-2.5-Pro on that snapshot. LMArena re-cuts constantly, so treat the ranking as a dated datapoint. Baidu also claims 40-plus benchmark wins across language and multimodal understanding versus Gemini-2.5-Pro and GPT-5-High.
ERNIE 5.0's agentic side comes from end-to-end reinforcement learning coupling reasoning with action and tool use, plus a "mentor" program of 835 domain experts who calibrate the model.
The reasoning sibling: ERNIE X1
Baidu's reasoning track is ERNIE X1, a "deep-thinking" model that extends internal chains of thought before answering and can call tools (search, code interpreter) mid-reason. Baidu launched X1 claiming performance comparable to DeepSeek R1 at roughly half the price — a cost-reduction pitch that predates and frames the 5.0 era. X1.1 followed in September 2025 with a 64K context and tighter factuality.
From model to product
ERNIE 5.0 is reachable through Baidu's Qianfan platform and the Wenxiaozhu (文心一言) app; Baidu said its assistant's monthly active users had passed 200 million before launch. A separate ERNIE-Image model (8B DiT, open weights) shipped 15 April 2026, and ERNIE 5.1 followed on 9 May 2026 at roughly a third of 5.0's parameters and about 6% of its pretraining compute — Baidu's efficiency play.
Qianfan list pricing for ERNIE 5.0 runs about US$0.89 per million input tokens and US$3.54 per million output; ERNIE 5.1 is cheaper at US$0.59 / US$2.65.
What Baidu's flagship says about the model race
ERNIE 5.0 is less a single product than a statement of where Baidu is betting. A native omni-modal flagship plus a cheaper 5.1 sibling is the same two-tier play every major Chinese lab now runs: one model for the capability demo, one for the cost-sensitive customer. Baidu's edge is distribution — its search, cloud, and enterprise accounts give ERNIE a built-in audience that pure labs lack. The open-weight ERNIE-Image and the mentor program show it still wants developer mindshare, not just enterprise contracts. For readers, the useful read is to ignore the leaderboard snapshot and watch whether ERNIE's omni-modal claim survives contact with real mixed-media workloads, where late fusion usually breaks.
Honest limitations
The 22 January 2026 release, 2.4T parameters, native omni-modal architecture, <3% activation, the 15 January LMArena score of 1,460 (No.1 China / No.8 global), the 200M MAU, the 835-expert mentor program, ERNIE-Image (15 April 2026, 8B, open weights), and ERNIE 5.1 (9 May 2026) come from Baidu's official ERNIE blog, stcn, cnstock, and Baidu Baike. The "40-plus benchmark wins vs Gemini-2.5-Pro / GPT-5-High" claim is Baidu's and unverified independently. The LMArena 1,460 ranking is a leaderboard snapshot that shifts constantly. ERNIE X1's "comparable to DeepSeek R1 at ~half price" is Baidu's launch positioning, not an audited comparison. The US$0.89 / US$3.54 and US$0.59 / US$2.65 prices are from model aggregators (Qianfan list) and vary by tier/region. No renminbi figure required conversion here; all cited pricing is in USD per published aggregator data.
What readers can do now
- If you need one model for text+image+audio+video: pilot ERNIE 5.0 on Qianfan and test the native-omni claim on mixed-modality tasks your late-fusion stack fumbles.
- If cost is the constraint: start on ERNIE 5.1 (1/3 params, 6% pretrain cost) and reserve 5.0 for jobs that need the full omni-modal range.
- If you benchmark: re-run LMArena-equivalent tasks on today's cut before trusting the January 1,460 snapshot; leaderboards move weekly.
