---
title: "Baichuan open-sourced a 235B medical model (大模型) trained to question you like a doctor"
date: 2026-09-30
category: Foundation Models
site: NeuroAI
canonical: https://neuroai.site/a/na-baichuan-m3-medical-open
language: en
---

# Baichuan open-sourced a 235B medical model (大模型) trained to question you like a doctor

> On 13 January 2026, Baichuan AI (百川智能) released Baichuan-M3 as open weights — a 235-billion-parameter medical large model (大模型) it says leads the HealthBench benchmark. The notable design choice is a model that asks follow-up questions before it commits to an answer.

Most medical chatbots are built to answer fast. A Beijing lab made the opposite bet: its new model is trained to interrupt you with questions, the way a clinician does, before it will commit to a conclusion.

On 13 January 2026, Baichuan AI (百川智能), the company founded by former Sogou chief Wang Xiaochuan (王小川), open-sourced Baichuan-M3, a 235-billion-parameter medical large model (大模型). The weights are public on Hugging Face and GitHub, and the company says the model is meant for clinical decision support, not as a replacement for physicians.

## Why question-first, not answer-first

The headline feature is what Baichuan calls "serious consultation" (严肃问诊) — an end-to-end dialogue mode where the model actively collects information instead of guessing from a short prompt. Baichuan describes a "SCAN" principle behind it: safety stratification, clarity (information confirmation), association and inquiry, and a normative protocol. The point is to mirror how doctors work — take the history, rule things in and out, then reason.

That is a different product philosophy from a question-answering bot. Rather than "here is the likely diagnosis," the model is tuned to surface the missing history and risk signals first, then reason on a fuller picture. Baichuan says its own evaluations put the model's questioning ability above the average human-doctor baseline on the dimensions it measures.

The training approach matters here. Baichuan frames medical dialogue as a pipeline — initial inquiry, differential diagnosis, lab work, final diagnosis — and inserts a quality gate so the model cannot advance to the next stage until the current one meets a clinical standard. A new algorithm it calls SPAR (step-penalty advantage relative baseline) gives fine-grained reward signals at every step of a consultation, which the company says reduces "reward hacking" and stabilizes long, multi-turn conversations.

## What the benchmarks say

Baichuan reports that M3 scored **65.1** on HealthBench — a medical AI evaluation suite originally released by OpenAI — and **44.4** on the harder HealthBench Hard subset. The company says both are the highest published scores at the time and that M3 is the first medical model to beat GPT-5.2 on those tests. It also reports a medical hallucination rate of **3.5%**, which it describes as the lowest among comparable models.

Those numbers matter because hallucination — a model inventing a drug, a dose, or a fact — is the single largest barrier to trusting AI in medicine. Baichuan says it pushed fact-consistency into the training loop itself, using a dynamic verifier system rather than relying only on retrieval or post-hoc checks. The verifier decomposes a claim into atomic statements and validates them against a two-level cache.

## The lineage: from M2 to M3

M3 is not a one-off. Baichuan open-sourced an earlier medical model, M2, in August 2025, which it says already beat open models such as DeepSeek-R1 on HealthBench Hard. Over the following five months the team rebuilt the reinforcement-learning feedback system from a semi-static patient simulator into what it calls a fully dynamic verifier that evolves as the model improves, and added the SPAR algorithm for long-dialogue stability.

The company has framed medical AI as its core mission since its 2023 founding. Wang Xiaochuan has repeatedly said the long-term direction is life sciences and medicine, and Baichuan has kept its earlier medical checkpoints (M1, M2) open and commercially usable. That consistency is part of why M3 reads less like a demo and more like a research program with a stated destination.

## How you would actually use it

The weights are on Hugging Face at baichuan-inc/Baichuan-M3-235B and mirrored on GitHub. Baichuan's own assistant app, 百小应 (Baixiaoying), has been wired to M3 so both doctors and patients can use it — doctors to reason through a case, patients to understand a diagnosis, treatment, or prognosis.

Like most open models, M3 is positioned as a reasoning and decision-support layer. The company's own disclaimers are explicit that it is for research and reference, not a substitute for professional diagnosis or treatment.

## Honest limitations

- The benchmark scores (HealthBench 65.1, HealthBench Hard 44.4, 3.5% hallucination) are **Baichuan's own reported figures**, not independently reproduced results at the time of this writing. HealthBench is an OpenAI benchmark, and leaderboard positions shift as new models land.

- The "beats GPT-5.2" comparison is a vendor claim tied to specific test versions; treat it as directional, not settled.

- The exact open-source license terms for M3 were not re-confirmed from a primary legal source in this review. Baichuan has historically released medical models as open weights with commercial use, but readers planning commercial deployment should read the license file on the repository directly.

- This is a clinical-support tool. No open medical model should be used for autonomous diagnosis or treatment decisions.

## What readers can do now

- **Pull the weights and test them** on a domain task you understand — the model is on Hugging Face and GitHub, and 百小应 gives a no-code way to try the consultation flow.

- **Benchmark against peers before any deployment** — compare M3 with other open medical models on your own cases rather than trusting a single vendor leaderboard.

- **Keep a human in the loop** — if you work in healthcare, treat the output strictly as decision support and validate it against clinical guidelines and a qualified clinician.

---

Published by NeuroAI (https://neuroai.site/) — https://neuroai.site/a/na-baichuan-m3-medical-open
Free to quote with attribution and a link to the original.
