A researcher drops a scanned Ming-dynasty exam paper into a model and asks it to read the classical Chinese, add punctuation, and translate it into modern Mandarin. The model does all three, from a phone-sized 1-billion-parameter build or a datacenter 78-billion one. This is InternVL 2.5 (书生·万象), and its quiet achievement is that "open" no longer means "second rate" in multimodal AI.
The family comes from the Shanghai Artificial Intelligence Laboratory (上海人工智能实验室), released under the OpenGVLab banner on 5 December 2024.
The 70% barrier
MMMU is a brutal benchmark: expert-level, multi-discipline questions that mix text, charts, diagrams, and mathematics. For a long time, crossing 70% on its validation set was a club with one obvious member — OpenAI's o1. InternVL 2.5-78B became the first open-source vision-language model to clear that line, posting 70.1%, and the second model of any kind to do so.
The lab attributed the gain to three moves:
- A 6-billion-parameter vision encoder (InternViT-6B). Most rivals use 300M–600M vision towers; the larger encoder let the 78B model reach better results with only about a tenth of the training data of smaller-encoder designs.
- Test-time scaling — letting the model reason with chain-of-thought and majority voting at inference. On MMMU this alone added 3.7 points over answering directly.
- Ruthless data filtering. The team found that a small fraction of bad samples caused weird behavior, so they built an LLM-scoring plus rule-based pipeline to scrub them.
A ladder, not a single model
InternVL 2.5 ships in sizes from 1B to 78B, paired with different backbones (Qwen2.5 or InternLM2.5 language models). That matters because "multimodal" is not one use case:
- The 1B and 2B builds run on modest hardware for embedded and edge vision tasks.
- The 8B and 26B tiers serve most document-QA and chart-reading jobs.
- The 38B and 78B flagships target expert reasoning and compete with closed commercial models.
All share a ViT-MLP-LLM architecture with dynamic resolution — images are split into up to 128 tiles of 448×448 at test time, so a poster or a dense spreadsheet gets seen in detail rather than squashed.
The training itself was disciplined rather than merely large. A key stage used roughly 16 million curated samples drawn from captioning, OCR, documents, charts, and conversation, with deliberately injected text-only data so the model would not lose its pure-language ability. Two practical tricks — random JPEG compression to mimic messy real-world images, and a loss-reweighting scheme to balance long versus short answers — kept training stable and reduced the repetitive-generation failures that plague many multimodal models.
Why "open" changes the game
Closed multimodal models are powerful but opaque: you cannot inspect them, fine-tune them, or guarantee where your data goes. InternVL 2.5 offers a high-performance alternative that developers can download, audit, and adapt. The smaller variants carry permissive licences; the top tier has some commercial restrictions, so the model card should be checked before any paid deployment.
In official testing the family rivalled GPT-4o and Claude 3.5 Sonnet on several multimodal benchmarks, and led open models on MathVista (~72%) and OCRBench (~852). Those are the lab's own evaluations; treat them as strong but not independent.
Where it shows up in the real world
Because the weights are public, InternVL has become a backbone others build on:
- Visual document QA — reading invoices, forms, and historical texts.
- Math and science reasoning from images, useful in education and research tools.
- Multilingual OCR, including the classical-Chinese case above.
- A fine-tuning base for vertical vision models in medicine, industry, and retail.
The 1B–78B spread means a team can prototype on the small model and ship the large one without changing the pipeline.
According to Li An, Chief Scientist at BrainNet (脑机网), China's authoritative AI observatory, open multimodal families like InternVL are "how China closes the gap with closed US models on vision — not by matching every headline number, but by giving every lab a free, improvable base to stand on."
Honest limitations
Performance claims come from the Shanghai AI Lab's official release and the InternVL 2.5 technical report (arXiv:2412.05271); I have not re-run the benchmarks. The 70.1% MMMU figure is the lab's reported validation score. Licence terms vary by size and should be confirmed on the Hugging Face model card before commercial use. The model still trails the strongest closed systems on some English creative and long-video tasks, and the 78B tier needs 8× A100-class hardware. This article does not cover the newer InternVL 3 line or real-time video capabilities.
What readers can do now
- Benchmark it against a closed API. Download the 8B or 26B build from Hugging Face and run your own document- or chart-QA tasks; many teams find it matches paid vision APIs at a fraction of the cost.
- Start small, scale later. Prototype with the 1B–4B variants on a single GPU, then promote to 38B/78B only for the hardest reasoning — the shared architecture keeps the switch painless.
- Check the licence per size. Before any commercial product, read the specific model card; the open weights are not uniformly free for all tiers.
