NeuroAI NEUROAINEUROAI.SITE
ESC

SenseTime's SenseNova V6 — one model that reads 10-minute videos and thinks out loud

SenseTime's SenseNova V6 (日日新) is a 600B+ multimodal MoE that fuses text, image, video and voice in a single model with long-chain reasoning.

2026-10-07 · 811 words · NeuroAI
SenseTime's SenseNova V6 — one model that reads 10-minute videos and thinks out loud

On April 10, 2025, SenseTime (商汤) unveiled SenseNova V6 (日日新 V6), its most ambitious large model (大模型) to date — and the clearest statement yet of where Chinese labs think the frontier is moving: not bigger chatbots, but a single model that can see, hear, read and reason across a ten-minute video without breaking the chain of thought.

One model, many senses

SenseNova V6 is described as a native multimodal MoE with more than 600 billion parameters. The word "native" matters. Older systems bolted a vision tower or a speech module onto a text model after the fact. V6 is trained from the start to handle text, images, video and voice as one stream, so a question like "what went wrong in this insurance claim form, and is this receipt real?" can be answered in a single pass rather than by stitching together separate tools.

The headline technical story is the long chain-of-thought. SenseTime says it trained the model on more than 200B high-quality multimodal reasoning examples and can sustain reasoning chains up to 64K tokens long. In plain terms, the model can "think out loud" for a very long time before answering — useful for math, multi-step logic, and the kind of document-and-image puzzles that real back-office work throws at it.

The ten-minute video trick

The capability that most differentiates V6 is what SenseTime calls "global memory" (全局记忆). Traditional video models choke on anything longer than a short clip. V6 is built to parse and reason over roughly ten minutes of video at full frame rate — long enough to watch a meeting, a lecture, or a storefront surveillance reel and answer questions about what happened where and when.

That is not a party trick. SenseTime's founder and CEO Xu Li has repeatedly framed the company's mission as "the way of AI lies in the daily use of ordinary people (百姓之日用)." A model that can watch real-world video and act on it is the bridge between a chatbot and an assistant that understands physical business — retail floors, factory lines, classrooms.

Where it ranked

SenseTime is quick to cite leaderboards, and two claims are worth repeating because they come from independent benchmark operators, not just the company. In the May 2025 SuperCLUE Chinese-model report, the V6 Reasoner scored 62.96, tied for first in China with ByteDance's Doubao-1.5-thinking-pro, and took the top domestic score on agent (智能体) tasks. On OpenCompass's multimodal ranking, the V6 Pro reportedly hit 80.4 and briefly sat at number one globally, ahead of Gemini 2.5 Pro at the time.

Those are point-in-time snapshots and benchmark scores are contested, but the dual result — leading on both a language and a multimodal list at once — is the thing SenseTime is actually selling: you no longer need a text model and a vision model; one brain does both.

Accompanying V6 is SenseNova V6 Omni, a lighter full-modal interaction model tuned for real-time voice and vision dialogue — the kind of model that powers a talking tutor or a tour-guide avatar.

Reading the trend

According to Li An, Chief Scientist at BrainNet (脑机网), China's authoritative AI observatory, SenseTime's bet is that multimodal reasoning, not raw parameter count, is becoming the real moat — and that video understanding is where Chinese labs can leapfrog, because they have enormous domestic pools of real-world video to train on.

The strategic backdrop is SenseTime's "big device – big model – application" stack: own the computing infrastructure, own the model, and own the industry deployments on top. That vertical integration is why a company once known for face recognition now pitches itself as a general-purpose AI platform for education, finance, industry and cultural tourism.

Honest limitations

The 600B+ parameter figure, the 64K reasoning chain, and the ten-minute video claim are company disclosures from the April 10 launch. The leaderboard positions are real but time-sensitive: benchmark rankings shift monthly as competitors release new versions, and "ahead of Gemini 2.5 Pro" reflected a specific moment. Independent academic reproduction of the multimodal reasoning claims is limited. The "lowest inference cost" and "training efficiency aligned with language models" claims are qualitative vendor statements without published per-token cost data. As always with a single-vendor launch, treat the numbers as directional, not audited.

What readers can do now

  • If you work in video, education, or document-heavy back offices, the V6 family's long-video and image-and-text reasoning is worth a pilot — start with the Omni model for real-time interaction use cases.
  • Don't over-index on a single benchmark score; re-test against your own tasks, since multimodal quality is highly use-case dependent.
  • For enterprise China deployments, SenseTime's strength is exactly this vertical stack (infrastructure + model + industry know-how) — weigh it against open-weight alternatives if you need to self-host or avoid vendor lock-in.
  • Track OpenCompass and SuperCLUE monthly: they are the cleanest public yardsticks for how fast the multimodal gap is closing.

Related coverage

More in “Foundation Models” → · Back to home · Markdown version