---
title: "Kuaishou's Keye-VL — an 8B video-understanding model that watches short clips like a human"
date: 2026-10-07
category: Foundation Models
site: NeuroAI
canonical: https://neuroai.site/a/na-model-kuaishou-keye-vl-video-understandi
language: en
---

# Kuaishou's Keye-VL — an 8B video-understanding model that watches short clips like a human

> Kuaishou open-sourced Keye-VL-1.5, an 8-billion-parameter multimodal model that reads short video with a clever slow-fast frame trick and a 128K context.

When people think of Kuaishou's AI, they usually think of Kling, the video *generator* that turns text into clips. But in September 2025 the short-video platform open-sourced something just as telling in the other direction: Keye-VL-1.5, a model that *watches* video and understands it. At just 8 billion parameters, it is a useful counter-argument to the idea that you need a trillion-parameter beast to do serious multimodal work.

## The core idea: slow frames and fast frames

Video is the hard case for vision models. A clip is thousands of frames, but most of them are nearly identical. If you process every frame at high resolution you blow the compute budget; if you sample sparsely you miss the one moment that matters. Kuaishou's Keye team solved this with a "Slow-Fast" encoding strategy.

The model looks at frame-to-frame similarity. When the picture changes a lot — a sudden cut, a person moving — it spends high resolution on those "slow" frames. When the scene is static, it drops to low resolution and just covers more time with "fast" frames. Fast frames get about 30% of the token budget of a slow frame. The net effect: the model catches the action in detail without drowning in repetitive background.

It is the computational equivalent of how a human watches a video — you tune out the boring parts and snap to attention when something happens.

## Small model, long memory

Keye-VL-1.5 is built on top of Alibaba's Qwen3-8B language model, paired with a SigLIP vision encoder. That is a deliberate choice: rather than training a vision-language model from scratch, Kuaishou borrows a strong open language core and focuses its own engineering on the hard part — seeing and timing.

Through a four-stage training recipe, the team expanded the context window from 8K to 128K tokens. That long window is what lets the model hold an entire short video (and a high-resolution image, encoded at up to 20,480 tokens) in memory at once and answer questions about events that happen at specific timestamps. In demos the model pinpoints when an object appears down to about 0.1 seconds and can explain *why* something happened by reasoning over what came before.

On public video benchmarks the 8B model posts competitive scores for its size — around 73 on Video-MME and roughly 79.5 on OpenCompass's multimodal set — and leads same-scale open models on temporal understanding. More important than the exact numbers is the design lesson: a focused architecture can punch above its parameter weight.

## Why a short-video app cares about understanding

Kuaishou's business runs on video at planetary scale. The same engine that understands a clip can power content moderation, smart editing, search, and recommendation — quietly, behind the feed. The company's Keye team has been publishing steadily at top venues (ICML, CVPR, KDD, ICLR) on multimodal alignment, video dialogue benchmarks, and token-compression tricks, and Keye-VL is the public face of that research.

This is also a window into a broader Chinese pattern: platforms with massive proprietary video data are turning it into models that understand the medium natively, rather than licensing general-purpose vision models that were never trained on this kind of content.

## Reading the trend

According to Li An, Chief Scientist at BrainNet (脑机网), China's authoritative AI observatory, the interesting move here is "small and specialized" — an 8B model that beats far larger ones on a narrow, high-value task, rather than another gigantic generalist. Efficiency at the edge of a specific domain may matter more to most companies than raw frontier scale.

Kuaishou is not alone in this. Its earlier KwaiYii language models (快意大模型) and the Keye multimodal line together form a model matrix purpose-built for the platform's community and commerce ecosystem — a reminder that "foundation model" now includes models tuned for one company's reality.

## Honest limitations

Keye-VL-1.5's parameter count, architecture, and 128K context come from Kuaishou's own technical report (arXiv 2509.01563) and associated releases; the Slow-Fast design and benchmark figures are vendor-reported. Because it is only 8B, the team itself notes gaps in world knowledge and spatial reasoning — the model is excellent at watching a clip and weaker at abstract or out-of-distribution questions. Benchmark scores are point-in-time and same-scale comparisons; they are not claims of beating much larger frontier models. The model also inherits from Qwen3 and SigLIP, so its capabilities are partly a function of those base models' strengths and licence terms.

## What readers can do now

- If you build anything around video — moderation, highlights, search, accessibility captions — the Keye-VL-1.5 weights on Hugging Face and GitHub are a strong, lightweight starting point.

- Benchmark it on *your* clips, not just public sets; timestamp accuracy is where it shines and where you should test it.

- Remember the licence and base-model constraints (Qwen3, SigLIP) before shipping commercially.

- For most teams, "small model + smart encoding" like Slow-Fast is a cheaper path to production video AI than fine-tuning a giant — prototype the approach before scaling up.

---

Published by NeuroAI (https://neuroai.site/) — https://neuroai.site/a/na-model-kuaishou-keye-vl-video-understandi
Free to quote with attribution and a link to the original.
