NeuroAI NEUROAINEUROAI.SITE
ESC

Kuaishou's Kling became the AI video model that out-ranked Google Veo and Pika

Kling AI (可灵AI), the video-generation large model (大模型) built inside short-video platform Kuaishou (快手), launched in June 2024 and by early 2025 was topping independent leaderboards for image-to-video generation — ahead of Google Veo 2 and Pika. Its bigger milestone was commercial: it turned AI video from a free demo into a paid product line with annualized revenue in the hundreds of millions of dollars.

2026-09-28 · 1008 words · NeuroAI
Kuaishou's Kling became the AI video model that out-ranked Google Veo and Pika

For most of 2024, AI video was a parade of impressive clips that nobody could actually use for work. A model from inside China's Kuaishou changed that — not by generating the prettiest demo, but by shipping something stable enough that people would pay for it.

Kling AI (可灵AI) is now the clearest example of a Chinese generative-video product that competes head-on with the best American labs and wins on at least one public yardstick.

The model that shipped before the hype cooled

Kling launched in June 2024 as, by Kuaishou's account, the first publicly usable real-image-level video-generation model — built on a 3D spatiotemporal joint-attention architecture, supporting 1080p resolution and videos up to two minutes long. That was early. OpenAI's Sora was still not broadly available, and most rivals were shipping short, unstable clips.

What set Kling apart was consistency. By the time it reached its 2.0 generation, the model was producing video with coherent motion and characters that held together across a shot — the dull, unglamorous property that actually matters for production use.

The 2.0 leap and the leaderboard result

On 15 April 2025, at a Beijing launch event called "Inspiration Made Real (灵感成真)", Kuaishou released the Kling 2.0 video model and the Kolors 2.0 image model globally. The company said 2.0 led on motion quality, semantic responsiveness and visual aesthetics, and introduced a new interaction idea called Multi-modal Visual Language (MVL) — letting users combine text with reference images or video clips to specify identity, style, scene and camera movement precisely.

The external validation landed a few weeks earlier. On 27 March 2025, the benchmarking group Artificial Analysis ranked Kling 1.6 Pro first in the image-to-video (图生视频) Arena ELO at a score of 1,000, ahead of Google Veo 2 and Pika. For a Chinese model to top that particular global leaderboard was a genuine marker.

By April 2025, Kuaishou reported Kling had over 22 million global users, a 25× growth in monthly actives over the ten months since launch, and more than 20 model iterations. It had generated 168 million videos and 344 million images, with over 15,000 developers and enterprise customers on its API.

Why "image-to-video" is the battleground

Kling's strength in image-to-video is not accidental. Starting from a reference image forces the model to preserve identity and layout — the failure mode that makes most AI video unusable for branding, e-commerce and film prep. Kuaishou said image-to-video already accounted for the large majority of Kling's video creation, which is why the 2.0 Master Edition leaned hard into controllable editing: adding, removing or swapping elements inside a generated clip.

Enterprise adoption followed the quality. Kuaishou has named partners across advertising, film, animation and gaming — including Xiaomi, Amazon Web Services, Alibaba Cloud, Freepik and BlueFocus — using Kling's API in production workflows.

From free toy to paid infrastructure

The part that separates Kling from most AI-video experiments is revenue. In its March 2025 disclosure, Kuaishou said Kling's annualized revenue run rate (ARR) had surpassed US$100 million by March 2025 — the tenth month after launch. The company also said monthly subscription bookings exceeded RMB 100 million (≈ US$14M / HK$109M) in both April and May 2025.

By January 2026, Kuaishou reported that Kling's December 2025 monthly revenue exceeded US$20 million, implying an ARR of US$240 million, serving over 60 million creators and 30,000+ enterprise users with more than 600 million videos generated. For the full year 2025, Kling's revenue was reported at roughly RMB 1.04 billion (≈ US$150M / HK$1.13B) — a real, audited-scale business line, not a research project.

Where it still struggles

Kling is not flawless, and Kuaishou says so openly. The company's own vice president has noted that stability of generated content and the precise transmission of complex creative intent remain "significant challenges." Longer, narratively coherent videos, reliable physics, and fine control over multi-character scenes are still uneven. And the economics are demanding: running frontier video generation at scale is expensive, which is why monetization speed matters as much as model quality.

What readers can do now

  • If you produce content, test Kling's image-to-video workflow for product, ad or storyboard drafts where identity consistency matters. Start from a reference image and compare against Runway or Veo on your own material, not on vendor reels.
  • If you build with APIs, Kling's API is a credible non-US option for image-to-video at scale; evaluate latency, quota and the MVL input format before committing a pipeline.
  • If you follow the sector, watch ARR growth versus compute cost. Kling's story is proof that AI video can be a business — but only if revenue scales faster than the GPU bill.

Honest limitations

Core facts (Kuaishou Technology, HKEX 01024 / 81024; Kling AI 可灵AI launched June 2024 as a publicly usable real-image-level video model on a 3D spatiotemporal joint-attention architecture, 1080p, up to 2 minutes; Kling 2.0 and Kolors 2.0 released 15 April 2025 at the "Inspiration Made Real (灵感成真)" Beijing event; Artificial Analysis ranked Kling 1.6 Pro #1 in image-to-video Arena ELO at 1,000 on 27 March 2025 ahead of Google Veo 2 and Pika; >22M users, 25× MAU growth, 168M videos / 344M images, 15,000+ API customers by April 2025; ARR >US$100M by March 2025 and >US$240M by December 2025; RMB 100M monthly bookings in April–May 2025; ~RMB 1.04B full-year 2025 revenue; 60M creators, 600M videos, 30,000+ enterprise users by December 2025; MVL interaction concept; partners including Xiaomi, AWS, Alibaba Cloud, Freepik, BlueFocus) are drawn from Kuaishou's official investor-relations releases (PRNewswire / ir.kuaishou.com, 5 June 2025 and 13 January 2026), GlobeNewswire/Kuaishou IR Chinese releases, Artificial Analysis' published leaderboard, and corroborating Chinese financial media (上海证券报, 中国日报, 环球网). RMB figures (RMB 100M monthly bookings; RMB 1.04B full-year revenue) are converted at roughly 7.2 RMB/USD and 1.09 HKD/RMB. Leaderboard rankings and benchmark scores are third-party assessments, not Kuaishou's own claims, but Kuaishou cites them. No political or ideological content is included. Facts are current to 28 September 2026 and this article is not investment advice.

Related coverage

More in “Generative Media” → · Back to home · Markdown version