NeuroAI NEUROAINEUROAI.SITE
ESC

Kuaishou's Kling 3.0 can now put words in a character's mouth — and it's already a US$500M business

Kuaishou's Kling (可灵) AI shipped its 3.0 model family on 5 February 2026 with native audio, 15-second clips and reference-based consistency, and the unit passed US$500M in annualized revenue by March as enterprise and creator adoption scaled.

2026-09-29 · 737 words · NeuroAI
Kuaishou's Kling 3.0 can now put words in a character's mouth — and it's already a US$500M business

For years, AI video looked like a silent film: gorgeous frames with no voice, no lip-sync, no consistent character across shots. Kuaishou's Kling (可灵) changed that chapter on 5 February 2026, when it launched the 3.0 family with audio generated in the same model that makes the pictures. Behind the demo is a business already running at a US$500M annualized pace.

The launch that added sound

Kling AI 3.0 arrived via Kuaishou's investor relations (5 February 2026, GlobeNewswire) as four models — Video 3.0, Video 3.0 Omni, Image 3.0 and Image 3.0 Omni. The framing is All-in-One (全模态): text, image, audio and video flow in and out of one multimodal architecture.

The upgrades that matter:

  • Native audio across English, Chinese, Japanese, Korean, Spanish and various accents and Chinese dialects; multi-character dialogue where each speaker uses a different language.
  • Up to 15-second video, enough for a real beat of narrative rather than a single gulp.
  • 2K and 4K image output for professional asset work.
  • Video 3.0 Omni extracts a character's look and voice from a reference clip and reuses them across new scenes, with a multi-shot storyboard you can direct shot by shot.
  • A "图生视频 + 主体参考" (image-to-video + subject reference) path that keeps one person or object coherent from frame to frame.

Why audio matters more than the pretty frames

Before 3.0, AI video lived in what Kuaishou itself calls the "silent era" (默片模式) bottleneck — you generated visuals, then dubbed or lip-synced separately, badly. Native Audio closes that loop end-to-end, with environment sound aligned to action frame by frame. It also keeps text inside images sharp (logos, captions), which is exactly what e-commerce advertisers need. The result is a shift from "generate a clip" to "run a production workflow."

The business behind it

The user base is already large. As of the February 2026 launch, Kuaishou said Kling had served:

  • Over 60 million creators worldwide.
  • More than 600 million videos generated.
  • Over 30,000 enterprise clients.

The money followed. In the first quarter of 2026, Kling's revenue topped RMB 650 million (≈ US$90M / HK$709M), up more than 300% year on year. By March 2026 its annualized revenue run-rate (ARR) reached nearly US$500M, four times the US$100M it recorded a year earlier, according to Kuaishou's HKEX filing. CEO Cheng Yixiao (程一笑) attributes the growth to a two-engine model: B-end enterprise API calls and P-end paid creator subscriptions.

The models are also leaking into real production. Kling supported virtual scenes and effects in the Chinese historical drama 太平年 (Taiping Year) and generated hundreds of shots — including large battle sequences — for the Hollywood series 大卫王朝 (House of David), at roughly a third of traditional production cost, per Kuaishou.

What it cannot do yet

The 15-second cap still forces long stories to be stitched together, and consistency, strong as it is, is not flawless across very long runs. Kuaishou's group profit actually dipped in Q1 (adjusted net income RMB 3.37 billion, down 26.3%) as it funds AI capital spending of about RMB 26 billion in 2026. And the field is crowded: ByteDance's Seedance 2.0, Runway and Google's Veo are all racing the same problem.

What readers can do now

  • If you are a creator or studio: pilot Kling 3.0 Omni for ad and microdrama (微短剧) pipelines where voice consistency and character locking are the hard part.
  • If you are an enterprise: use the API to localize one script into many languages with native audio, instead of dubbing after the fact.
  • If you track the sector: watch ARR and AI-microdrama (AI漫剧) ad spend — which peaked above RMB 20 million a day in March — as the real adoption signal, not viewer counts.

Honest limitations

Sources: Kuaishou's official IR release (ir.kuaishou.com / GlobeNewswire, 5 Feb 2026: 3.0 family, native-audio languages, 15s clips, 2K/4K, 60M creators / 600M videos / 30,000 enterprises), the HKEX filing (hkexnews.hk, 2026 Q1 results: Kling revenue >RMB 650M, +300%, March ARR near US$500M, up from US$100M in March 2025), and cnstock / stcn (Q1 Kling figures, 太平年 / 大卫王朝 use, AI-microdrama ad spend >RMB 20M/day peak in March). RMB converted at ≈7.2 RMB/USD and ≈1.09 RMB/HKD (RMB 650M ≈ US$90M / HK$709M). The "60M creators / 600M videos" are company disclosures as of the Feb 2026 launch. ARR is Kuaishou's own run-rate, not audited annual revenue, and the "one-third of traditional cost" Hollywood claim is Kuaishou's. This is not investment advice.

Related coverage

More in “Generative Media” → · Back to home · Markdown version