---
title: "The open-source image model that finally understands Chinese characters"
date: 2026-10-01
category: Generative Media
site: NeuroAI
canonical: https://neuroai.site/a/na-aigc-kuaishou-kolors-opensource
language: en
---

# The open-source image model that finally understands Chinese characters

> Kuaishou's (快手) Kolors (可图) text-to-image model, open-sourced in July 2024, is bilingual, renders Chinese and English text inside images, and scored second only to DALL·E 3 on a major subjective benchmark.

Type a prompt in Chinese asking for a street sign, a book cover or a restaurant menu, and most Western image models quietly fail — they mangle the characters or invent plausible-looking gibberish. For the world's largest language community, that has been the quiet wall around generative image tools. A model that can actually write "北京" on a billboard is more than a party trick; it is the difference between usable and not.

That is the opening Kuaishou (快手) built with Kolors (可图), its text-to-image large model (大模型) that the company fully open-sourced at the World Artificial Intelligence Conference (WAIC) on July 6, 2024.

## What makes Kolors (可图) different

Most open image models lean on an English CLIP text encoder, which is why they stumble on long, complex or Chinese-language prompts. Kolors took a different path:

- It uses the **ChatGLM3** large language model for Chinese–English text representation, supporting prompts up to **256 characters** — far beyond the 77-character limit of classic CLIP-based models.

- It is the first natively open model to render **Chinese and English text inside images** without extra control tricks.

- It is built on a U-Net latent-diffusion backbone trained on billions of text–image pairs, with a two-stage regime: concept learning, then quality refinement.

- It was released with **weights and full code** on Hugging Face and GitHub (and ModelScope), under an Apache-2.0 license, free for individual developers.

## How the training was shaped

Beyond the ChatGLM3 text encoder, the Kolors team used the multimodal model **CogVLM** to regenerate richer descriptions for training images, which improved how faithfully the model follows detailed prompts. Training ran in two explicit phases — a concept-learning stage for broad entity coverage, then a quality-refinement stage using curated, high-aesthetic data and optimized high-resolution techniques. A purpose-built, category-balanced benchmark (KolorsPrompts) was built to steer evaluation rather than rely on generic scores. None of this is exotic, but the combination — a bilingual encoder, long-prompt support, native Chinese text rendering and a fully open release — is what made Kolors stand out among open image models.

## The benchmark that got attention

On the BeiZhen (智源) FlagEval text-to-image leaderboard, Kolors posted a subjective comprehensive score of **75.23**, ranking **second globally** and behind only the closed-source DALL·E 3. On image quality specifically, the company's own expert panel rated it at Midjourney-v6 level.

Those are vendor-furnished and benchmark-specific numbers, so read them as "competitive with the best open models and some closed ones," not as an absolute ranking. Still, for a model released openly and tuned for Chinese, the result reframed the open-source image landscape.

## Why Chinese-language rendering is a real moat

Image generation is not only about pretty faces. The commercially useful cases in China — e-commerce banners, short-drama (短剧) posters, local ads, educational graphics, cultural content — all need correct, legible Chinese text. A model that gets the characters right, understands Chinese semantics, and handles long prompts removes a whole class of manual fixes.

Kolors has already been wired into Kuaishou products (AI playtests, camera effects, video-editing tools) and spawned practical add-ons like IP customization, AI portraits and virtual try-on. The open release let outside developers build ComfyUI wrappers and acceleration tools within days.

## The open-weights (开源权重) angle

Kolors is part of a broader Chinese pattern: release a strong model's weights openly to seed a developer ecosystem. For researchers and small teams, that means a bilingual, text-rendering base model they can fine-tune without a US$ API bill. For Kuaishou, it is a talent and standards play as much as a product one.

## Honest limitations

- The "Midjourney-v6 level" and FlagEval #2 claims come from Kuaishou and a single vendor-organized expert panel; we did not run independent benchmarks.

- Kolors is a 2024 release; the fast-moving image-model field has since advanced, so "second only to DALL·E 3" should be understood as a point-in-time result, not today's ranking.

- Benchmarks measure subjective preference and specific prompts; real-world performance on niche or adversarial prompts may differ.

- Open-weight use still carries content-safety and licensing obligations, especially for commercial deployment above the stated user threshold.

## What readers can do now

- **Try it directly** — pull Kolors (可图) from Hugging Face or GitHub and test a Chinese prompt with on-image text; compare it to your current tool on character accuracy.

- **Use it for localized design** — if you make Chinese-language posters, ads or short-drama (短剧) thumbnails, prototype with Kolors before paying for a closed model.

- **Fine-tune, don't just prompt** — because the weights are open, small teams can adapt Kolors to a brand's visual style instead of relying on generic outputs.

---

Published by NeuroAI (https://neuroai.site/) — https://neuroai.site/a/na-aigc-kuaishou-kolors-opensource
Free to quote with attribution and a link to the original.
