Type a prompt in Chinese asking for a street sign, a book cover or a restaurant menu, and most Western image models quietly fail — they mangle the characters or invent plausible-looking gibberish. For the world's largest language community, that has been the quiet wall around generative image tools. A model that can actually write "北京" on a billboard is more than a party trick; it is the difference between usable and not.
That is the opening Kuaishou (快手) built with Kolors (可图), its text-to-image large model (大模型) that the company fully open-sourced at the World Artificial Intelligence Conference (WAIC) on July 6, 2024.
What makes Kolors (可图) different
Most open image models lean on an English CLIP text encoder, which is why they stumble on long, complex or Chinese-language prompts. Kolors took a different path:
- It uses the ChatGLM3 large language model for Chinese–English text representation, supporting prompts up to 256 characters — far beyond the 77-character limit of classic CLIP-based models.
- It is the first natively open model to render Chinese and English text inside images without extra control tricks.
- It is built on a U-Net latent-diffusion backbone trained on billions of text–image pairs, with a two-stage regime: concept learning, then quality refinement.
- It was released with weights and full code on Hugging Face and GitHub (and ModelScope), under an Apache-2.0 license, free for individual developers.
How the training was shaped
Beyond the ChatGLM3 text encoder, the Kolors team used the multimodal model CogVLM to regenerate richer descriptions for training images, which improved how faithfully the model follows detailed prompts. Training ran in two explicit phases — a concept-learning stage for broad entity coverage, then a quality-refinement stage using curated, high-aesthetic data and optimized high-resolution techniques. A purpose-built, category-balanced benchmark (KolorsPrompts) was built to steer evaluation rather than rely on generic scores. None of this is exotic, but the combination — a bilingual encoder, long-prompt support, native Chinese text rendering and a fully open release — is what made Kolors stand out among open image models.
The benchmark that got attention
On the BeiZhen (智源) FlagEval text-to-image leaderboard, Kolors posted a subjective comprehensive score of 75.23, ranking second globally and behind only the closed-source DALL·E 3. On image quality specifically, the company's own expert panel rated it at Midjourney-v6 level.
Those are vendor-furnished and benchmark-specific numbers, so read them as "competitive with the best open models and some closed ones," not as an absolute ranking. Still, for a model released openly and tuned for Chinese, the result reframed the open-source image landscape.
Why Chinese-language rendering is a real moat
Image generation is not only about pretty faces. The commercially useful cases in China — e-commerce banners, short-drama (短剧) posters, local ads, educational graphics, cultural content — all need correct, legible Chinese text. A model that gets the characters right, understands Chinese semantics, and handles long prompts removes a whole class of manual fixes.
Kolors has already been wired into Kuaishou products (AI playtests, camera effects, video-editing tools) and spawned practical add-ons like IP customization, AI portraits and virtual try-on. The open release let outside developers build ComfyUI wrappers and acceleration tools within days.
The open-weights (开源权重) angle
Kolors is part of a broader Chinese pattern: release a strong model's weights openly to seed a developer ecosystem. For researchers and small teams, that means a bilingual, text-rendering base model they can fine-tune without a US$ API bill. For Kuaishou, it is a talent and standards play as much as a product one.
Honest limitations
- The "Midjourney-v6 level" and FlagEval #2 claims come from Kuaishou and a single vendor-organized expert panel; we did not run independent benchmarks.
- Kolors is a 2024 release; the fast-moving image-model field has since advanced, so "second only to DALL·E 3" should be understood as a point-in-time result, not today's ranking.
- Benchmarks measure subjective preference and specific prompts; real-world performance on niche or adversarial prompts may differ.
- Open-weight use still carries content-safety and licensing obligations, especially for commercial deployment above the stated user threshold.
What readers can do now
- Try it directly — pull Kolors (可图) from Hugging Face or GitHub and test a Chinese prompt with on-image text; compare it to your current tool on character accuracy.
- Use it for localized design — if you make Chinese-language posters, ads or short-drama (短剧) thumbnails, prototype with Kolors before paying for a closed model.
- Fine-tune, don't just prompt — because the weights are open, small teams can adapt Kolors to a brand's visual style instead of relying on generic outputs.
