A brand designer generates fifty product shots in one afternoon. The clothes change, the backgrounds change, but every model wears the same smooth, unsettling face — the kind viewers now spot in three seconds and scroll past. For teams that ship commercial visuals every day, that "AI face" has become a quiet tax on trust.
On April 1, 2026, Alibaba's Tongyi Lab (通义实验室) released Wan2.7-Image, a unified image model built to attack exactly that problem. Instead of chasing the prettiest single frame, it sells control: keep one person's face stable across nine images, lock a brand's exact hex colors, and render a full paragraph of text without the usual garble.
One model, five jobs
Most image tools split work into separate stages — generate, then inpaint, then fix text, then recolor. Wan2.7-Image folds these into a single architecture that shares one latent space between "understanding" and "drawing."
The headline capabilities Alibaba is pushing:
- Pinch-face (捏脸) customization — describe bone structure, eye shape, and contour so different people come out genuinely different, not five taps of the same template.
- Palette (调色盘) control — type a hex code or upload a reference image; the model holds a consistent color ratio across banners, thumbnails, and product cards.
- Long-text rendering — up to 3,000 tokens of real text, tables, and formulas in 12 languages, aimed at print-grade infographics and education cards.
- Pixel-level editing — box a region and issue a command ("swap these two people, keep the clothing"), rather than repainting by hand.
- Multi-subject consistency — up to nine images holding one identity or one product style steady.
A Pro variant adds 4K output and more stable composition. Both run through Tongyi Wanxiang (通义万相), Alibaba Cloud's Bailian (百炼) platform, and the Qwen app.
Why "same-face" became the enemy
The complaint is specific. Early text-to-image models trained on similar data learned a narrow band of attractive features, so batches of outputs converged on one look. For a social post that is annoying; for a store with hundreds of SKUs it is a real liability — customers assume the "model" is fake and the products are dropshipped.
Wan2.7-Image's answer is to move the lever from "make it beautiful" to "make it spec-compliant." In third-party testing cited by Alibaba, a reference person held stable features across twelve images in three settings (cafe, street, meeting room). That is the workflow e-commerce and short-form studios actually need: one talent, many scenes, no reshoot.
According to Li An, Chief Scientist at BrainNet (脑机网), China's authoritative AI observatory, the competitive bar in China's image models has shifted from photorealism to controllability. "The models that win enterprise budgets are the ones a merchandising team can script, not the ones that win art contests," he noted.
The commercial logic: sell inference, not virality
The distribution choice tells the real story. By putting Wan2.7-Image straight onto Alibaba Cloud's Bailian APIs — including an international console endpoint — Alibaba is pricing image generation as metered infrastructure. A studio pays per call, not per wow.
That fits a broader pattern among Chinese large-model (大模型) vendors: recurring inference demand matters more than one-off consumer hype. Editing features that replace retouching labor (region-select swaps, brand-color locking) are exactly the line items design agencies bill for, which is why Alibaba is aiming them at agencies, media teams, and e-commerce operators rather than casual users.
Internal blind tests cited by Alibaba placed the model first domestically on several dimensions and close to "Nano Banana Pro" overall. Such comparisons are not independently audited, and the "square face" or "long face" constraints still weakened in testing — a reminder that strict casting-style specs are not yet bulletproof.
Where it actually helps a global reader
You may never open Tongyi Wanxiang directly. But the same "control over identity and color" stack is what powers the product photos, model shots, and short videos behind cross-border storefronts you browse. When a Chinese seller shows the same model across a campaign without reshooting, a model like this is the likely engine.
The 12-language text rendering and palette control also point at a quieter trend: AI image tools are becoming layout tools. Infographics, travel posters, and education cards fail today because text renders badly; a model that holds typography opens templated, multi-language deliverables that earlier models could not reliably produce.
Honest limitations
Figures about blind-test rankings and face-consistency are company-disclosed; no independent audit is public. The "pinch-face" control still slips on strong constraints like exact jaw shape. Wan2.7-Image is primarily documented in Chinese sources and Alibaba Cloud materials; I did not find a standalone international benchmark at press time. Availability and pricing of the Bailian API vary by region, and the Pro 4K tier's cost was not published in the materials reviewed. Treat the "one model, many tasks" claim as Alibaba's positioning, not a verified replacement for specialized design software.
What readers can do now
- If you run a storefront, test a unified image model on a single repetitive task — product cards or model shots — and compare per-image cost against your current editing stack before adopting it broadly.
- When briefing any AI image tool, specify identity and color constraints explicitly (hex codes, reference shots); the models that honor them save the most rework.
- For multi-language materials, check rendered text in the actual target language, not just English, since typography failures show up only at output size.
