---
title: "Erase the lamp, extend the sky — China's multimodal image editors take on Photoshop's empty canvas"
date: 2026-10-06
category: Generative Media
site: NeuroAI
canonical: https://neuroai.site/a/na-aigc-multimodal-image-editing
language: en
---

# Erase the lamp, extend the sky — China's multimodal image editors take on Photoshop's empty canvas

> Domestic editors like Jimeng and Tongyi Wanxiang ship inpainting and outpainting in one click, challenging Photoshop and open-source FLUX workflows on simplicity.

A travel blogger uploads a photo taken through a hotel window and wants the air-conditioner gone, the sky widened, and a second version styled like ink wash. In Photoshop that is three tools and a tutorial. In a Chinese app it is three text boxes and a wait. The question is no longer whether the machine can edit, but whose editing language you already speak.

## What inpainting and outpainting actually do

Two capabilities define the current wave of image editing, and both are "multimodal" because they mix image and language input.

- **Inpainting** means selecting a region and telling the model what should be there instead — remove the tourist, add a hat, replace the background.

- **Outpainting** means extending the canvas beyond its original borders — widen a portrait to landscape, grow the room past the wall, continue the horizon.

Both were research curiosities until consumer tools made them one-click. The interesting competition is now between three lineages: Adobe's Photoshop with Firefly, the open-source FLUX ecosystem, and a cluster of Chinese editors shipping the same tricks inside everyday apps.

## Jimeng: editing baked into the creation app

Jimeng (即梦), built by the Jianying (剪映, the CapCut team in China) group at ByteDance, started life as Dreamina and was renamed in May 2024. It is less a photo editor than a creative workbench that happens to edit. Its image tools include local redraw (局部重绘), one-click expand (一键扩图, outpainting), image erase (图像消除, inpainting), and smart cutout, with output up to 8K resolution.

The design philosophy is integration. A user generates an image, drops it onto a layered canvas, fixes a finger or a stray object with the erase brush, then expands the frame to fit a poster ratio — without leaving the app or opening a desktop suite. For short-video creators who live in Jianying already, that continuity is the point; the editing is a step in a pipeline, not a separate craft.

## Tongyi Wanxiang: Alibaba's aesthetic engine

Alibaba's (阿里) Tongyi Wanxiang (通义万相, also called Wanx) takes the same capabilities and tilts them toward brand and cultural aesthetics. It supports image-to-image (图生图) and local editing, and reviewers single out its handling of Chinese visual styles — ink wash, embroidery, blue-and-white porcelain — where prompt-following on local detail tends to hold up.

Wanxiang is also where the open-source story connects. Alibaba's creative team released FLUX-Controlnet-Inpainting, an inpainting ControlNet checkpoint built for the FLUX.1-dev model, optimized around 768×768 inference. That matters because it puts a domestic Lab's inpainting weights directly into the hands of developers who run FLUX locally — the same community workflow that powers a large share of custom image editing outside the app stores.

## Baidu, Doubao, and the long tail

The field is wider than two names. Baidu's (百度) Wenxin Yige (文心一格) leans into Chinese-style illustration; ByteDance's Doubao (豆包) folds smart retouch and old-photo repair into a chat interface; smaller tools like Zutang (佐糖) and the Duiz reactor target one job — cutout, denoise, restore — and do it with near-zero learning curve.

The common thread is a Chinese-first prompt language. Describing "a rainy alley in Hangzhou at dusk, wet stone, lantern glow" tends to land more reliably in these tools than in systems trained primarily on English captions, which is the quiet advantage domestic editors hold for local creators.

## How they compare to Photoshop and FLUX

Photoshop's Generative Fill, powered by Adobe Firefly, remains the reference point for professionals. Its strengths are control, layers, and a copyright-clean training posture that enterprises trust. Its weakness is that it assumes you already think in Photoshop — panels, masks, selections.

The open-source FLUX path assumes you think in code or ComfyUI graphs. It is the most flexible and the least friendly: you wire nodes, manage checkpoints, and own the pipeline. FLUX-Controlnet-Inpainting is exactly this kind of power-user lever.

The Chinese consumer editors sit in the middle and win on friction. They trade raw control for a text box and a button. For the creator who wants the lamp gone and the sky wider, the domestic app is faster; for the retoucher who wants to govern every pixel, Photoshop or a FLUX node graph still wins.

## Where each one fits

- **Photoshop + Firefly** — production studios, brand-safe commercial work, pixel-level control.

- **FLUX + ControlNet inpainting** — developers and power users building custom or batch pipelines.

- **Jimeng / Wanxiang / Doubao** — creators, e-commerce merchants, and social-media teams who need good-enough edits in seconds.

None replaces the others outright. They are different points on the control-versus-speed line, and the line keeps sliding as each side borrows from the others.

## Honest limitations

Capabilities described here (Jimeng's local redraw, one-click expand, image erase, 8K; Wanxiang's local editing and Chinese-style strength; the FLUX-Controlnet-Inpainting checkpoint at ~768×768 from Alibaba's creative team) are drawn from vendor documentation, product pages, and developer repositories, which are authoritative for "what the tool claims" but not for independent quality benchmarks. No authoritative side-by-side test here measures output fidelity across Photoshop, FLUX, and the domestic editors on a common image set, so any "better or worse" claim is directional. Market-size and user-count figures circulating for China's image-tools space come from third-party trackers and are omitted here as unverified. Copyright, deepfake, and likeness-consent risks apply to all inpainting/outpainting tools and are governed by platform policy rather than solved by the technology.

## What readers can do now

- If you edit for social or e-commerce, run one real task (remove an object, extend a frame) through Jimeng and Wanxiang and note where the one-click result saves you versus where you still reach for Photoshop.

- If you build pipelines, test Alibaba's FLUX-Controlnet-Inpainting checkpoint against your current inpainting node to see whether the open weights match your fidelity bar before committing.

- If you publish edited images, keep a record of what was generated versus captured, because inpainting and outpainting quietly rewrite what a photo "shows" and that provenance now matters.

---

Published by NeuroAI (https://neuroai.site/) — https://neuroai.site/a/na-aigc-multimodal-image-editing
Free to quote with attribution and a link to the original.
