Перейти к содержимому
Text → imageImage → image

Qwen Image

Posters, covers, and packaging with multilingual in-image text — English, Chinese, and more

About the model

Qwen Image is Alibaba’s multimodal model, with an architecture packing up to 20 billion parameters. Our service offers it in two modes at once: text-to-image for generating from scratch and image-to-image for controlled edits.

The model’s strong suit is accurate text rendering inside the image, including several languages at once. That makes it a solid pick for posters, covers, packaging, and any asset where typography plays as big a role as the illustration.

The image-to-image mode supports both semantic edits (pose changes, restyling) and pinpoint visual tweaks (adding or removing objects) while keeping the source’s overall composition intact.

Strengths

  • Accurate text rendering in multiple languages, including English and Chinese
  • Semantic edits in image-to-image mode (restyling, pose changes)
  • Preserves the integrity of the source image while editing
  • A wide range of styles — from photorealism to animation
  • Generation time of 8–20 seconds

Best for

  • Posters and covers with typography
  • Prototyping design concepts
  • Social media posts with multilingual captions
  • Branding and localization for different markets
  • Fast iterations on creative projects

Limitations

  • Complex multi-step edits may take several iterations
  • Output quality is sensitive to how detailed the prompt is

Prompting tips

  • Put in-image text in quotes, specifying the language and font style
  • Break long prompts into blocks (subject, background, typography, style)
  • For editing, describe what to keep first, then what to change
  • Handle complex edits as a series of short iterations