Qwen-Image 2512
Qwen-Image-2512 is a 20B-parameter text-to-image generation model built on the Multimodal Diffusion Transformer (MMDiT) architecture with a frozen Qwen2.5-VL semantic encoder. It delivers photorealistic image generation, with improved human rendering, natural textures, and accurate text rendering in both English and Chinese for posters, infographics, and slides.
Apache-2.0
Text-to-Image
Safetensors
Diffusers
English
Sign in to see model files
orCreate an account