qwen-image-2512

Qwen-Image 2512

Qwen-Image-2512 is a 20B-parameter text-to-image generation model built on the Multimodal Diffusion Transformer (MMDiT) architecture with a frozen Qwen2.5-VL semantic encoder. It delivers photorealistic image generation, with improved human rendering, natural textures, and accurate text rendering in both English and Chinese for posters, infographics, and slides.

Apache-2.0
Text-to-Image
Safetensors
Diffusers
English
by @AIOZAI
56
0

Last updated: 22 days ago


Details
Files
Discussions
0

No discussions yet. Start the first one.

New Discussion