Ghost Palette

local/qwen-image-2512-20b — ImageBench V1

ImageBench V1 —
192 evaluations across 6 categories
Benchmark V1 verdicts are produced by VLM judges and can contain mistakes. Treat PASS/FAIL labels as machine-assisted assessments, and inspect the images yourself. Learn more about the methodology.

Generation Details

Source-backed model context, size, cost, and request settings for this run.

MakerAlibaba
FamilyQwen Image
Model size20B20B params
Costself-hostedfree
Latency80.2s median
Run targetlocal/qwen-image-2512-20b
Effective requestmodel: qwen-image-2512-20b · size: 1024x1024 · seed: 42
SourcesModel cardWeights & checkpoint
Text58%Spatial43%Human36%Pro Studio63%Graphical50%Truth49%Preference67%Latency0%

All 192 generations

Every challenge, grouped by category and subcategory. Solid-bordered tiles with a check passed the VLM judge; dashed tiles with a cross failed. Images are placeholders until generated.

Text Rendering

58%
Typography Style33%
Writing accuracy58%

Spatial Reasoning

43%
Attributes Binding33%
Compositionality33%
Counting56%
Negation33%
Relative Position25%
Scale & Proportions44%

Human realism

36%
Faces & Expressions33%
Full Body33%
Hands33%
Multi-Subject33%

Professional Studio

63%
Camera & Lighting42%
Color Precision67%
Photorealism33%

Graphical design

50%
Data Visualisation33%
Layout & Design44%
Style Diversity33%

Truthfulness

49%
Photorealism0%
Physics & Reflections33%
World Knowledge58%