Ghost Palette

local/flux-2-klein-9b — ImageBench V1

ImageBench V1 —
192 evaluations across 6 categories
Benchmark V1 verdicts are produced by VLM judges and can contain mistakes. Treat PASS/FAIL labels as machine-assisted assessments, and inspect the images yourself. Learn more about the methodology.

Generation Details

Source-backed model context, size, cost, and request settings for this run.

MakerBlack Forest Labs
FamilyFLUX.2
Model size9B9B params
Costself-hostedfree
Latency8.5s median
Run targetlocal/flux-2-klein-9b
Effective requestmodel: flux-2-klein-9b · size: 1024x1024 · seed: 42
SourcesModel cardWeights & checkpoint
Text63%Spatial58%Human56%Pro Studio60%Graphical44%Truth37%Preference60%Latency72%

All 192 generations

Every challenge, grouped by category and subcategory. Solid-bordered tiles with a check passed the VLM judge; dashed tiles with a cross failed. Images are placeholders until generated.

Text Rendering

63%
Typography Style33%
Writing accuracy42%

Spatial Reasoning

58%
Attributes Binding56%
Compositionality67%
Counting56%
Negation67%
Relative Position58%
Scale & Proportions67%

Human realism

56%
Faces & Expressions42%
Full Body50%
Hands58%
Multi-Subject50%

Professional Studio

60%
Camera & Lighting58%
Color Precision42%
Photorealism67%

Graphical design

44%
Data Visualisation33%
Layout & Design44%
Style Diversity33%

Truthfulness

37%
Photorealism33%
Physics & Reflections25%
World Knowledge25%