Ghost Palette

local/flux-2-klein-4b — ImageBench V1

ImageBench V1 —
192 evaluations across 6 categories
Benchmark V1 verdicts are produced by VLM judges and can contain mistakes. Treat PASS/FAIL labels as machine-assisted assessments, and inspect the images yourself. Learn more about the methodology.

Generation Details

Source-backed model context, size, cost, and request settings for this run.

MakerBlack Forest Labs
FamilyFLUX.2
Model size4B4B params
Costself-hostedfree
Latency4.5s median
Run targetlocal/flux-2-klein-4b
Effective requestmodel: flux-2-klein-4b · size: 1024x1024 · seed: 42
SourcesModel cardWeights & checkpoint
Text36%Spatial54%Human63%Pro Studio66%Graphical36%Truth36%Preference55%Latency85%

All 192 generations

Every challenge, grouped by category and subcategory. Solid-bordered tiles with a check passed the VLM judge; dashed tiles with a cross failed. Images are placeholders until generated.

Text Rendering

36%
Typography Style33%
Writing accuracy25%

Spatial Reasoning

54%
Attributes Binding33%
Compositionality44%
Counting56%
Negation44%
Relative Position50%
Scale & Proportions44%

Human realism

63%
Faces & Expressions50%
Full Body67%
Hands67%
Multi-Subject50%

Professional Studio

66%
Camera & Lighting50%
Color Precision58%
Photorealism67%

Graphical design

36%
Data Visualisation67%
Layout & Design33%
Style Diversity17%

Truthfulness

36%
Photorealism67%
Physics & Reflections33%
World Knowledge42%