Ghost Palette

local/sefi-image-2b-turbo — ImageBench V1

ImageBench V1 —
192 evaluations across 6 categories
Benchmark V1 verdicts are produced by VLM judges and can contain mistakes. Treat PASS/FAIL labels as machine-assisted assessments, and inspect the images yourself. Learn more about the methodology.

Generation Details

Source-backed model context, size, cost, and request settings for this run.

MakerSefi
FamilySefi Image
Model size2B2B params
Costself-hostedfree
Latency3.6s median
Run targetlocal/sefi-image-2b-turbo
Effective requestmodel: sefi-image-2b-turbo · size: 1024x1024 · seed: 42
SourcesModel cardWeights & checkpoint
Text30%Spatial40%Human49%Pro Studio52%Graphical47%Truth30%Preference26%Latency88%

All 192 generations

Every challenge, grouped by category and subcategory. Solid-bordered tiles with a check passed the VLM judge; dashed tiles with a cross failed. Images are placeholders until generated.

Text Rendering

30%
Typography Style33%
Writing accuracy17%

Spatial Reasoning

40%
Attributes Binding44%
Compositionality33%
Counting33%
Negation22%
Relative Position58%
Scale & Proportions33%

Human realism

49%
Faces & Expressions50%
Full Body42%
Hands50%
Multi-Subject50%

Professional Studio

52%
Camera & Lighting33%
Color Precision50%
Photorealism33%

Graphical design

47%
Data Visualisation33%
Layout & Design22%
Style Diversity25%

Truthfulness

30%
Photorealism0%
Physics & Reflections17%
World Knowledge42%