local/sefi-image-2b-turbo — ImageBench V1
ImageBench V1 —
192 evaluations across 6 categoriesBenchmark V1 verdicts are produced by VLM judges and can contain mistakes. Treat PASS/FAIL labels as machine-assisted assessments, and inspect the images yourself. Learn more about the methodology.
Generation Details
Source-backed model context, size, cost, and request settings for this run.
MakerSefi
FamilySefi Image
Model size2B2B params
Costself-hostedfree
Latency3.6s median
Run targetlocal/sefi-image-2b-turbo
Effective requestmodel: sefi-image-2b-turbo · size: 1024x1024 · seed: 42
SourcesModel cardWeights & checkpoint
All 192 generations
Every challenge, grouped by category and subcategory. Solid-bordered tiles with a check passed the VLM judge; dashed tiles with a cross failed. Images are placeholders until generated.
Text Rendering
30%Typography Style33%
Writing accuracy17%
Spatial Reasoning
40%Attributes Binding44%
Compositionality33%
Counting33%
Negation22%
Relative Position58%
Scale & Proportions33%
Human realism
49%Faces & Expressions50%
Full Body42%
Hands50%
Multi-Subject50%
Professional Studio
52%Camera & Lighting33%
Color Precision50%
Photorealism33%
Graphical design
47%Data Visualisation33%
Layout & Design22%
Style Diversity25%
Truthfulness
30%Photorealism0%
Physics & Reflections17%
World Knowledge42%