Ghost Palette

openai/gpt-image-2 — ImageBench V1

ImageBench V1 —
192 evaluations across 6 categories
Benchmark V1 verdicts are produced by VLM judges and can contain mistakes. Treat PASS/FAIL labels as machine-assisted assessments, and inspect the images yourself. Learn more about the methodology.

Generation Details

Source-backed model context, size, cost, and request settings for this run.

MakerOpenAI
FamilyGPT Image
Model sizenot disclosedundisclosed
Costtoken-basedvariable
Latency45.3s median
Run targetopenai/gpt-image-2
Effective requestmodel: gpt-image-2 · size: 1024x1024 · seed: 42
SourcesOpenAI model docsOpenAI pricing
Text100%Spatial91%Human55%Pro Studio100%Graphical100%Truth74%Preference73%Latency0%

All 192 generations

Every challenge, grouped by category and subcategory. Solid-bordered tiles with a check passed the VLM judge; dashed tiles with a cross failed. Images are placeholders until generated.

Text Rendering

100%
Typography Style100%
Writing accuracy83%

Spatial Reasoning

91%
Attributes Binding89%
Compositionality89%
Counting78%
Negation89%
Relative Position92%
Scale & Proportions89%

Human realism

55%
Faces & Expressions58%
Full Body50%
Hands67%
Multi-Subject50%

Professional Studio

100%
Camera & Lighting83%
Color Precision83%
Photorealism100%

Graphical design

100%
Data Visualisation100%
Layout & Design100%
Style Diversity83%

Truthfulness

74%
Photorealism67%
Physics & Reflections75%
World Knowledge58%