openai/gpt-image-2 — ImageBench V1
ImageBench V1 —
192 evaluations across 6 categoriesBenchmark V1 verdicts are produced by VLM judges and can contain mistakes. Treat PASS/FAIL labels as machine-assisted assessments, and inspect the images yourself. Learn more about the methodology.
Generation Details
Source-backed model context, size, cost, and request settings for this run.
MakerOpenAI
FamilyGPT Image
Model sizenot disclosedundisclosed
Costtoken-basedvariable
Latency45.3s median
Run targetopenai/gpt-image-2
Effective requestmodel: gpt-image-2 · size: 1024x1024 · seed: 42
SourcesOpenAI model docsOpenAI pricing
All 192 generations
Every challenge, grouped by category and subcategory. Solid-bordered tiles with a check passed the VLM judge; dashed tiles with a cross failed. Images are placeholders until generated.
Text Rendering
100%Typography Style100%
Writing accuracy83%
Spatial Reasoning
91%Attributes Binding89%
Compositionality89%
Counting78%
Negation89%
Relative Position92%
Scale & Proportions89%
Human realism
55%Faces & Expressions58%
Full Body50%
Hands67%
Multi-Subject50%
Professional Studio
100%Camera & Lighting83%
Color Precision83%
Photorealism100%
Graphical design
100%Data Visualisation100%
Layout & Design100%
Style Diversity83%
Truthfulness
74%Photorealism67%
Physics & Reflections75%
World Knowledge58%