Ghost Palette

External rankings and benchmark glossary.

Industry reference

External image model benchmarks

Curated cross-metric rankings from published third-party leaderboards — human preference, automated suites, speed, and cost. Use this to orient; use Ghost Palette's live scores to verify on our stack.

Benchmark glossary

Six widely used evaluation frameworks — what each measures and when to trust it.

Artificial Analysis Image Arena

Arena Elo

Blind pairwise human votes aggregated with Bradley-Terry maximum likelihood estimation. The most widely cited overall quality signal for text-to-image models.

Best for: Overall human preference and cross-model ranking

ImageBench.ai

Pass rate

64 tests × 3 prompt variants (192 images), graded pass/fail by category-routed VLM judges. Fixed challenge CSV, deterministic routing, and every output published — the benchmark you can reproduce and inspect.

Best for: Task-specific diagnostics — text rendering, spatial reasoning, human realism, studio control, design layout, truthfulness

GenEval

Overall accuracy

Object-detection-based compositional accuracy on structured prompts. Measures whether generated images contain the right objects, counts, colors, and positions.

Best for: Prompt adherence, object counts, and spatial layout

Arena.ai

Arena Elo

Community-driven blind comparisons for text-to-image and image editing. Bradley-Terry ratings reflect real-world taste rather than narrow automated scores.

Best for: Community preference across generation and editing tasks

T2I-CompBench++

Category scores

Fine-grained compositional evaluation across attribute binding, spatial relations, numeracy, and complex prompts using specialized automated metrics.

Best for: Compositional prompt stress-testing

ImagenHub / GenAI-Arena

Human consistency scores

Standardized inference pipelines and human evaluation guidelines across seven conditional image generation tasks. Research-grade reproducibility.

Best for: Research reproducibility and multi-task evaluation

Industry comparison table

26 models across metrics where public data exists. Rows marked in Ghost Palette are available in the app today.

#ModelProviderArena EloImageBench (pub.)GenEvalGen time$/1k imgsOpenIn GP
1GPT Image 2OpenAI1,33995.3%84%45.3s$211No
2GPT Image 1.5OpenAI1,26591.2%82%38.0s$133No
3HiDream-O1-ImageHiDream-ai1,26581%12.0sFreeYes
4Nano Banana 2Google DeepMind1,25595.3%79%28.1s$67No
5Cosmos3-Super-Text2ImageNVIDIA1,23076%15.0sFreeYes
6Nano Banana ProGoogle DeepMind1,21492.2%77%23.4s$30No
7Recraft V4.1 ProRecraft1,20472%8.0s$40No
8FLUX.2 [max]Black Forest Labs1,20091.1%74%26.7s$50No
9Seedream 4.0ByteDance1,18573%10.0s$35NoYes
10Riverflow 2.0 ProRiverflow1,18070%14.0s$42No
11Qwen-Image-2512Alibaba1,16670.3%65%80.2s$20YesYes
12FLUX.2 [pro]Black Forest Labs1,15383.9%71%11.8s$55NoYes
13Ideogram 4Ideogram1,14082.8%70%16.6s$45No
14FLUX.2 [dev]Black Forest Labs1,12068%6.5s$25YesYes
15Recraft V3Recraft1,10066%7.0s$30NoYes
16Krea 2 MediumKrea1,09063%9.0s$28No
17FLUX.2 Klein 9BBlack Forest Labs1,08078.6%64%4.1s$10Yes
18FLUX.1 [dev]Black Forest Labs1,05062%5.5s$15Yes
19FLUX.2 Klein 4BBlack Forest Labs1,02073.4%60%3.8s$5Yes
20SD 3.5 LargeStability AI98058%8.0s$12YesYes
21Nano Banana 2 LiteGoogle DeepMindNoYes
22Z-Image Turbo 6BLocal / open74.0%59%18.1sFreeYes
23Nucleus Image 17BLocal / open65.6%55%39.1sFreeYes
24HiDream-I1 Full 17BLocal / open57.8%52%FreeYes
25Sana 1.5 1.6BLocal / open52.1%48%11.1sFreeYes
26Ideogram V3Ideogram67.7%61%$35NoYes

External reference data as of 2026-06. Arena Elo from Artificial Analysis Image Arena, published ImageBench pass rates from ImageBench.ai. Ghost Palette does not reproduce these numbers — see the live leaderboard for GP-reproduced scores.

Reproduce ImageBench on Ghost Palette

Industry tables use published scores. See the full ImageBench V1 reference — methodology, routing, and official leaderboard — on the ImageBench docs, or run the same suite here to compare on our stack.