Industry reference
External image model benchmarks
Curated cross-metric rankings from published third-party leaderboards — human preference, automated suites, speed, and cost. Use this to orient; use Ghost Palette's live scores to verify on our stack.
Benchmark glossary
Six widely used evaluation frameworks — what each measures and when to trust it.
Artificial Analysis Image Arena
Arena EloBlind pairwise human votes aggregated with Bradley-Terry maximum likelihood estimation. The most widely cited overall quality signal for text-to-image models.
Best for: Overall human preference and cross-model ranking
ImageBench.ai
Pass rate64 tests × 3 prompt variants (192 images), graded pass/fail by category-routed VLM judges. Fixed challenge CSV, deterministic routing, and every output published — the benchmark you can reproduce and inspect.
Best for: Task-specific diagnostics — text rendering, spatial reasoning, human realism, studio control, design layout, truthfulness
GenEval
Overall accuracyObject-detection-based compositional accuracy on structured prompts. Measures whether generated images contain the right objects, counts, colors, and positions.
Best for: Prompt adherence, object counts, and spatial layout
Arena.ai
Arena EloCommunity-driven blind comparisons for text-to-image and image editing. Bradley-Terry ratings reflect real-world taste rather than narrow automated scores.
Best for: Community preference across generation and editing tasks
T2I-CompBench++
Category scoresFine-grained compositional evaluation across attribute binding, spatial relations, numeracy, and complex prompts using specialized automated metrics.
Best for: Compositional prompt stress-testing
ImagenHub / GenAI-Arena
Human consistency scoresStandardized inference pipelines and human evaluation guidelines across seven conditional image generation tasks. Research-grade reproducibility.
Best for: Research reproducibility and multi-task evaluation
Industry comparison table
26 models across metrics where public data exists. Rows marked in Ghost Palette are available in the app today.
| # | Model | Provider | Arena Elo | ImageBench (pub.) | GenEval | Gen time | $/1k imgs | Open | In GP |
|---|---|---|---|---|---|---|---|---|---|
| 1 | GPT Image 2 | OpenAI | 1,339 | 95.3% | 84% | 45.3s | $211 | No | — |
| 2 | GPT Image 1.5 | OpenAI | 1,265 | 91.2% | 82% | 38.0s | $133 | No | — |
| 3 | HiDream-O1-Image | HiDream-ai | 1,265 | — | 81% | 12.0s | Free | Yes | — |
| 4 | Nano Banana 2 | Google DeepMind | 1,255 | 95.3% | 79% | 28.1s | $67 | No | — |
| 5 | Cosmos3-Super-Text2Image | NVIDIA | 1,230 | — | 76% | 15.0s | Free | Yes | — |
| 6 | Nano Banana Pro | Google DeepMind | 1,214 | 92.2% | 77% | 23.4s | $30 | No | — |
| 7 | Recraft V4.1 Pro | Recraft | 1,204 | — | 72% | 8.0s | $40 | No | — |
| 8 | FLUX.2 [max] | Black Forest Labs | 1,200 | 91.1% | 74% | 26.7s | $50 | No | — |
| 9 | Seedream 4.0 | ByteDance | 1,185 | — | 73% | 10.0s | $35 | No | Yes |
| 10 | Riverflow 2.0 Pro | Riverflow | 1,180 | — | 70% | 14.0s | $42 | No | — |
| 11 | Qwen-Image-2512 | Alibaba | 1,166 | 70.3% | 65% | 80.2s | $20 | Yes | Yes |
| 12 | FLUX.2 [pro] | Black Forest Labs | 1,153 | 83.9% | 71% | 11.8s | $55 | No | Yes |
| 13 | Ideogram 4 | Ideogram | 1,140 | 82.8% | 70% | 16.6s | $45 | No | — |
| 14 | FLUX.2 [dev] | Black Forest Labs | 1,120 | — | 68% | 6.5s | $25 | Yes | Yes |
| 15 | Recraft V3 | Recraft | 1,100 | — | 66% | 7.0s | $30 | No | Yes |
| 16 | Krea 2 Medium | Krea | 1,090 | — | 63% | 9.0s | $28 | No | — |
| 17 | FLUX.2 Klein 9B | Black Forest Labs | 1,080 | 78.6% | 64% | 4.1s | $10 | Yes | — |
| 18 | FLUX.1 [dev] | Black Forest Labs | 1,050 | — | 62% | 5.5s | $15 | Yes | — |
| 19 | FLUX.2 Klein 4B | Black Forest Labs | 1,020 | 73.4% | 60% | 3.8s | $5 | Yes | — |
| 20 | SD 3.5 Large | Stability AI | 980 | — | 58% | 8.0s | $12 | Yes | Yes |
| 21 | Nano Banana 2 Lite | Google DeepMind | — | — | — | — | — | No | Yes |
| 22 | Z-Image Turbo 6B | Local / open | — | 74.0% | 59% | 18.1s | Free | Yes | — |
| 23 | Nucleus Image 17B | Local / open | — | 65.6% | 55% | 39.1s | Free | Yes | — |
| 24 | HiDream-I1 Full 17B | Local / open | — | 57.8% | 52% | — | Free | Yes | — |
| 25 | Sana 1.5 1.6B | Local / open | — | 52.1% | 48% | 11.1s | Free | Yes | — |
| 26 | Ideogram V3 | Ideogram | — | 67.7% | 61% | — | $35 | No | Yes |
External reference data as of 2026-06. Arena Elo from Artificial Analysis Image Arena, published ImageBench pass rates from ImageBench.ai. Ghost Palette does not reproduce these numbers — see the live leaderboard for GP-reproduced scores.
Reproduce ImageBench on Ghost Palette
Industry tables use published scores. See the full ImageBench V1 reference — methodology, routing, and official leaderboard — on the ImageBench docs, or run the same suite here to compare on our stack.