AI Image Model Leaderboard
Benchmark V1 · ranked by pass rate on 192 prompts
32 models ranked by Overall — capability (graded pass/fail by VLM judges across 192 prompts in 6 categories) blended 50/50 with aesthetic Estimated Preference. Every generated image is published so you can judge with your own eyes.
V1 Leaderboard
Click any model to see every image it generated, or open two models to compare them side-by-side.
| # | Model | |||||
|---|---|---|---|---|---|---|
| 1 | openai/gpt-image-2api | 78.5 | 83.9% | 73.0 | token-based | 45.3s |
| 2 | fal/google/nano-banana-2api | 73.0 | 73.4% | 72.6 | $0.08 / image | 28.1s |
| 3 | fal/google/nano-banana-proapi | 66.3 | 73.4% | 59.1 | $0.15 / image | 23.4s |
| 4 | local/boogu-image-turbolocal | 62.0 | 57.3% | 66.6 | N/A | 9.9s |
| 5 | bfl/flux-2-proapi | 60.2 | 63.0% | 57.3 | from $0.03 / image | 11.8s |
| 6 | bfl/flux-2-maxapi | 59.6 | 62.5% | 56.6 | from $0.07 / image | 26.7s |
| 7 | fal/bytedance/seedream-v4api | 59.4 | 62.5% | 56.2 | $0.03 / image | 14.1s |
| 8 | local/qwen-image-2512-20blocal | 58.6 | 50.0% | 67.1 | N/A | 80.2s |
| 9 | local/flux-2-klein-9blocal | 56.6 | 53.1% | 60.1 | N/A | 8.5s |
| 10 | local/krea-2-turbolocal | 56.5 | 57.8% | 55.2 | N/A | 73.2s |
| 11 | bfl/flux-2-klein-9bapi | 55.4 | 53.6% | 57.1 | from $0.015 / image | 4.1s |
| 12 | local/krea-2-turbo-no-filterlocal | 53.8 | 53.6% | 54.0 | N/A | 71.5s |
| 13 | local/flux-2-klein-4blocal | 51.5 | 48.4% | 54.6 | N/A | 4.5s |
| 14 | bfl/flux-2-klein-4bapi | 49.8 | 46.4% | 53.2 | from $0.014 / image | 3.8s |
| 15 | local/sefi-image-5b-baselocal | 49.7 | 58.9% | 40.4 | N/A | 131s |
| 16 | local/z-image-turbo-6blocal | 49.2 | 48.4% | 49.9 | N/A | 18.1s |
| 17 | fal/krea/v2-medium-turboapi | 48.7 | 56.8% | 40.5 | not found | 15.5s |
| 18 | fal/ideogram/v3api | 48.6 | 43.8% | 53.3 | $0.06 / image | 12.9s |
| 19 | fal/bria/fastapi | 47.9 | 51.0% | 44.8 | $0.028 / generation | 12.4s |
| 20 | local/sefi-image-5b-rllocal | 47.7 | 56.3% | 39.1 | N/A | 131s |
| 21 | local/z-image-6blocal | 47.4 | 53.1% | 41.7 | N/A | 131s |
| 22 | fal/krea/v2-mediumapi | 46.9 | 58.9% | 34.9 | $0.030 / image | 18.5s |
| 23 | local/bonsai-image-ternary-4blocal | 46.6 | 43.2% | 49.9 | N/A | 4.1s |
| 24 | local/prxpixel-t2i-7blocal | 46.4 | 42.7% | 50.0 | N/A | 64.7s |
| 25 | fal/ideogram/v4api | 46.2 | 58.3% | 34.0 | $0.015 / MP | 16.6s |
| 26 | fal/krea/v2-largeapi | 45.9 | 59.9% | 31.9 | $0.060 / image | 30.1s |
| 27 | local/hidream-i1-full-17blocal | 42.9 | 35.9% | 49.9 | N/A | 91.3s |
| 28 | local/krea-2-rawlocal | 42.1 | 57.3% | 26.9 | N/A | 912s |
| 29 | local/sefi-image-5b-turbolocal | 41.9 | 52.6% | 31.2 | N/A | 5.9s |
| 30 | local/nucleus-image-17b-a2blocal | 36.6 | 41.7% | 31.5 | N/A | 39.1s |
| 31 | local/sefi-image-2b-turbolocal | 33.6 | 41.1% | 26.1 | N/A | 3.6s |
| 32 | local/sana-1.5-1.6blocal | 33.0 | 31.8% | 34.2 | N/A | 11.1s |
Local model analysis: Quality vs. size
Overall score against model size for the open-weight models — smaller models that still score highly are the sweet spot.
API model analysis: Quality vs. price
Overall score against API price per image for the hosted models with a flat per-image price — cheaper models that still score highly are the sweet spot.
RealBench V1
Realism leaderboard — can a model's output pass as a real photograph? Scored by human votes.
See the full realism benchmark| # | Model | Realism score | Rated real / votes | Images |
|---|---|---|---|---|
| 1 | fal/google/nano-banana-pro | 53% | 1,907 / 3,609 | 141 |
| 2 | fal/bytedance/seedream-v4 | 47% | 1,679 / 3,539 | 139 |
| 3 | openai/gpt-image-2 | 46% | 1,572 / 3,418 | 139 |
| 4 | local/z-image-turbo-6b | 46% | 1,646 / 3,603 | 140 |
| 5 | bfl/flux-2-max | 41% | 1,473 / 3,592 | 141 |
RealBench V1
Can a model's output pass as a real photograph? Scored by real human votes.
ExploreGallery
Browse generated images across models and prompts, side by side.
BrowseRun the suite
Reproduce any score — generate against the fixed prompt set and grade with a VLM judge.
RunFrequently asked questions
A generative-image benchmark that shows the images. Benchmark V1 scores 32 models on 192 prompts across six capability categories; RealBench V1 scores how photoreal each model looks. Every output is published so you can judge with your own eyes which model fits your use case, budget, and quality bar.