Ghost Palette

AI Image Model Leaderboard

Benchmark V1 · ranked by pass rate on 192 prompts

32 models ranked by Overall — capability (graded pass/fail by VLM judges across 192 prompts in 6 categories) blended 50/50 with aesthetic Estimated Preference. Every generated image is published so you can judge with your own eyes.


V1 Leaderboard

Click any model to see every image it generated, or open two models to compare them side-by-side.

32/32
#Model
1openai/gpt-image-2api78.583.9%73.0token-based45.3s
2fal/google/nano-banana-2api73.073.4%72.6$0.08 / image28.1s
3fal/google/nano-banana-proapi66.373.4%59.1$0.15 / image23.4s
4local/boogu-image-turbolocal62.057.3%66.6N/A9.9s
5bfl/flux-2-proapi60.263.0%57.3from $0.03 / image11.8s
6bfl/flux-2-maxapi59.662.5%56.6from $0.07 / image26.7s
7fal/bytedance/seedream-v4api59.462.5%56.2$0.03 / image14.1s
8local/qwen-image-2512-20blocal58.650.0%67.1N/A80.2s
9local/flux-2-klein-9blocal56.653.1%60.1N/A8.5s
10local/krea-2-turbolocal56.557.8%55.2N/A73.2s
11bfl/flux-2-klein-9bapi55.453.6%57.1from $0.015 / image4.1s
12local/krea-2-turbo-no-filterlocal53.853.6%54.0N/A71.5s
13local/flux-2-klein-4blocal51.548.4%54.6N/A4.5s
14bfl/flux-2-klein-4bapi49.846.4%53.2from $0.014 / image3.8s
15local/sefi-image-5b-baselocal49.758.9%40.4N/A131s
16local/z-image-turbo-6blocal49.248.4%49.9N/A18.1s
17fal/krea/v2-medium-turboapi48.756.8%40.5not found15.5s
18fal/ideogram/v3api48.643.8%53.3$0.06 / image12.9s
19fal/bria/fastapi47.951.0%44.8$0.028 / generation12.4s
20local/sefi-image-5b-rllocal47.756.3%39.1N/A131s
21local/z-image-6blocal47.453.1%41.7N/A131s
22fal/krea/v2-mediumapi46.958.9%34.9$0.030 / image18.5s
23local/bonsai-image-ternary-4blocal46.643.2%49.9N/A4.1s
24local/prxpixel-t2i-7blocal46.442.7%50.0N/A64.7s
25fal/ideogram/v4api46.258.3%34.0$0.015 / MP16.6s
26fal/krea/v2-largeapi45.959.9%31.9$0.060 / image30.1s
27local/hidream-i1-full-17blocal42.935.9%49.9N/A91.3s
28local/krea-2-rawlocal42.157.3%26.9N/A912s
29local/sefi-image-5b-turbolocal41.952.6%31.2N/A5.9s
30local/nucleus-image-17b-a2blocal36.641.7%31.5N/A39.1s
31local/sefi-image-2b-turbolocal33.641.1%26.1N/A3.6s
32local/sana-1.5-1.6blocal33.031.8%34.2N/A11.1s

Local model analysis: Quality vs. size

Overall score against model size for the open-weight models — smaller models that still score highly are the sweet spot.

204060800B4B8B12B16B20B↖ smaller & higher qualitylocal/boogu-image-turbo: 8B, 62local/qwen-image-2512-20b: 20B, 58.6local/flux-2-klein-9b: 9B, 56.6local/krea-2-turbo: 12B, 56.5local/krea-2-turbo-no-filter: 12B, 53.8local/flux-2-klein-4b: 4B, 51.5local/sefi-image-5b-base: 5B, 49.7local/z-image-turbo-6b: 6B, 49.2local/sefi-image-5b-rl: 5B, 47.7local/z-image-6b: 6B, 47.4local/bonsai-image-ternary-4b: 4B, 46.6local/prxpixel-t2i-7b: 7B, 46.4local/hidream-i1-full-17b: 17B, 42.9local/krea-2-raw: 12B, 42.1local/sefi-image-5b-turbo: 5B, 41.9local/nucleus-image-17b-a2b: 17B, 36.6local/sefi-image-2b-turbo: 2B, 33.6local/sana-1.5-1.6b: 1.6B, 33Model size (billions of parameters)Overall score

API model analysis: Quality vs. price

Overall score against API price per image for the hosted models with a flat per-image price — cheaper models that still score highly are the sweet spot.

20406080100$0.00$0.04$0.08$0.12$0.16↖ cheaper & higher qualityfal/google/nano-banana-2: $0.08, 73fal/google/nano-banana-pro: $0.15, 66.3bfl/flux-2-pro: $0.03, 60.2bfl/flux-2-max: $0.07, 59.6fal/bytedance/seedream-v4: $0.03, 59.4bfl/flux-2-klein-9b: $0.01, 55.4bfl/flux-2-klein-4b: $0.01, 49.8fal/ideogram/v3: $0.06, 48.6fal/bria/fast: $0.03, 47.9fal/krea/v2-medium: $0.03, 46.9fal/krea/v2-large: $0.06, 45.9API price per image (USD)Overall score

RealBench V1

Realism leaderboard — can a model's output pass as a real photograph? Scored by human votes.

See the full realism benchmark
#ModelRealism scoreRated real / votesImages
1fal/google/nano-banana-pro53%1,907 / 3,609141
2fal/bytedance/seedream-v447%1,679 / 3,539139
3openai/gpt-image-246%1,572 / 3,418139
4local/z-image-turbo-6b46%1,646 / 3,603140
5bfl/flux-2-max41%1,473 / 3,592141


Frequently asked questions

A generative-image benchmark that shows the images. Benchmark V1 scores 32 models on 192 prompts across six capability categories; RealBench V1 scores how photoreal each model looks. Every output is published so you can judge with your own eyes which model fits your use case, budget, and quality bar.