Ghost Palette

Blog

Analysis, methodology, and findings from the Ghost Palette benchmark project.

01 6 min

Does the score match your eye?

A blind test of the aesthetic metric: look at the images, pick your favourite, and see how often you disagree with the score — a critique of what a single preference number can tell you.

02 7 min

V1.1: Grading the Benchmark with a Frontier Judge

Swapping local open-VLM judges for a frontier model with specialized per-prompt questions — a stricter, more consistent grader. Capability drops ~22 points; the ranking holds.

03 8 min

Estimated Preference Score

A second axis for Benchmark V1 — an aesthetic human-preference score, combined with capability into a new Overall score.

04 7 min

Tuning a Turbo Model for Better Prompt Following

A controlled strength sweep of one conditioning weight — prompt adherence rises to a peak around strength 1, then degrades.

05 9 min

Why a Few Good Prompts Are Enough

The theory behind the benchmark — Item Response Theory and tinyBenchmarks explain why a small, discriminative prompt set can rank models almost as well as a full benchmark.

06 6 min

Can Local VLMs Recognize Public Figures?

A 90-image public-figure recognition study across global icons, field-famous specialists, and long-tail notables.

07 5 min

Which VLM Should Judge Style Diversity?

A VLM calibration study for Style Diversity, selecting the most reliable route judge.

08 3 min

RealBench V1 Methodology

How RealBench measures photorealism — paired real/AI images and human votes from the community.

09 4 min

Which VLM Best Detects Bad Hands?

A detector calibration study using strong and weak hand outputs to pick the most reliable judge.

10 5 min

GPT Image 2 vs Flux 2 vs Nano Banana 2

Side-by-side benchmark results for the leading AI image generation models.

11 5 min

Best AI Image Generator for Text Rendering

Which models handle spelling, posters, labels, typography, and small text best.

12 5 min

Benchmark V1 Methodology

How Benchmark V1 is designed, scored, and reported across 64 tests and 6 capability categories.

13 7 min

Quality Metrics Across 10 Models

MUSIQ, NIQE, NIMA, and TOPIQ scores for the benchmark models — and why the results are still preliminary.

14 5 min

Glossary

A–Z reference of every term used across the guides.

15 10 min

Safety & Bias

NSFW content, demographic bias, IP concerns, red teaming, and the regulatory landscape.

16 7 min

Consistency & Reproducibility

Same prompt, different outputs — measuring variance and why it matters for production.

17 11 min

Prompt Fidelity & Compositionality

Does the image match the text? Measuring attribute binding, spatial reasoning, and counting.

18 9 min

Comparing Image Models

Quality, speed, cost, consistency — how to compare fairly and find the Pareto frontier.

19 10 min

Human Evaluation

ELO rankings, pairwise preference, rater calibration, and LLM-as-a-Judge approaches.

20 12 min

Automated Metrics

FID, CLIP Score, LPIPS, VQAScore — what they measure, when to use them, and common pitfalls.

21 8 min

Introduction to Image Evaluation

Why evaluating image generation matters, and the two sides: automated metrics vs human judgment.