Blog
Analysis, methodology, and findings from the Ghost Palette benchmark project.
Does the score match your eye?
A blind test of the aesthetic metric: look at the images, pick your favourite, and see how often you disagree with the score — a critique of what a single preference number can tell you.
V1.1: Grading the Benchmark with a Frontier Judge
Swapping local open-VLM judges for a frontier model with specialized per-prompt questions — a stricter, more consistent grader. Capability drops ~22 points; the ranking holds.
Estimated Preference Score
A second axis for Benchmark V1 — an aesthetic human-preference score, combined with capability into a new Overall score.
Tuning a Turbo Model for Better Prompt Following
A controlled strength sweep of one conditioning weight — prompt adherence rises to a peak around strength 1, then degrades.
Why a Few Good Prompts Are Enough
The theory behind the benchmark — Item Response Theory and tinyBenchmarks explain why a small, discriminative prompt set can rank models almost as well as a full benchmark.
Can Local VLMs Recognize Public Figures?
A 90-image public-figure recognition study across global icons, field-famous specialists, and long-tail notables.
Which VLM Should Judge Style Diversity?
A VLM calibration study for Style Diversity, selecting the most reliable route judge.
RealBench V1 Methodology
How RealBench measures photorealism — paired real/AI images and human votes from the community.
Which VLM Best Detects Bad Hands?
A detector calibration study using strong and weak hand outputs to pick the most reliable judge.
GPT Image 2 vs Flux 2 vs Nano Banana 2
Side-by-side benchmark results for the leading AI image generation models.
Best AI Image Generator for Text Rendering
Which models handle spelling, posters, labels, typography, and small text best.
Benchmark V1 Methodology
How Benchmark V1 is designed, scored, and reported across 64 tests and 6 capability categories.
Quality Metrics Across 10 Models
MUSIQ, NIQE, NIMA, and TOPIQ scores for the benchmark models — and why the results are still preliminary.
Glossary
A–Z reference of every term used across the guides.
Safety & Bias
NSFW content, demographic bias, IP concerns, red teaming, and the regulatory landscape.
Consistency & Reproducibility
Same prompt, different outputs — measuring variance and why it matters for production.
Prompt Fidelity & Compositionality
Does the image match the text? Measuring attribute binding, spatial reasoning, and counting.
Comparing Image Models
Quality, speed, cost, consistency — how to compare fairly and find the Pareto frontier.
Human Evaluation
ELO rankings, pairwise preference, rater calibration, and LLM-as-a-Judge approaches.
Automated Metrics
FID, CLIP Score, LPIPS, VQAScore — what they measure, when to use them, and common pitfalls.
Introduction to Image Evaluation
Why evaluating image generation matters, and the two sides: automated metrics vs human judgment.