Ghost Palette
Back to blog

Human Evaluation

ELO rankings, pairwise preference, rater calibration, and LLM-as-a-Judge approaches.

The full write-up is published alongside the generated benchmark images. In the meantime, see the Benchmark V1 methodology and the live leaderboard.