Human Evaluation
ELO rankings, pairwise preference, rater calibration, and LLM-as-a-Judge approaches.
The full write-up is published alongside the generated benchmark images. In the meantime, see the Benchmark V1 methodology and the live leaderboard.