V1.1: Grading the Benchmark with a Frontier Judge
Swapping local open-VLM judges for a frontier model with specialized per-prompt questions — a stricter, more consistent grader. Capability drops ~22 points; the ranking holds.
The full write-up is published alongside the generated benchmark images. In the meantime, see the Benchmark V1 methodology and the live leaderboard.