Ghost Palette
Back to blog

Why a Few Good Prompts Are Enough

It feels like a good benchmark needs thousands of prompts. In practice, a small, carefully chosen set can rank models almost as well as a full suite — the theory behind it comes from Item Response Theory and the tinyBenchmarks line of work.

Discriminative prompts do the work

Most prompts are either trivial (every model passes) or hopeless (every model fails); neither separates models. The prompts that matter are the discriminative ones — those a good model passes and a weak model fails. Concentrating on them lets 64 well-designed tests reproduce the ordering of a much larger set.

Why variance still matters

A small set is more exposed to luck, which is why each test uses three prompt variants and why per-category pass rates are published — so a single unlucky generation can't quietly move a rank.