Methodology
How Ghost Palette evaluates models
Fair comparison rules and workflow design aligned with how industry benchmarks normalize runs — so your personal evals are as rigorous as public leaderboards.
Fair-run rules
Ghost Palette follows the same normalization principles used by Artificial Analysis and research benchmarks like ImagenHub.
- Same promptEvery model in a run receives identical prompt text — no per-model prompt engineering.
- Fixed seed (optional)Use a shared seed when models support it to isolate model differences from randomness.
- Consistent resolutionGenerate at the same aspect ratio and resolution where the API allows, matching industry normalization (typically 1024×1024).
- Model defaultsUse each provider's documented defaults for steps, guidance, and safety — no hidden tuning per model.
Workflow mapping
Each Ghost Palette workflow targets a different creation and evaluation dimension.
Create
Human preferenceRun one prompt across multiple models, compare outputs side by side, and pick a winner — blind or named. Mirrors Arena.ai and Artificial Analysis.
Refine
Task-specific fitProvide a reference direction and compare how each model interprets the same refinement instruction.
Gallery
ReproducibilityEvery run is saved automatically with prompt, model, and outputs so decisions are documented and repeatable.
External benchmark methodology
For public leaderboard scores, refer to each source's published methodology.
Put methodology into practice
Start with a fair comparison run in the Create studio.