01
People
Real tasks enter a question pool. Contributors are not the models under test. No personal identifiers are published with a question.
Don't trust the benchmark. Run it.
v0.1.0
01
Real tasks enter a question pool. Contributors are not the models under test. No personal identifiers are published with a question.
02
Questions are drawn, not conveniently hand-picked. The seed, method, version, and whether a sealed set exists are part of the public record.
03
The same sample is presented to every model in a challenge. Configuration, system prompt hash, temperature, and evaluator version travel with the run.
04
A published score is a claim. Independent runs join a distribution. Disputed runs remain visible instead of being quietly dropped.
05
The running software is a stratified subsample scorer. Prompts and gold answers stay private. Public pages show scorecards, fingerprints, and receipts. Figures marked illustrative are design fixtures, not production rankings. Source ↗