Lock the comparison
A fair comparison starts with one versioned challenge, one prompt, one method, one attempt policy, and one environment.
Capture the run
Commands, duration, access mode, environment, files, interventions, screenshots, and available usage data remain attached.
Review the evidence
Tests, accessibility, performance, visible defects, missing data, and reviewer notes are checked before scoring.
Withhold incomplete scores
Active or incomplete runs show Review pending or Score withheld. They never resemble published results.
Publish a scoped verdict
A published verdict describes one recorded test. It does not claim universal model superiority.
Correct transparently
Method changes are versioned. Corrections preserve the earlier record and state what changed.