Documentation
README
The auditable method wins the argument and loses the score
Given a fixed clock and a scored predictions file, there is a recurring choice between two methods:
- one you can build in twenty minutes, cross-validate cleanly, explain fully, and defend against every question a reviewer asks;
- one the field actually uses to get the published number, which needs an hour of setup you have not done, has failure modes you cannot fully enumerate, and might not finish.
Every incentive inside a rigour-checking pipeline points at the first. The gates reward an auditable choice. The reviewer is easier to satisfy. The write-up is cleaner. And on a benchmark that grades predictions, none of that is measured.
This is the opening of the README. Read the full README on GitHub.