Documentation
README
Finetuning
Priors, not rules. Only firm guardrails: held-out eval you never train on, no leakage, trust evo's recorded numbers over the run's self-report. Override anything else against the gate.
Pick the technique by reward shape
Decide on the reward first, technique second. Choosing the comfortable technique over the matching one is the most common failure.
This is the opening of the README. Read the full README on GitHub.