Documentation
README
ACL Experiments
Use this while the experimental story can still change. The ACL evidence bar is not "beats the baseline once": it is a defensible measurement of a language capability, with the failure modes examined.
Baseline honesty
- Include the strongest cheap baseline: a well-prompted current LLM has become mandatory context for most tasks — a method beating only pre-LLM systems invites the "does this matter now?" review.
- Tune baselines with the same care as your method (same search budget, same data); reviewers explicitly probe for asymmetric tuning.
- Report the trivial baselines (majority class, copy input, retrieval-only) when they contextualize how hard the task actually is.
This is the opening of the README. Read the full README on GitHub.