Documentation
README
Eval Design
Design high-quality evaluation datasets that measure what actually matters for your
application. You extract evaluation dimensions from business context, structure them
into stratified test cases, and output datasets ready for OpenJudge GradingRunner.
When to Activate
- User has agent traces / production logs and wants to build an eval set from them
- User has evaluation principles but needs properly stratified test data
- User wants to generate adversarial examples that stress-test their system
- User needs coverage analysis โ are they testing all the right things?
- User wants a labeling guide for human annotators
This is the opening of the README. Read the full README on GitHub.