agentscope-ai/OpenJudge
DirSkills catalogs 18 skills from this repository, across 2 categories: AI Engineering, Quality.
⚖️
1w ago
Align Human
Align Human measures how well an automatic judge agrees with human labels, including TPR/TNR, kappa, AC1, and bias checks. Use it to validate a grader, find disagreement patterns, and decide how much human review can be reduced.
AI Engineering
79764
🏟️
1w ago
Auto Arena
Auto Arena automatically generates test queries, collects responses, and runs pairwise judging to compare multiple models or agents on a custom task. Use it to benchmark endpoints, resume evaluations, add endpoints incrementally, or swap the judge model.
AI Engineering
79764
📚
1w ago
BibTeX Verification
BibTeX Verification checks each .bib entry against CrossRef, arXiv, and DBLP to find fake or mis-cited references. Use it to audit bibliography files and report entries as verified, suspect, or not found.
Quality
79764
🧪
1w ago
Bootstrap
Bootstrap helps you create a first evaluation when you have no labels, no traces, and no criteria yet. It generates a v0 grader, synthetic test inputs, and a roadmap for collecting human labels and calibrating the evaluator.
AI Engineering
79764
🕵️
1w ago
Claude Authenticity
Claude Authenticity checks whether an API endpoint is backed by genuine Claude using weighted rule-based signals. Use it to audit Claude API keys, third-party providers, and optionally extract an injected system prompt.
AI Engineering
79764
🧪
1w ago
Eval Design
Eval Design helps create stratified evaluation datasets, adversarial test cases, and labeling guides from traces, specs, or interviews. It outputs OpenJudge-compatible datasets for GradingRunner.
AI Engineering
79764
📊
1w ago
Eval Report
Eval Report synthesizes evaluation runs into a maturity assessment, cross-skill signals, weaknesses, and prioritized actions. Use it for ship-readiness reviews, evaluation audits, and health checks of the eval system itself.
AI Engineering
79764
🧩
1w ago
Find Skills Combo
Find Skills Combo decomposes complex requests into subtasks and recommends combinations of agent skills. It is used when a task spans multiple domains or when you need a best-fit set of skills instead of one skill.
AI Engineering
79764
🧭
1w ago
Meta Eval
Meta Eval routes users to the right evaluation workflow for LLM and agent apps. It asks about data, labels, stakes, and domain knowledge, then recommends the next sub-skill to use.
AI Engineering
79764
📏
1w ago
Metric Design
Metric Design helps choose graders, design evaluation metrics, write LLM-as-judge prompts, and combine scores into an OpenJudge grading pipeline. Use it when you need to evaluate outputs automatically with executable grading code.
AI Engineering
79764
🎛️
1w ago
Mmx Cli
Mmx Cli generates text, images, video, speech, music, and web search results through MiniMax AI. Use it when you need AI-generated media or narration from prompts.
AI Engineering
79764
⚖️
1w ago
OpenJudge
OpenJudge builds evaluation pipelines for LLM outputs using graders, batch runners, aggregators, and analysis tools. Use it to score responses, compare models, and inspect results statistically.
AI Engineering
79764
📄
1w ago
Paper Review
Paper Review reviews academic papers for correctness, quality, novelty, and reference accuracy using OpenJudge's multi-stage pipeline. Use it when you need to assess a PDF, LaTeX source package, or BibTeX file for a research paper.
Quality
79764
📊
1w ago
Prompt Regression
Prompt Regression compares a baseline and candidate prompt to see whether the change improved results on each evaluation dimension. It is used for prompt A/B tests, regression checks, and other prompt optimization decisions with statistical confidence.
AI Engineering
79764
🧩
1w ago
RAG Eval
RAG Eval diagnoses RAG systems by separating retrieval problems from generation problems. Use it to check faithfulness, retrieval quality, hallucinations, and chunking changes with a diagnostic matrix.
AI Engineering
79764
🏆
1w ago
RL Reward Construction
RL Reward Construction builds reward signals with OpenJudge for RLHF and RLAIF. Use it to choose pointwise, pairwise, tournament, or listwise strategies for GRPO, DPO, Best-of-N, and reward normalization.
AI Engineering
79764
🛡️
1w ago
Red Teaming
Red Teaming tests an LLM or agent application for jailbreaks, prompt injection, PII extraction, harmful content, and evaluator gaming. Use it to measure ASR paired with over-refusal rate and produce an audit document before deployment.
AI Engineering
79764
📚
1w ago
Reference Hallucination Arena
Reference Hallucination Arena evaluates LLMs on recommending real academic papers by checking cited references against Crossref, PubMed, arXiv, and DBLP. Use it to benchmark citation accuracy, hallucination rate, and year/field constraints, including tool-augmented search mode.
Quality
79764