agentscope-ai/OpenJudge

DirSkills catalogs 18 skills from this repository, across 2 categories: AI Engineering, Quality.

797 stars64 forksView on GitHub
⚖️
1w ago

Align Human

Align Human measures how well an automatic judge agrees with human labels, including TPR/TNR, kappa, AC1, and bias checks. Use it to validate a grader, find disagreement patterns, and decide how much human review can be reduced.
AI Engineering
79764
🏟️
1w ago

Auto Arena

Auto Arena automatically generates test queries, collects responses, and runs pairwise judging to compare multiple models or agents on a custom task. Use it to benchmark endpoints, resume evaluations, add endpoints incrementally, or swap the judge model.
AI Engineering
79764
📚
1w ago

BibTeX Verification

BibTeX Verification checks each .bib entry against CrossRef, arXiv, and DBLP to find fake or mis-cited references. Use it to audit bibliography files and report entries as verified, suspect, or not found.
Quality
79764
🧪
1w ago

Bootstrap

Bootstrap helps you create a first evaluation when you have no labels, no traces, and no criteria yet. It generates a v0 grader, synthetic test inputs, and a roadmap for collecting human labels and calibrating the evaluator.
AI Engineering
79764
🕵️
1w ago

Claude Authenticity

Claude Authenticity checks whether an API endpoint is backed by genuine Claude using weighted rule-based signals. Use it to audit Claude API keys, third-party providers, and optionally extract an injected system prompt.
AI Engineering
79764
🧪
1w ago

Eval Design

Eval Design helps create stratified evaluation datasets, adversarial test cases, and labeling guides from traces, specs, or interviews. It outputs OpenJudge-compatible datasets for GradingRunner.
AI Engineering
79764
📊
1w ago

Eval Report

Eval Report synthesizes evaluation runs into a maturity assessment, cross-skill signals, weaknesses, and prioritized actions. Use it for ship-readiness reviews, evaluation audits, and health checks of the eval system itself.
AI Engineering
79764
🧩
1w ago

Find Skills Combo

Find Skills Combo decomposes complex requests into subtasks and recommends combinations of agent skills. It is used when a task spans multiple domains or when you need a best-fit set of skills instead of one skill.
AI Engineering
79764
🧭
1w ago

Meta Eval

Meta Eval routes users to the right evaluation workflow for LLM and agent apps. It asks about data, labels, stakes, and domain knowledge, then recommends the next sub-skill to use.
AI Engineering
79764
📏
1w ago

Metric Design

Metric Design helps choose graders, design evaluation metrics, write LLM-as-judge prompts, and combine scores into an OpenJudge grading pipeline. Use it when you need to evaluate outputs automatically with executable grading code.
AI Engineering
79764
🎛️
1w ago

Mmx Cli

Mmx Cli generates text, images, video, speech, music, and web search results through MiniMax AI. Use it when you need AI-generated media or narration from prompts.
AI Engineering
79764
⚖️
1w ago

OpenJudge

OpenJudge builds evaluation pipelines for LLM outputs using graders, batch runners, aggregators, and analysis tools. Use it to score responses, compare models, and inspect results statistically.
AI Engineering
79764
📄
1w ago

Paper Review

Paper Review reviews academic papers for correctness, quality, novelty, and reference accuracy using OpenJudge's multi-stage pipeline. Use it when you need to assess a PDF, LaTeX source package, or BibTeX file for a research paper.
Quality
79764
📊
1w ago

Prompt Regression

Prompt Regression compares a baseline and candidate prompt to see whether the change improved results on each evaluation dimension. It is used for prompt A/B tests, regression checks, and other prompt optimization decisions with statistical confidence.
AI Engineering
79764
🧩
1w ago

RAG Eval

RAG Eval diagnoses RAG systems by separating retrieval problems from generation problems. Use it to check faithfulness, retrieval quality, hallucinations, and chunking changes with a diagnostic matrix.
AI Engineering
79764
🏆
1w ago

RL Reward Construction

RL Reward Construction builds reward signals with OpenJudge for RLHF and RLAIF. Use it to choose pointwise, pairwise, tournament, or listwise strategies for GRPO, DPO, Best-of-N, and reward normalization.
AI Engineering
79764
🛡️
1w ago

Red Teaming

Red Teaming tests an LLM or agent application for jailbreaks, prompt injection, PII extraction, harmful content, and evaluator gaming. Use it to measure ASR paired with over-refusal rate and produce an audit document before deployment.
AI Engineering
79764
📚
1w ago

Reference Hallucination Arena

Reference Hallucination Arena evaluates LLMs on recommending real academic papers by checking cited references against Crossref, PubMed, arXiv, and DBLP. Use it to benchmark citation accuracy, hallucination rate, and year/field constraints, including tool-augmented search mode.
Quality
79764