Browse Skills
11972 skills across 8 categories
๐
4w ago
Relay 80 To 100 Workflow
Relay 80 To 100 Workflow defines agent-relay processes that validate features end to end before merge. It is used for workflows that need repair loops, test-fix-rerun gates, and fresh-eyes review before commit.
AI Engineering
80164
๐
4w ago
Relay Workflows
Relay Workflows helps build multi-agent workflows with @relayflows/core using DAG steps, agent coordination, and output chaining. Use it for review/fix loops, verification gates, and chat-native or pipeline-style orchestration.
AI Engineering
80164
๐
4w ago
Review Fix Signoff Loop
Review Fix Signoff Loop coordinates repeated review, fix, and validation cycles with fresh agent context until independent signoff agents both approve the result. It is used for high-stakes workflows that need deterministic gates, repair passes, and final signoff reporting.
AI Engineering
80164
๐
4w ago
Review Fix Signoff Loop
Review Fix Signoff Loop defines a workflow for implementation tasks that must cycle through review, repair, validation, and fresh-context signoff until independent reviewers agree the work is complete. Use it for high-stakes agent workflows with repairable gates, dual-verdict contracts, and PR signoff reporting.
AI Engineering
80164
โ๏ธ
4w ago
Align Human
Align Human measures how well an automatic judge agrees with human labels, including TPR/TNR, kappa, AC1, and bias checks. Use it to validate a grader, find disagreement patterns, and decide how much human review can be reduced.
AI Engineering
79764
๐๏ธ
4w ago
Auto Arena
Auto Arena automatically generates test queries, collects responses, and runs pairwise judging to compare multiple models or agents on a custom task. Use it to benchmark endpoints, resume evaluations, add endpoints incrementally, or swap the judge model.
AI Engineering
79764
๐งช
4w ago
Bootstrap
Bootstrap helps you create a first evaluation when you have no labels, no traces, and no criteria yet. It generates a v0 grader, synthetic test inputs, and a roadmap for collecting human labels and calibrating the evaluator.
AI Engineering
79764
๐ต๏ธ
4w ago
Claude Authenticity
Claude Authenticity checks whether an API endpoint is backed by genuine Claude using weighted rule-based signals. Use it to audit Claude API keys, third-party providers, and optionally extract an injected system prompt.
AI Engineering
79764
๐งช
4w ago
Eval Design
Eval Design helps create stratified evaluation datasets, adversarial test cases, and labeling guides from traces, specs, or interviews. It outputs OpenJudge-compatible datasets for GradingRunner.
AI Engineering
79764
๐
4w ago
Eval Report
Eval Report synthesizes evaluation runs into a maturity assessment, cross-skill signals, weaknesses, and prioritized actions. Use it for ship-readiness reviews, evaluation audits, and health checks of the eval system itself.
AI Engineering
79764
๐งฉ
4w ago
Find Skills Combo
Find Skills Combo decomposes complex requests into subtasks and recommends combinations of agent skills. It is used when a task spans multiple domains or when you need a best-fit set of skills instead of one skill.
AI Engineering
79764
๐งญ
4w ago
Meta Eval
Meta Eval routes users to the right evaluation workflow for LLM and agent apps. It asks about data, labels, stakes, and domain knowledge, then recommends the next sub-skill to use.
AI Engineering
79764