๐Ÿงช
QualityPython

Agent Evaluation

by seb1n

Agent Evaluation is a Quality skill for Claude Code, published by seb1n in awesome-ai-agent-skills.

174 stars32 forkson seb1n/awesome-ai-agent-skillsAdded 2026/09/07+2% in starsRepository updated 2026/08/09
agent-skillsai-agent-skillsai-agentsawesome-listclaude-codeclaude-code-skillsclaude-skillscodexcodex-skillscontext-engineeringcursorcursor-skillsgemini-cligemini-skillsgithub-copilotmcpopenai-codexskill-mdskillswindsurf
Install in seconds
Install Agent Evaluation
Copy Agent Evaluation into your Claude Code skills folder. Run the command in your terminal, or review the source on GitHub before installing.
terminal
npx degit https://github.com/seb1n/awesome-ai-agent-skills/tree/main/agent-engineering/agent-evaluation ~/.claude/skills/agent-evaluation

Requires Node.js. Downloads this skill only โ€” not the rest of the repository โ€” into your Claude Code skills folder.

Without Node.js

git clone https://github.com/seb1n/awesome-ai-agent-skills.git

Clones the whole repository, then copy the skillโ€™s own directory into your skills folder yourself.

In this catalog

Source file
agent-engineering/agent-evaluation/SKILL.md in seb1n/awesome-ai-agent-skills
Installs to
~/.claude/skills/agent-evaluation
Collection
One of 25 skills cataloged from this repository
Category
Quality โ€” 1817 skills

What Agent Evaluation does

Agent Evaluation designs reproducible tests for AI agents using task sets, rubrics, graders, baselines, and release gates. Use it when comparing prompts or models, validating a release, or investigating regressions.

Agent Evaluation is cataloged under Quality on DirSkills. Agent Evaluation comes from a repository tagged agent-skills, ai-agent-skills, ai-agents, awesome-list and claude-code.

Documentation

README

Agent Evaluation

Build evidence that can inform a release owner, not a showcase of favorable examples or a safety certification.

Use when

  • Define quality before building or changing an agent.
  • Compare prompts, models, tools, memory strategies, or orchestration patterns.
  • Convert production failures into regression cases.
  • Establish a repeatable release gate or human-review plan.

Inputs

Collect the agent objective, users, supported tasks, unacceptable outcomes, current baseline, execution environment, available traces, and evaluation budget. State assumptions when an input is unavailable.

Output contract

Produce:

This is the opening of the README. Read the full README on GitHub.

Frequently asked about Agent Evaluation

  • What else does seb1n publish alongside Agent Evaluation?

    Agent Evaluation is one of 25 skills that DirSkills catalogs from seb1n/awesome-ai-agent-skills, the repository it ships in. Its siblings there include API Design, API Integration and Agent Observability. Each one is a separate skill with its own page in this directory, installs the same way Agent Evaluation does, and is maintained by seb1n in that same repository. The rest of the collection is listed on the seb1n/awesome-ai-agent-skills page.

  • How does Agent Evaluation compare to other Quality skills?

    Agent Evaluation ranks #1716 by stars among the 1817 Quality skills in this catalog. The most-starred ones next to it are Benchmark, Benchmark Optimization Loop and API Design Patterns. DirSkills ranks by the star count of the repository each skill ships in, so that order reflects how popular those repositories are rather than any review of Agent Evaluation against them. Open each page to compare what they document and how they install.

More from seb1n/awesome-ai-agent-skills

Agent Evaluation is one of 25 skills cataloged on DirSkills from seb1n/awesome-ai-agent-skills.

See all 25 skills โ†’
๐Ÿงฉ
2h ago

API Design

API Design helps design production-quality RESTful APIs with resource modeling, HTTP methods, status codes, pagination, versioning, and OpenAPI docs. Use it when you need endpoint definitions, schemas, or a REST API specification.
AI Engineering
17432
๐Ÿ”Œ
2h ago

API Integration

API Integration builds reliable clients for external APIs using REST, webhooks, polling, or SDK wrappers. It handles authentication, rate limits, retries, and error handling when you need production-ready integration code.
AI Engineering
17432
๐Ÿ“ˆ
2h ago

Agent Observability

Agent Observability designs privacy-aware telemetry for AI agents using traces, spans, metrics, dashboards, and alerts. Use it to debug failures, measure latency and cost, define SLOs, and audit agent behavior.
AI Engineering
17432
๐Ÿ›ก๏ธ
2h ago

Agent Red Teaming

Agent Red Teaming plans, executes, documents, and retests authorized security assessments of AI agents and multi-agent workflows. Use it to test prompt injection, tool and identity boundaries, memory or cross-agent attacks, and remediation in approved environments.
AI Engineering
17432
๐Ÿ“
2h ago

Code Documentation

Code Documentation generates docstrings, API references, and README files from source code. Use it when you need inline docs, usage guides, or project documentation for existing codebases.
Writing
17432
๐Ÿ”
2h ago

Code Review

Code Review performs structured reviews of files, diffs, or pull requests for bugs, security issues, performance problems, and maintainability concerns. Use it when you need actionable feedback on code changes.
Quality
17432