🧪
QualityPython

Agent Evaluation

by davila7

Agent Evaluation is a Quality skill for Claude Code, published by davila7 in claude-code-templates.

30.2K stars3.4K forkson davila7/claude-code-templatesAdded 2026/08/14Repository updated 2026/08/14
anthropicanthropic-claudeclaudeclaude-code
Install in seconds
Install Agent Evaluation
Copy Agent Evaluation into your Claude Code skills folder. Run the command in your terminal, or review the source on GitHub before installing.
terminal
npx degit https://github.com/davila7/claude-code-templates/tree/main/cli-tool/components/skills/ai-research/agent-evaluation ~/.claude/skills/agent-evaluation

Requires Node.js. Downloads this skill only — not the rest of the repository — into your Claude Code skills folder.

Without Node.js

git clone https://github.com/davila7/claude-code-templates.git

Clones the whole repository, then copy the skill’s own directory into your skills folder yourself.

In this catalog

Source file
cli-tool/components/skills/ai-research/agent-evaluation/SKILL.md in davila7/claude-code-templates
Installs to
~/.claude/skills/agent-evaluation
Collection
One of 25 skills cataloged from this repository
Category
Quality1354 skills

What Agent Evaluation does

Agent Evaluation tests and benchmarks LLM agents using behavioral contracts, capability assessments, reliability metrics, and adversarial testing to catch issues before production. Use it when evaluating agent reliability, designing benchmarks, or monitoring production agents.

Agent Evaluation is cataloged under Quality on DirSkills. Agent Evaluation comes from a repository tagged anthropic, anthropic-claude, claude and claude-code.

Documentation

README

Agent Evaluation

You're a quality engineer who has seen agents that aced benchmarks fail spectacularly in production. You've learned that evaluating LLM agents is fundamentally different from testing traditional software—the same input can produce different outputs, and "correct" often has no single answer.

You've built evaluation frameworks that catch issues before production: behavioral regression tests, capability assessments, and reliability metrics. You understand that the goal isn't 100% test pass rate—it

Capabilities

  • agent-testing
  • benchmark-design
  • capability-assessment
  • reliability-metrics
  • regression-testing

This is the opening of the README. Read the full README on GitHub.

Frequently asked about Agent Evaluation

  • What else does davila7 publish alongside Agent Evaluation?

    Agent Evaluation is one of 25 skills that DirSkills catalogs from davila7/claude-code-templates, the repository it ships in. Its siblings there include AI Agents Architect, Agent Management and Agent Manager. Each one is a separate skill with its own page in this directory, installs the same way Agent Evaluation does, and is maintained by davila7 in that same repository. The rest of the collection is listed on the davila7/claude-code-templates page.

  • How does Agent Evaluation compare to other Quality skills?

    Agent Evaluation ranks #91 by stars among the 1354 Quality skills in this catalog. The most-starred ones next to it are Benchmark, Benchmark Optimization Loop and API Design Patterns. DirSkills ranks by the star count of the repository each skill ships in, so that order reflects how popular those repositories are rather than any review of Agent Evaluation against them. Open each page to compare what they document and how they install.

More from davila7/claude-code-templates

Agent Evaluation is one of 25 skills cataloged on DirSkills from davila7/claude-code-templates.

See all 25 skills
🤖
2w ago

AI Agents Architect

AI Agents Architect designs and builds autonomous AI agents, covering tool use, memory systems, planning strategies, and multi-agent orchestration. Use it when building or debugging AI agents that need function calling, planning loops, and controlled autonomy.
AI Engineering
30.2K3.4K
🤖
2w ago

Agent Management

Agent Management creates, manages, and orchestrates AI agents through the AI Maestro CLI, covering agent lifecycle tasks such as create, hibernate, wake, rename, export/import, and plugin management.
AI Engineering
30.2K3.4K
🤖
2w ago

Agent Manager

Agent Manager starts, stops, monitors, and assigns tasks to multiple local CLI agents running in tmux sessions, with cron-friendly scheduling. Use it when you need to run agents in parallel and tail their logs.
Automation
30.2K3.4K
🧠
2w ago

Agent Memory MCP

Agent Memory MCP provides a persistent, searchable memory bank for AI agents, exposing MCP tools to search, write, read, and analyze project knowledge. Use it when an agent needs long-term memory synced with project documentation.
AI Engineering
30.2K3.4K
🧠
2w ago

Agent Memory Systems

Agent Memory Systems describes architectures for short-term, long-term, and working memory in AI agents, including vector store selection, chunking strategies, and retrieval patterns. Use it when designing or debugging agent memory to prevent retrieval failures that look like intelligence failures.
AI Engineering
30.2K3.4K
✉️
2w ago

Agent Messaging

Agent Messaging sends and receives cryptographically signed messages between AI agents using the Agent Messaging Protocol (AMP). Use when the user asks to send a message to an agent, check agent inbox, message another agent, reply to a message, notify an agent, or any inter-agent communication task.
AI Engineering
30.2K3.4K