⚖️
AI EngineeringJavaScript

Agent Eval

by affaan-m

Agent Eval is an AI Engineering skill for Claude Code, published by affaan-m in ECC.

239.8K stars36.4K forkson affaan-m/ECCAdded 2026/08/13Repository updated 2026/08/12
ai-agentsanthropicclaudeclaude-codedeveloper-toolsllmmcpproductivity
Install in seconds
Install Agent Eval
Copy Agent Eval into your Claude Code skills folder. Run the command in your terminal, or review the source on GitHub before installing.
terminal
npx degit https://github.com/affaan-m/ECC/tree/main/skills/agent-eval ~/.claude/skills/agent-eval

Requires Node.js. Downloads this skill only — not the rest of the repository — into your Claude Code skills folder.

Without Node.js

git clone https://github.com/affaan-m/ECC.git

Clones the whole repository, then copy the skill’s own directory into your skills folder yourself.

In this catalog

Source file
skills/agent-eval/SKILL.md in affaan-m/ECC
Installs to
~/.claude/skills/agent-eval
Collection
One of 25 skills cataloged from this repository
Category
AI Engineering2451 skills

What Agent Eval does

Agent Eval compares coding agents head-to-head on reproducible tasks, measuring pass rate, cost, time, and consistency. Use it to choose between agents or validate changes with data instead of impressions.

Agent Eval is cataloged under AI Engineering on DirSkills. Agent Eval comes from a repository tagged ai-agents, anthropic, claude, claude-code and developer-tools.

Documentation

README

Agent Eval Skill

A lightweight CLI tool for comparing coding agents head-to-head on reproducible tasks. Every "which coding agent is best?" comparison runs on vibes — this tool systematizes it.

When to Activate

  • Comparing coding agents (Claude Code, Aider, Codex, etc.) on your own codebase
  • Measuring agent performance before adopting a new tool or model
  • Running regression checks when an agent updates its model or tooling
  • Producing data-backed agent selection decisions for a team

Installation

Note: Install agent-eval from its repository after reviewing the source.

Core Concepts

YAML Task Definitions

Define tasks declaratively. Each task specifies what to do, which files to touch, and how to judge success:

This is the opening of the README. Read the full README on GitHub.

Frequently asked about Agent Eval

  • What else does affaan-m publish alongside Agent Eval?

    Agent Eval is one of 25 skills that DirSkills catalogs from affaan-m/ECC, the repository it ships in. Its siblings there include AI Regression Testing, AI-First Engineering and API Connector Builder. Each one is a separate skill with its own page in this directory, installs the same way Agent Eval does, and is maintained by affaan-m in that same repository. The rest of the collection is listed on the affaan-m/ECC page.

  • How does Agent Eval compare to other AI Engineering skills?

    Agent Eval ranks #3 by stars among the 2451 AI Engineering skills in this catalog. The most-starred ones next to it are Architecture Decision Records, AI-First Engineering and Agentic OS. DirSkills ranks by the star count of the repository each skill ships in, so that order reflects how popular those repositories are rather than any review of Agent Eval against them. Open each page to compare what they document and how they install.

More from affaan-m/ECC

Agent Eval is one of 25 skills cataloged on DirSkills from affaan-m/ECC.

See all 25 skills
🧪
3w ago

AI Regression Testing

AI Regression Testing provides testing patterns for AI-assisted development, catching blind spots when the same model writes and reviews code. Use it for sandbox-mode API tests and regression coverage after AI code changes.
Quality
239.8K36.4K
🤖
3w ago

AI-First Engineering

AI-First Engineering defines process, review, and architecture standards for teams where AI agents produce most implementation output. Use it to set ownership rules, quality gates, and evaluation criteria for agent-generated code.
AI Engineering
239.8K36.4K
🔌
3w ago

API Connector Builder

API Connector Builder builds a new API connector or provider by matching an existing integration pattern in the codebase exactly. Use it when adding a new integration to a project without introducing a second architecture.
DevOps
239.8K36.4K
🔌
3w ago

API Design Patterns

API Design Patterns provides conventions and best practices for designing consistent REST APIs. Use it when creating endpoints, reviewing contracts, handling pagination, error responses, status codes, filtering, or versioning.
Quality
239.8K36.4K
3w ago

Accessibility

Accessibility helps design, implement, and audit inclusive digital products using WCAG 2.2 Level AA standards for keyboard, contrast, and screen-reader support. Use it when building or auditing UI that must meet WCAG 2.2 Level AA or reviewing changes for accessibility.
Frontend
239.8K36.4K
🔍
3w ago

Agent Architecture Audit

Agent Architecture Audit diagnoses failures in agent and LLM applications by auditing the 12-layer agent stack for wrapper regression, memory pollution, tool discipline failures, hidden repair loops, and rendering corruption. Use it when an agent misbehaves and the failing layer is unknown, or before shipping an agent stack.
AI Engineering
239.8K36.4K