🧪
AI EngineeringJavaScript

Eval Harness

by xu-xiang

Eval Harness is an AI Engineering skill for Claude Code, published by xu-xiang in everything-claude-code-zh.

1.9K stars307 forkson xu-xiang/everything-claude-code-zhAdded 2026/08/18+1% in starsRepository updated 2026/03/05
ai-agentsanthropicclaudeclaude-codedeveloper-toolseverything-claude-codellmmcponeskillproductivity
Install in seconds
Install Eval Harness
Copy Eval Harness into your Claude Code skills folder. Run the command in your terminal, or review the source on GitHub before installing.
terminal
npx degit https://github.com/xu-xiang/everything-claude-code-zh/tree/main/skills/eval-harness ~/.claude/skills/eval-harness

Requires Node.js. Downloads this skill only — not the rest of the repository — into your Claude Code skills folder.

Without Node.js

git clone https://github.com/xu-xiang/everything-claude-code-zh.git

Clones the whole repository, then copy the skill’s own directory into your skills folder yourself.

In this catalog

Source file
skills/eval-harness/SKILL.md in xu-xiang/everything-claude-code-zh
Installs to
~/.claude/skills/eval-harness
Collection
One of 25 skills cataloged from this repository
Category
AI Engineering2451 skills

What Eval Harness does

Eval Harness provides a formal evaluation framework for Claude Code sessions, implementing eval-driven development (EDD) principles. Use it to define pass/fail criteria, measure agent reliability with pass@k, and create regression test suites.

Eval Harness is cataloged under AI Engineering on DirSkills. Eval Harness comes from a repository tagged ai-agents, anthropic, claude, claude-code and developer-tools.

Documentation

README

评测工具链(Eval Harness)技能(Skill)

一个用于 Claude Code 会话的正式评测框架,实现了评测驱动开发(Eval-Driven Development, EDD)原则。

何时激活

  • 为 AI 辅助工作流设置评测驱动开发(EDD)
  • 为 Claude Code 任务完成定义通过/失败的标准
  • 使用 pass@k 指标衡量智能体(Agent)的可靠性
  • 为提示词(Prompt)或智能体(Agent)的变更创建回归测试套件
  • 跨模型版本对智能体(Agent)性能进行基准测试

核心理念

评测驱动开发(Eval-Driven Development)将评测视为“AI 开发的单元测试”:

  • 在实现之定义预期行为
  • 在开发过程中持续运行评测
  • 追踪每次变更带来的回归(Regression)
  • 使用 pass@k 指标进行可靠性衡量

评测类型

能力评测(Capability Evals)

测试 Claude 是否能够完成之前无法完成的任务:

[CAPABILITY EVAL: feature-name]
任务:描述 Claude 应该完成的目标
成功标准:
  - [ ] 标准 1
  - [ ] 标准 2
  - [ ] 标准 3
预期输出:对预期结果的描述

This is the opening of the README. Read the full README on GitHub.

Commands Eval Harness provides

Slash commands named in this skill’s SKILL.md, listed in the order they first appear.

  • /eval

Frequently asked about Eval Harness

  • What else does xu-xiang publish alongside Eval Harness?

    Eval Harness is one of 25 skills that DirSkills catalogs from xu-xiang/everything-claude-code-zh, the repository it ships in. Its siblings there include API Design Patterns, Article Writing and Autonomous Loops. Each one is a separate skill with its own page in this directory, installs the same way Eval Harness does, and is maintained by xu-xiang in that same repository. The rest of the collection is listed on the xu-xiang/everything-claude-code-zh page.

  • How does Eval Harness compare to other AI Engineering skills?

    Eval Harness ranks #914 by stars among the 2451 AI Engineering skills in this catalog. The most-starred ones next to it are Architecture Decision Records, AI-First Engineering and Agentic OS. DirSkills ranks by the star count of the repository each skill ships in, so that order reflects how popular those repositories are rather than any review of Eval Harness against them. Open each page to compare what they document and how they install.

More from xu-xiang/everything-claude-code-zh

Eval Harness is one of 25 skills cataloged on DirSkills from xu-xiang/everything-claude-code-zh.

See all 25 skills
📐
2w ago

API Design Patterns

API Design Patterns provides conventions and best practices for designing consistent, developer-friendly REST APIs, including resource naming, HTTP status codes, pagination, filtering, error responses, versioning, and rate limiting.
DevOps
1.9K307
✍️
2w ago

Article Writing

Article Writing drafts long-form content such as articles, guides, tutorials, and newsletters in a specific voice captured from examples or brand guidelines. Use it when you need polished copy longer than a paragraph with consistent tone, structure, and credibility.
Writing
1.9K307
🔁
2w ago

Autonomous Loops

Autonomous Loops provides patterns and architectures for running Claude Code without human intervention, from simple claude -p pipelines to RFC-driven multi-agent DAG orchestration. Use it to set up autonomous development workflows, CI/CD pipelines, parallel agents, and quality gates.
Automation
1.9K307
⚙️
2w ago

Backend Patterns

Backend Patterns provides architectural patterns and best practices for building scalable server-side applications, covering REST API design, repository/service layers, database optimization, caching, and error handling. Use it when implementing Node.js, Express, or Next.js API routes.
DevOps
1.9K307
🛠️
2w ago

C++ Coding Standards

C++ Coding Standards provides a set of rules and best practices based on the C++ Core Guidelines for writing modern, safe, and idiomatic C++ code. Use it when writing, reviewing, or refactoring C++ code to enforce type safety, resource safety, immutability, and clarity.
Quality
1.9K307
🧪
2w ago

C++ Testing

C++ Testing guides writing and debugging C++17/20 tests with GoogleTest/Mock, CMake/CTest, coverage, and sanitizers. Use it when creating or fixing C++ tests, configuring CI test gates, or diagnosing flaky failures.
Quality
1.9K307