๐Ÿงช
AI EngineeringTypeScript

Agent Evaluation

by Prism-Shadow

Agent Evaluation is an AI Engineering skill for Claude Code, published by Prism-Shadow in penguin-harness.

1.5K stars149 forkson Prism-Shadow/penguin-harnessAdded 2026/08/19+21% in starsRepository updated 2026/08/19
agentagentic-aiaibuild-toolclaude-codedeepseekdeepseek-harnessdesktopharnessllmrsiself-evolving
Install in seconds
Install Agent Evaluation
Copy Agent Evaluation into your Claude Code skills folder. Run the command in your terminal, or review the source on GitHub before installing.
terminal
npx degit https://github.com/Prism-Shadow/penguin-harness/tree/main/packages/skills/skills/agent-evaluation ~/.claude/skills/agent-evaluation

Requires Node.js. Downloads this skill only โ€” not the rest of the repository โ€” into your Claude Code skills folder.

Without Node.js

git clone https://github.com/Prism-Shadow/penguin-harness.git

Clones the whole repository, then copy the skillโ€™s own directory into your skills folder yourself.

In this catalog

Source file
packages/skills/skills/agent-evaluation/SKILL.md in Prism-Shadow/penguin-harness
Installs to
~/.claude/skills/agent-evaluation
Collection
One of 21 skills cataloged from this repository
Category
AI Engineering โ€” 2451 skills

What Agent Evaluation does

Agent Evaluation runs one specified test agent on one benchmark case exactly once, privately scores the execution, and returns one protocol result. Use it when a run_subagent caller needs a single isolated agent evaluation without managing loops or scoreboards.

Agent Evaluation is cataloged under AI Engineering on DirSkills. Agent Evaluation comes from a repository tagged agent, agentic-ai, ai, build-tool and claude-code.

Documentation

README

Agent Evaluation

Handle one evaluation request from a run_subagent caller: run the specified Test Agent on one Benchmark Case once, score that execution privately, and return one protocol result.

The top-level Benchmark Designer or Optimizer owns all Case and Run loops, concurrency, and follow-up handling. This worker handles no other Case or Run, launches no evaluator or subagent, modifies no Agent or Benchmark, and never writes scoreboard.yaml. Use the Penguin CLI only to launch the specified Test Agent; do not use it to create another phase, designer, optimizer, or evaluator.

This is the opening of the README. Read the full README on GitHub.

Frequently asked about Agent Evaluation

  • What else does Prism-Shadow publish alongside Agent Evaluation?

    Agent Evaluation is one of 21 skills that DirSkills catalogs from Prism-Shadow/penguin-harness, the repository it ships in. Its siblings there include Agent Creation, Agent Optimization and AgentHub Models. Each one is a separate skill with its own page in this directory, installs the same way Agent Evaluation does, and is maintained by Prism-Shadow in that same repository. The rest of the collection is listed on the Prism-Shadow/penguin-harness page.

  • How does Agent Evaluation compare to other AI Engineering skills?

    Agent Evaluation ranks #998 by stars among the 2451 AI Engineering skills in this catalog. The most-starred ones next to it are Architecture Decision Records, AI-First Engineering and Agentic OS. DirSkills ranks by the star count of the repository each skill ships in, so that order reflects how popular those repositories are rather than any review of Agent Evaluation against them. Open each page to compare what they document and how they install.

More from Prism-Shadow/penguin-harness

Agent Evaluation is one of 21 skills cataloged on DirSkills from Prism-Shadow/penguin-harness.

See all 21 skills โ†’
๐Ÿค–
2w ago

Agent Creation

Agent Creation turns a user requirement into a working agent configuration by writing AGENTS.md, setting identity metadata, and installing only needed skills. Use it to create or configure an Agent State under the agents directory.
AI Engineering
1.5K149
๐Ÿงช
2w ago

Agent Optimization

Agent Optimization improves one Test Agent through an evidence โ†’ hypothesis โ†’ Candidate โ†’ evaluation โ†’ accept or rollback loop. It uses a frozen Benchmark, Scoreboard, and public Test Traces as black-box feedback while delegating all evaluation to an agent-evaluation subagent.
AI Engineering
1.5K149
๐Ÿค–
2w ago

AgentHub Models

AgentHub Models calls model APIs through @prismshadow/agenthub for streaming text, image generation, speech synthesis, embeddings, and the supported-model registry with one client. Use it when building AI apps that need a single TypeScript client for multiple model providers.
AI Engineering
1.5K150
๐Ÿงช
2w ago

Benchmark Design

Benchmark Design builds and calibrates a multi-Case capability benchmark for one Test Agent and establishes a traceable Formal Baseline. Use it when you need to design, pilot, and freeze a benchmark without changing the test agent.
AI Engineering
1.5K150
๐Ÿ“ฝ๏ธ
2w ago

Bento Slides

Bento Slides creates and edits Bento presentations โ€” self-contained .bento.html decks whose document is JSON. Use it when the user wants a slide deck or presentation, whether from scratch, from source material, or by improving an existing file.
Frontend
1.5K150
๐Ÿ“Š
2w ago

Data Analysis

Data Analysis completes data-analysis tasks with bounded inspection, correct data semantics, native artifact handling, complete delivery, and risk-based verification. Use it when an agent must deliver a requested dataset or report artifact at an exact path with correct grain, units, and format.
Data
1.5K150