๐Ÿงช
AI EngineeringTypeScript

Benchmark Design

by Prism-Shadow

Benchmark Design is an AI Engineering skill for Claude Code, published by Prism-Shadow in penguin-harness.

1.5K stars150 forkson Prism-Shadow/penguin-harnessAdded 2026/08/19+21% in starsRepository updated 2026/08/19
agentagentic-aiaibuild-toolclaude-codedeepseekdeepseek-harnessdesktopharnessllmrsiself-evolving
Install in seconds
Install Benchmark Design
Copy Benchmark Design into your Claude Code skills folder. Run the command in your terminal, or review the source on GitHub before installing.
terminal
npx degit https://github.com/Prism-Shadow/penguin-harness/tree/main/packages/skills/skills/benchmark-design ~/.claude/skills/benchmark-design

Requires Node.js. Downloads this skill only โ€” not the rest of the repository โ€” into your Claude Code skills folder.

Without Node.js

git clone https://github.com/Prism-Shadow/penguin-harness.git

Clones the whole repository, then copy the skillโ€™s own directory into your skills folder yourself.

In this catalog

Source file
packages/skills/skills/benchmark-design/SKILL.md in Prism-Shadow/penguin-harness
Installs to
~/.claude/skills/benchmark-design
Collection
One of 21 skills cataloged from this repository
Category
AI Engineering โ€” 2451 skills

What Benchmark Design does

Benchmark Design builds and calibrates a multi-Case capability benchmark for one Test Agent and establishes a traceable Formal Baseline. Use it when you need to design, pilot, and freeze a benchmark without changing the test agent.

Benchmark Design is cataloged under AI Engineering on DirSkills. Benchmark Design comes from a repository tagged agent, agentic-ai, ai, build-tool and claude-code.

Documentation

README

Benchmark Design

Build a multi-Case Benchmark for one Test Agent, calibrate its difficulty with one Run per Case, and record the selected frozen Pilot as the Formal Baseline.

This Skill changes the Benchmark, never the Test Agent. It does not run or score the Test Agent. Delegate every evaluation with run_subagent, and tell each worker to use agent-evaluation. Stop after the Baseline; do not begin optimization.

Before you start

This is the opening of the README. Read the full README on GitHub.

Frequently asked about Benchmark Design

  • What else does Prism-Shadow publish alongside Benchmark Design?

    Benchmark Design is one of 21 skills that DirSkills catalogs from Prism-Shadow/penguin-harness, the repository it ships in. Its siblings there include Agent Creation, Agent Evaluation and Agent Optimization. Each one is a separate skill with its own page in this directory, installs the same way Benchmark Design does, and is maintained by Prism-Shadow in that same repository. The rest of the collection is listed on the Prism-Shadow/penguin-harness page.

  • How does Benchmark Design compare to other AI Engineering skills?

    Benchmark Design ranks #990 by stars among the 2451 AI Engineering skills in this catalog. The most-starred ones next to it are Architecture Decision Records, AI-First Engineering and Agentic OS. DirSkills ranks by the star count of the repository each skill ships in, so that order reflects how popular those repositories are rather than any review of Benchmark Design against them. Open each page to compare what they document and how they install.

More from Prism-Shadow/penguin-harness

Benchmark Design is one of 21 skills cataloged on DirSkills from Prism-Shadow/penguin-harness.

See all 21 skills โ†’
๐Ÿค–
2w ago

Agent Creation

Agent Creation turns a user requirement into a working agent configuration by writing AGENTS.md, setting identity metadata, and installing only needed skills. Use it to create or configure an Agent State under the agents directory.
AI Engineering
1.5K149
๐Ÿงช
2w ago

Agent Evaluation

Agent Evaluation runs one specified test agent on one benchmark case exactly once, privately scores the execution, and returns one protocol result. Use it when a run_subagent caller needs a single isolated agent evaluation without managing loops or scoreboards.
AI Engineering
1.5K149
๐Ÿงช
2w ago

Agent Optimization

Agent Optimization improves one Test Agent through an evidence โ†’ hypothesis โ†’ Candidate โ†’ evaluation โ†’ accept or rollback loop. It uses a frozen Benchmark, Scoreboard, and public Test Traces as black-box feedback while delegating all evaluation to an agent-evaluation subagent.
AI Engineering
1.5K149
๐Ÿค–
2w ago

AgentHub Models

AgentHub Models calls model APIs through @prismshadow/agenthub for streaming text, image generation, speech synthesis, embeddings, and the supported-model registry with one client. Use it when building AI apps that need a single TypeScript client for multiple model providers.
AI Engineering
1.5K150
๐Ÿ“ฝ๏ธ
2w ago

Bento Slides

Bento Slides creates and edits Bento presentations โ€” self-contained .bento.html decks whose document is JSON. Use it when the user wants a slide deck or presentation, whether from scratch, from source material, or by improving an existing file.
Frontend
1.5K150
๐Ÿ“Š
2w ago

Data Analysis

Data Analysis completes data-analysis tasks with bounded inspection, correct data semantics, native artifact handling, complete delivery, and risk-based verification. Use it when an agent must deliver a requested dataset or report artifact at an exact path with correct grain, units, and format.
Data
1.5K150