๐Ÿงช
QualityPython

Agent Testing Framework

by nWave-ai

Agent Testing Framework is a Quality skill for Claude Code, published by nWave-ai in nWave.

598 stars61 forkson nWave-ai/nWaveAdded 2026/08/25+1% in starsRepository updated 2026/06/27
agentic-aiagentic-codingagentic-frameworkagentic-workflowaiatddbddclaude-codeclaude-code-cliclaude-code-commandsclaude-code-hooksclaude-code-skillsclaude-code-subagentsdevopslean-uxopencodesoftware-architecturesoftware-craftmanshiptdd
Install in seconds
Install Agent Testing Framework
Copy Agent Testing Framework into your Claude Code skills folder. Run the command in your terminal, or review the source on GitHub before installing.
terminal
npx degit https://github.com/nWave-ai/nWave/tree/main/nWave/skills/nw-agent-testing ~/.claude/skills/nw-agent-testing

Requires Node.js. Downloads this skill only โ€” not the rest of the repository โ€” into your Claude Code skills folder.

Without Node.js

git clone https://github.com/nWave-ai/nWave.git

Clones the whole repository, then copy the skillโ€™s own directory into your skills folder yourself.

In this catalog

Source file
nWave/skills/nw-agent-testing/SKILL.md in nWave-ai/nWave
Installs to
~/.claude/skills/nw-agent-testing
Collection
One of 25 skills cataloged from this repository
Category
Quality โ€” 1354 skills

What Agent Testing Framework does

Agent Testing Framework defines a 5-layer approach for validating agent outputs, handoffs, and adversarial cases. It is used to check prompt injection resistance and other security constraints in Claude Code agents.

Agent Testing Framework is cataloged under Quality on DirSkills. Agent Testing Framework comes from a repository tagged agentic-ai, agentic-coding, agentic-framework, agentic-workflow and ai.

Documentation

README

Agent Testing Framework

5-Layer Testing Approach

Layer 1: Output Quality (Unit-Level)

Validate agent produces correct, well-structured outputs for typical inputs.

Test: Agent follows workflow phases | Outputs match expected format/structure | Domain-specific rules correctly applied | Token efficiency within bounds

How: Manual invocation with representative inputs. Check against acceptance criteria in agent description.

Layer 2: Integration / Handoff Validation

Validate correct input/output between agents in workflows.

Test: Input parsing handles upstream format | Output format matches downstream expectations | Error signals propagate correctly | Subagent mode activation works (skip greet, execute autonomously)

This is the opening of the README. Read the full README on GitHub.

Frequently asked about Agent Testing Framework

  • What else does nWave-ai publish alongside Agent Testing Framework?

    Agent Testing Framework is one of 25 skills that DirSkills catalogs from nWave-ai/nWave, the repository it ships in. Its siblings there include AT Completeness Check, Acceptance Test Critique Dimensions and Agent Creation Workflow. Each one is a separate skill with its own page in this directory, installs the same way Agent Testing Framework does, and is maintained by nWave-ai in that same repository. The rest of the collection is listed on the nWave-ai/nWave page.

  • How does Agent Testing Framework compare to other Quality skills?

    Agent Testing Framework ranks #1129 by stars among the 1354 Quality skills in this catalog. The most-starred ones next to it are Benchmark, Benchmark Optimization Loop and API Design Patterns. DirSkills ranks by the star count of the repository each skill ships in, so that order reflects how popular those repositories are rather than any review of Agent Testing Framework against them. Open each page to compare what they document and how they install.

More from nWave-ai/nWave

Agent Testing Framework is one of 25 skills cataloged on DirSkills from nWave-ai/nWave.

See all 25 skills โ†’
โœ…
6d ago

AT Completeness Check

AT Completeness Check gates acceptance-test sets against a canonical 7-category taxonomy and 15-item checklist. Use it to spot missing boundary, state, lifecycle, flag, negative, and environment cases.
Quality
59861
โœ…
6d ago

Acceptance Test Critique Dimensions

Acceptance Test Critique Dimensions reviews acceptance tests for coverage, GWT structure, business-language wording, observable assertions, and traceability. Use it when peer-reviewing scenarios or walking skeletons before handoff.
Quality
59861
๐Ÿค–
6d ago

Agent Creation Workflow

Agent Creation Workflow guides agent design from requirements analysis through validation and refinement. Use it when creating a new agent or modifying an existing one with clear responsibilities, tools, and quality gates.
AI Engineering
59861
โœ…
6d ago

Agent Quality Critique Dimensions

Agent Quality Critique Dimensions provides checks for Claude Code agent definitions, including template compliance, safety, examples, loading, and priority. Use it when reviewing an agent file for correctness and fit.
Quality
59861
โœ…
6d ago

Agent Quality Critique Dimensions

Agent Quality Critique Dimensions provides review criteria for Claude Code agent definitions, including template compliance, safety, examples, and loading behavior. Use it when validating an agent before release or refactor.
Quality
59861
๐Ÿ—๏ธ
6d ago

Architectural Styles Tradeoffs

Architectural Styles Tradeoffs helps choose between common software architecture styles using decision trees, comparison matrices, and combination patterns. Use it when selecting or evaluating an architecture for a system.
AI Engineering
59861