๐Ÿ“Š
QualityC#

Agent Benchmark Framework

by vibeeval

Agent Benchmark Framework is a Quality skill for Claude Code, published by vibeeval in vibecosystem.

528 stars44 forkson vibeeval/vibecosystemAdded 2026/08/26Repository updated 2026/08/08
ai-agentsai-software-teamanthropicautomationclaudeclaude-codeclaude-skillsdeveloper-toolsdevtoolshooksmulti-agentopen-sourceself-learningvibe-codingvibecoding
Install in seconds
Install Agent Benchmark Framework
Copy Agent Benchmark Framework into your Claude Code skills folder. Run the command in your terminal, or review the source on GitHub before installing.
terminal
npx degit https://github.com/vibeeval/vibecosystem/tree/main/skills/agent-benchmark ~/.claude/skills/agent-benchmark

Requires Node.js. Downloads this skill only โ€” not the rest of the repository โ€” into your Claude Code skills folder.

Without Node.js

git clone https://github.com/vibeeval/vibecosystem.git

Clones the whole repository, then copy the skillโ€™s own directory into your skills folder yourself.

In this catalog

Source file
skills/agent-benchmark/SKILL.md in vibeeval/vibecosystem
Installs to
~/.claude/skills/agent-benchmark
Collection
One of 25 skills cataloged from this repository
Category
Quality โ€” 1354 skills

What Agent Benchmark Framework does

Agent Benchmark Framework measures agent response quality over time and compares runs against baselines. Use it before and after agent changes, for audits, and to catch regressions before release.

Agent Benchmark Framework is cataloged under Quality on DirSkills. Agent Benchmark Framework comes from a repository tagged ai-agents, ai-software-team, anthropic, automation and claude.

Documentation

README

Agent Benchmark Framework

Without benchmarks, we cannot know whether agent changes improve or degrade quality. This skill defines how to measure, track, and protect agent performance.

When to Activate

  • Before and after modifying any agent definition file
  • When adding a new skill that an agent depends on
  • Periodic quality audits (weekly/monthly)
  • When a user reports degraded agent output
  • Before promoting an agent from experimental to production

Core Concepts

Why Benchmarks Matter

Agent quality degrades silently. A prompt tweak that improves one response can break ten others. Without a baseline to compare against, every change is a guess. Benchmarks make quality visible and regressions detectable.

Benchmark Types

This is the opening of the README. Read the full README on GitHub.

Frequently asked about Agent Benchmark Framework

  • What else does vibeeval publish alongside Agent Benchmark Framework?

    Agent Benchmark Framework is one of 25 skills that DirSkills catalogs from vibeeval/vibecosystem, the repository it ships in. Its siblings there include AI Slop Cleaner, API Patterns and API Versioning Patterns. Each one is a separate skill with its own page in this directory, installs the same way Agent Benchmark Framework does, and is maintained by vibeeval in that same repository. The rest of the collection is listed on the vibeeval/vibecosystem page.

  • How does Agent Benchmark Framework compare to other Quality skills?

    Agent Benchmark Framework ranks #1194 by stars among the 1354 Quality skills in this catalog. The most-starred ones next to it are Benchmark, Benchmark Optimization Loop and API Design Patterns. DirSkills ranks by the star count of the repository each skill ships in, so that order reflects how popular those repositories are rather than any review of Agent Benchmark Framework against them. Open each page to compare what they document and how they install.

More from vibeeval/vibecosystem

Agent Benchmark Framework is one of 25 skills cataloged on DirSkills from vibeeval/vibecosystem.

See all 25 skills โ†’
๐Ÿงน
5d ago

AI Slop Cleaner

AI Slop Cleaner removes dead imports, debug code, redundant comments, and other AI-generated bloat after a feature is implemented. It works in small passes and runs tests after each pass to catch regressions.
Quality
52844
๐Ÿ”Œ
5d ago

API Patterns

API Patterns covers REST and GraphQL design, versioning, validation, pagination, error responses, and contract testing. Use it when building consistent APIs that need stable schemas and predictable behavior.
AI Engineering
52844
๐Ÿงฉ
5d ago

API Versioning Patterns

API Versioning Patterns defines versioning strategies, breaking-change checks, deprecation phases, and migration guide structure for APIs. Use it when you need to ship compatible changes and retire old versions safely.
DevOps
52844
โ˜๏ธ
5d ago

AWS Patterns

AWS Patterns covers Lambda best practices, S3 event handling, SQS/SNS fanout, and DynamoDB single-table access patterns for serverless AWS workloads. Use it when designing or reviewing managed-service backend architectures.
DevOps
52844
โ™ฟ
5d ago

Accessibility Patterns

Accessibility Patterns covers WCAG 2.2 AA requirements, ARIA patterns, keyboard navigation, and screen reader-friendly UI practices. Use it when building or auditing interfaces for accessibility.
Frontend
52844
โ™ฟ
5d ago

Accessibility Testing

Accessibility Testing adds axe-core checks, WCAG 2.2 AA checklists, keyboard navigation tests, screen reader patterns, and ARIA validation for web apps. Use it to verify components and end-to-end flows meet common accessibility requirements.
Quality
52844