๐Ÿงช
AI EngineeringShell

Agent Evals

by BagelHole

Agent Evals is an AI Engineering skill for Claude Code, published by BagelHole in DevOps-Security-Agent-Skills.

709 stars89 forkson BagelHole/DevOps-Security-Agent-SkillsAdded 2026/08/24+1343% in starsRepository updated 2026/05/22
agent-skillsagentic-aiai-agentsawscheatsheetscompliancedevopsdockerinfrastructure-as-codekubernetessecuritysreterraform
Install in seconds
Install Agent Evals
Copy Agent Evals into your Claude Code skills folder. Run the command in your terminal, or review the source on GitHub before installing.
terminal
npx degit https://github.com/BagelHole/DevOps-Security-Agent-Skills/tree/main/devops/ai/agent-evals ~/.claude/skills/agent-evals

Requires Node.js. Downloads this skill only โ€” not the rest of the repository โ€” into your Claude Code skills folder.

Without Node.js

git clone https://github.com/BagelHole/DevOps-Security-Agent-Skills.git

Clones the whole repository, then copy the skillโ€™s own directory into your skills folder yourself.

In this catalog

Source file
devops/ai/agent-evals/SKILL.md in BagelHole/DevOps-Security-Agent-Skills
Installs to
~/.claude/skills/agent-evals
Collection
One of 25 skills cataloged from this repository
Category
AI Engineering โ€” 2451 skills

What Agent Evals does

Agent Evals builds repeatable tests for AI agent behavior using golden datasets, tool-call checks, safety prompts, and judge-based scoring. Use it to validate prompt changes, catch regressions, and gate deployments on quality.

Agent Evals is cataloged under AI Engineering on DirSkills. Agent Evals comes from a repository tagged agent-skills, agentic-ai, ai-agents, aws and cheatsheets.

Documentation

README

Agent Evals

Create repeatable checks so agent behavior improves safely over time.

When to Use This Skill

Use this skill when:

  • Shipping new agent features or changing prompts
  • Adding CI gates for agent quality and safety
  • Building regression suites for tool-calling agents
  • Measuring LLM output quality at scale
  • Validating RAG retrieval accuracy

Prerequisites

  • Python 3.10+
  • An LLM API key (OpenAI, Anthropic, etc.)
  • pytest or a custom eval harness
  • Optional: Braintrust, Promptfoo, or LangSmith account

Evaluation Layers

Unit Evals โ€” Prompt-Level Correctness

Test individual prompt โ†’ response quality:

This is the opening of the README. Read the full README on GitHub.

Frequently asked about Agent Evals

  • What else does BagelHole publish alongside Agent Evals?

    Agent Evals is one of 25 skills that DirSkills catalogs from BagelHole/DevOps-Security-Agent-Skills, the repository it ships in. Its siblings there include AI Pipeline Orchestration, AI SRE Incident Response and AWS CloudTrail. Each one is a separate skill with its own page in this directory, installs the same way Agent Evals does, and is maintained by BagelHole in that same repository. The rest of the collection is listed on the BagelHole/DevOps-Security-Agent-Skills page.

  • How does Agent Evals compare to other AI Engineering skills?

    Agent Evals ranks #1705 by stars among the 2451 AI Engineering skills in this catalog. The most-starred ones next to it are Architecture Decision Records, AI-First Engineering and Agentic OS. DirSkills ranks by the star count of the repository each skill ships in, so that order reflects how popular those repositories are rather than any review of Agent Evals against them. Open each page to compare what they document and how they install.

More from BagelHole/DevOps-Security-Agent-Skills

Agent Evals is one of 25 skills cataloged on DirSkills from BagelHole/DevOps-Security-Agent-Skills.

See all 25 skills โ†’
๐Ÿง 
1w ago

AI Pipeline Orchestration

AI Pipeline Orchestration builds reliable workflows for document ingestion, batch inference, model training, and RAG indexing. Use it to schedule recurring AI jobs and manage dependencies in production pipelines.
AI Engineering
70989
๐Ÿšจ
1w ago

AI SRE Incident Response

AI SRE Incident Response builds runbooks and alerting for LLM outages, quality regressions, safety incidents, and runaway cost events. Use it when AI services need SRE-style monitoring, rollback, and escalation procedures.
DevOps
70989
๐Ÿชต
1w ago

AWS CloudTrail

AWS CloudTrail configures audit logging for AWS account activity, including organization trails and event selectors. Use it to investigate incidents, meet compliance needs, and alert on sensitive API calls.
DevOps
70989
๐Ÿ”
1w ago

Access Review

Access Review conducts periodic access certifications and reviews across identity systems like AWS IAM, GitHub, and Okta. Use it to find stale accounts, unused permissions, and generate audit evidence for compliance reviews.
DevOps
70989
๐Ÿ“ˆ
1w ago

Agent Observability

Agent Observability instruments AI agents with logs, traces, metrics, token usage, latency, and cost telemetry. Use it to debug reliability issues, set SLOs, and monitor agent or LLM workloads.
AI Engineering
70989
๐Ÿ“ฆ
1w ago

Asset Inventory

Asset Inventory maintains an IT asset inventory and CMDB data for hardware, software, cloud resources, and related configuration. Use it when discovering assets, enforcing tags, or preparing compliance audits.
DevOps
70989