⚖️
AI EngineeringPython

Advanced Evaluation

by sickn33

Advanced Evaluation is an AI Engineering skill for Claude Code, published by sickn33 in agentic-awesome-skills.

44.5K stars6.5K forkson sickn33/agentic-awesome-skillsAdded 2026/07/19+1% in starsRepository updated 2026/08/05
agent-skillsagentic-skillsai-agent-skillsai-agentsai-codingai-workflowsantigravityantigravity-skillsclaude-codeclaude-code-skillscodex-clicodex-skillscursorcursor-skillsdeveloper-toolsgemini-cligemini-skillskiromcpskill-library
Install in seconds
Install Advanced Evaluation
Copy Advanced Evaluation into your Claude Code skills folder. Run the command in your terminal, or review the source on GitHub before installing.
terminal
npx degit https://github.com/sickn33/agentic-awesome-skills/tree/main/plugins/agentic-awesome-skills-claude/skills/advanced-evaluation ~/.claude/skills/advanced-evaluation

Requires Node.js. Downloads this skill only — not the rest of the repository — into your Claude Code skills folder.

Without Node.js

git clone https://github.com/sickn33/agentic-awesome-skills.git

Clones the whole repository, then copy the skill’s own directory into your skills folder yourself.

In this catalog

Source file
plugins/agentic-awesome-skills-claude/skills/advanced-evaluation/SKILL.md in sickn33/agentic-awesome-skills
Installs to
~/.claude/skills/advanced-evaluation
Collection
One of 26 skills cataloged from this repository
Category
AI Engineering3670 skills

What Advanced Evaluation does

Implements LLM-as-judge techniques for evaluating model outputs with direct scoring, pairwise comparison, and bias mitigation. Use when building automated evaluation pipelines or comparing AI-generated responses.

Advanced Evaluation is cataloged under AI Engineering on DirSkills. Advanced Evaluation comes from a repository tagged agent-skills, agentic-skills, ai-agent-skills, ai-agents and ai-coding.

Documentation

README

Advanced Evaluation

This skill covers production-grade techniques for evaluating LLM outputs using LLMs as judges. It synthesizes research from academic papers, industry practices, and practical implementation experience into actionable patterns for building reliable evaluation systems.

Key insight: LLM-as-a-Judge is not a single technique but a family of approaches, each suited to different evaluation contexts. Choosing the right approach and mitigating known biases is the core competency this skill develops.

When to Use

Activate this skill when:

This is the opening of the README. Read the full README on GitHub.

Frequently asked about Advanced Evaluation

  • What else does sickn33 publish alongside Advanced Evaluation?

    Advanced Evaluation is one of 26 skills that DirSkills catalogs from sickn33/agentic-awesome-skills, the repository it ships in. Its siblings there include 007 Security Audit, 2Slides Presentation Generator and 3D Web Experience. Each one is a separate skill with its own page in this directory, installs the same way Advanced Evaluation does, and is maintained by sickn33 in that same repository. The rest of the collection is listed on the sickn33/agentic-awesome-skills page.

  • How does Advanced Evaluation compare to other AI Engineering skills?

    Advanced Evaluation ranks #113 by stars among the 3670 AI Engineering skills in this catalog. The most-starred ones next to it are Architecture Decision Records, AI-First Engineering and Agentic OS. DirSkills ranks by the star count of the repository each skill ships in, so that order reflects how popular those repositories are rather than any review of Advanced Evaluation against them. Open each page to compare what they document and how they install.

More from sickn33/agentic-awesome-skills

Advanced Evaluation is one of 26 skills cataloged on DirSkills from sickn33/agentic-awesome-skills.

See all 26 skills