⚖️
AI EngineeringPython

Arize Evaluator

by github

Arize Evaluator is an AI Engineering skill for Claude Code, published by github in awesome-copilot.

37.8K stars4.8K forkson github/awesome-copilotAdded 2026/08/13Repository updated 2026/08/13
agent-skillsagentsaiawesomecustom-agentsgithub-copilothacktoberfestprompt-engineering
Install in seconds
Install Arize Evaluator
Copy Arize Evaluator into your Claude Code skills folder. Run the command in your terminal, or review the source on GitHub before installing.
terminal
npx degit https://github.com/github/awesome-copilot/tree/main/skills/arize-evaluator ~/.claude/skills/arize-evaluator

Requires Node.js. Downloads this skill only — not the rest of the repository — into your Claude Code skills folder.

Without Node.js

git clone https://github.com/github/awesome-copilot.git

Clones the whole repository, then copy the skill’s own directory into your skills folder yourself.

In this catalog

Source file
skills/arize-evaluator/SKILL.md in github/awesome-copilot
Installs to
~/.claude/skills/arize-evaluator
Collection
One of 25 skills cataloged from this repository
Category
AI Engineering2451 skills

What Arize Evaluator does

Arize Evaluator creates and runs LLM-as-judge evaluators on Arize, including setting up evaluator templates, running evaluations on spans or experiments, and managing tasks and column mappings. Use it when working with hallucination, faithfulness, correctness, relevance, or other LLM quality metrics.

Arize Evaluator is cataloged under AI Engineering on DirSkills. Arize Evaluator comes from a repository tagged agent-skills, agents, ai, awesome and custom-agents.

Documentation

README

Arize Evaluator Skill

SPACE — All --space flags and the ARIZE_SPACE env var accept a space name (e.g., my-workspace) or a base64 space ID (e.g., U3BhY2U6...). Find yours with ax spaces list.

This skill covers designing, creating, and running LLM-as-judge evaluators on Arize. An evaluator defines the judge; a task is how you run it against real data.


Prerequisites

Proceed directly with the task — run the ax command you need. Do NOT check versions, env vars, or profiles upfront.

This is the opening of the README. Read the full README on GitHub.

Frequently asked about Arize Evaluator

  • What else does github publish alongside Arize Evaluator?

    Arize Evaluator is one of 25 skills that DirSkills catalogs from github/awesome-copilot, the repository it ships in. Its siblings there include AI Prompt Engineering Safety Review, AI Readiness Assessment and AI Ready. Each one is a separate skill with its own page in this directory, installs the same way Arize Evaluator does, and is maintained by github in that same repository. The rest of the collection is listed on the github/awesome-copilot page.

  • How does Arize Evaluator compare to other AI Engineering skills?

    Arize Evaluator ranks #156 by stars among the 2451 AI Engineering skills in this catalog. The most-starred ones next to it are Architecture Decision Records, AI-First Engineering and Agentic OS. DirSkills ranks by the star count of the repository each skill ships in, so that order reflects how popular those repositories are rather than any review of Arize Evaluator against them. Open each page to compare what they document and how they install.

More from github/awesome-copilot

Arize Evaluator is one of 25 skills cataloged on DirSkills from github/awesome-copilot.

See all 25 skills
🛡️
3w ago

AI Prompt Engineering Safety Review

AI Prompt Engineering Safety Review analyzes AI prompts for safety, bias, security vulnerabilities, and effectiveness, then provides structured improvement recommendations using systematic frameworks and testing methodologies.
AI Engineering
37.8K4.8K
📊
3w ago

AI Readiness Assessment

AI Readiness Assessment runs the AgentRC readiness scan on a repository and generates a static HTML dashboard at reports/index.html with maturity level, score, pillar breakdowns, and remediation plan. Use it when asked to audit, score, or assess the AI readiness of a codebase.
AI Engineering
37.8K4.8K
🤖
3w ago

AI Ready

AI Ready guides users to install and review the ai-ready skill. After installation, the skill analyzes the codebase and generates AGENTS.md, copilot-instructions.md, CI workflows, and issue templates tailored to the stack.
AI Engineering
37.8K4.8K
🤖
3w ago

AI Team Orchestration

AI Team Orchestration bootstraps and runs a lightweight multi-agent development team with producer, developer, and optional QA roles. Use it for project planning, implementation, independent testing, brainstorming, and preserving context across sessions.
AI Engineering
37.8K4.8K
📚
3w ago

Acquire Codebase Knowledge

Acquire Codebase Knowledge scans a repository and produces seven documentation files covering stack, structure, architecture, conventions, integrations, testing, and concerns. Use it when you need to map, document, or onboard into an existing codebase.
Writing
37.8K4.8K
📊
3w ago

Ad Campaign Analyzer

Ad Campaign Analyzer turns raw ad campaign performance data into clear decisions on what to cut, scale, and test. Use it to diagnose wasted ad spend, identify top performers, check statistical significance, and reallocate budget across channels.
Data
37.8K4.8K