πŸ§ͺ
AI EngineeringPython

Hugging Face Community Evals

by waybarrios

Hugging Face Community Evals is an AI Engineering skill for Claude Code, published by waybarrios in opencode-power-pack.

482 stars39 forkson waybarrios/opencode-power-packAdded 2026/08/26+2% in starsRepository updated 2026/08/24
ai-agentsanthropicclaude-codecode-reviewcodexcodex-clideveloper-toolsgradiohuggingfacellm-toolsmachine-learningml-trainingopenai-codexopencodepipi-coding-agentpi-packagepluginsecurity-auditskills
Install in seconds
Install Hugging Face Community Evals
Copy Hugging Face Community Evals into your Claude Code skills folder. Run the command in your terminal, or review the source on GitHub before installing.
terminal
npx degit https://github.com/waybarrios/opencode-power-pack/tree/main/skills/huggingface-community-evals ~/.claude/skills/huggingface-community-evals

Requires Node.js. Downloads this skill only β€” not the rest of the repository β€” into your Claude Code skills folder.

Without Node.js

git clone https://github.com/waybarrios/opencode-power-pack.git

Clones the whole repository, then copy the skill’s own directory into your skills folder yourself.

In this catalog

Source file
skills/huggingface-community-evals/SKILL.md in waybarrios/opencode-power-pack
Installs to
~/.claude/skills/huggingface-community-evals
Collection
One of 25 skills cataloged from this repository
Category
AI Engineering β€” 2451 skills

What Hugging Face Community Evals does

Hugging Face Community Evals runs inspect-ai and lighteval evaluations for Hugging Face Hub models on local hardware. Use it to choose an inference backend, smoke-test tasks, and run GPU evals with vLLM, Transformers, or accelerate.

Hugging Face Community Evals is cataloged under AI Engineering on DirSkills. Hugging Face Community Evals comes from a repository tagged ai-agents, anthropic, claude-code, code-review and codex.

Documentation

README

Overview

This skill is for running evaluations against models on the Hugging Face Hub on local hardware.

It covers:

  • inspect-ai with local inference
  • lighteval with local inference
  • choosing between vllm, Hugging Face Transformers, and accelerate
  • smoke tests, task selection, and backend fallback strategy

It does not cover:

  • Hugging Face Jobs orchestration
  • model-card or model-index edits
  • README table extraction
  • Artificial Analysis imports
  • .eval_results generation or publishing
  • PR creation or community-evals automation

If the user wants to run the same eval remotely on Hugging Face Jobs, submit the same script via hf jobs uv run (CLI) or the hf_jobs() MCP tool if configured, for remote GPU execution.

This is the opening of the README. Read the full README on GitHub.

Frequently asked about Hugging Face Community Evals

  • What else does waybarrios publish alongside Hugging Face Community Evals?

    Hugging Face Community Evals is one of 25 skills that DirSkills catalogs from waybarrios/opencode-power-pack, the repository it ships in. Its siblings there include AI Slop Rubric, AWS Context Discovery and Agentic Actions Auditor. Each one is a separate skill with its own page in this directory, installs the same way Hugging Face Community Evals does, and is maintained by waybarrios in that same repository. The rest of the collection is listed on the waybarrios/opencode-power-pack page.

  • How does Hugging Face Community Evals compare to other AI Engineering skills?

    Hugging Face Community Evals ranks #2205 by stars among the 2451 AI Engineering skills in this catalog. The most-starred ones next to it are Architecture Decision Records, AI-First Engineering and Agentic OS. DirSkills ranks by the star count of the repository each skill ships in, so that order reflects how popular those repositories are rather than any review of Hugging Face Community Evals against them. Open each page to compare what they document and how they install.

More from waybarrios/opencode-power-pack

Hugging Face Community Evals is one of 25 skills cataloged on DirSkills from waybarrios/opencode-power-pack.

See all 25 skills β†’
🧩
5d ago

AI Slop Rubric

AI Slop Rubric defines observable signs of generic or poorly grounded interface design, with severity levels, evidence needs, and repair actions. Use it to review marketing sites, product pages, dashboards, portfolios, and e-commerce layouts.
Frontend
48239
☁️
5d ago

AWS Context Discovery

AWS Context Discovery reads the local AWS profile, region, account, and caller identity before AWS or SageMaker work. Use it when cloud context is needed and you want to avoid guessing configuration.
DevOps
48239
πŸ›‘οΈ
5d ago

Agentic Actions Auditor

Agentic Actions Auditor reviews GitHub Actions workflows that invoke AI agents for prompt injection, unsafe interpolation, sandbox gaps, and permissive actor rules. Use it when auditing agentic CI workflows and cross-file references in Actions configs.
Quality
48239
πŸ“
5d ago

Agents MD Revise

Agents MD Revise captures recurring learnings from a session into AGENTS.md, CLAUDE.md, or a local override so future sessions have the same context. Use it when you need to remember project-specific commands, conventions, or gotchas.
Writing
48239
🧭
5d ago

Agents Md Improver

Agents Md Improver audits AGENTS.md, CLAUDE.md, and related rules files for coverage, clarity, and freshness. Use it when project rules need review or targeted updates after code changes.
AI Engineering
48239
πŸ—οΈ
5d ago

Code Architect

Code Architect analyzes an existing codebase’s patterns and conventions to produce an implementation blueprint for a non-trivial feature. Use it when planning architecture, file changes, component design, data flow, and build steps.
AI Engineering
48239