๐Ÿ“Š
AI EngineeringShell

LLM Evaluation Harness

by sangrokjung

LLM Evaluation Harness is an AI Engineering skill for Claude Code, published by sangrokjung in claude-forge.

806 stars175 forkson sangrokjung/claude-forgeAdded 2026/08/22Repository updated 2026/08/21
agentsai-assistantai-codingai-frameworkai-pair-programminganthropicautomationclaude-codeclaude-code-agentscli-toolsdeveloper-experiencedeveloper-toolsdotfileshooksllmmacosproductivityshellslash-commandsworkflow
Install in seconds
Install LLM Evaluation Harness
Copy LLM Evaluation Harness into your Claude Code skills folder. Run the command in your terminal, or review the source on GitHub before installing.
terminal
npx degit https://github.com/sangrokjung/claude-forge/tree/main/skills/evaluating-llms-harness ~/.claude/skills/evaluating-llms-harness

Requires Node.js. Downloads this skill only โ€” not the rest of the repository โ€” into your Claude Code skills folder.

Without Node.js

git clone https://github.com/sangrokjung/claude-forge.git

Clones the whole repository, then copy the skillโ€™s own directory into your skills folder yourself.

In this catalog

Source file
skills/evaluating-llms-harness/SKILL.md in sangrokjung/claude-forge
Installs to
~/.claude/skills/evaluating-llms-harness
Collection
One of 25 skills cataloged from this repository
Category
AI Engineering โ€” 2451 skills

What LLM Evaluation Harness does

LLM Evaluation Harness evaluates language models on 60+ academic benchmarks such as MMLU, GSM8K, and HumanEval. Use it to compare models, report results, or track training checkpoints over time.

LLM Evaluation Harness is cataloged under AI Engineering on DirSkills. LLM Evaluation Harness comes from a repository tagged agents, ai-assistant, ai-coding, ai-framework and ai-pair-programming.

Documentation

README

lm-evaluation-harness - LLM Benchmarking

Quick start

lm-evaluation-harness evaluates LLMs across 60+ academic benchmarks using standardized prompts and metrics.

Installation:

pip install lm-eval

Evaluate any HuggingFace model:

lm_eval --model hf \
  --model_args pretrained=meta-llama/Llama-2-7b-hf \
  --tasks mmlu,gsm8k,hellaswag \
  --device cuda:0 \
  --batch_size 8

View available tasks:

lm_eval --tasks list

Common workflows

Workflow 1: Standard benchmark evaluation

Evaluate model on core benchmarks (MMLU, GSM8K, HumanEval).

Copy this checklist:

Benchmark Evaluation:
- [ ] Step 1: Choose benchmark suite
- [ ] Step 2: Configure model
- [ ] Step 3: Run evaluation
- [ ] Step 4: Analyze results

This is the opening of the README. Read the full README on GitHub.

Frequently asked about LLM Evaluation Harness

  • What else does sangrokjung publish alongside LLM Evaluation Harness?

    LLM Evaluation Harness is one of 25 skills that DirSkills catalogs from sangrokjung/claude-forge, the repository it ships in. Its siblings there include BigCode Evaluation Harness, Blind Spot Pass and Build System. Each one is a separate skill with its own page in this directory, installs the same way LLM Evaluation Harness does, and is maintained by sangrokjung in that same repository. The rest of the collection is listed on the sangrokjung/claude-forge page.

  • How does LLM Evaluation Harness compare to other AI Engineering skills?

    LLM Evaluation Harness ranks #1584 by stars among the 2451 AI Engineering skills in this catalog. The most-starred ones next to it are Architecture Decision Records, AI-First Engineering and Agentic OS. DirSkills ranks by the star count of the repository each skill ships in, so that order reflects how popular those repositories are rather than any review of LLM Evaluation Harness against them. Open each page to compare what they document and how they install.

More from sangrokjung/claude-forge

LLM Evaluation Harness is one of 25 skills cataloged on DirSkills from sangrokjung/claude-forge.

See all 25 skills โ†’
๐Ÿงช
1w ago

BigCode Evaluation Harness

BigCode Evaluation Harness evaluates code generation models on HumanEval, MBPP, MultiPL-E, and other benchmarks with pass@k metrics. Use it to compare coding ability, multi-language support, or overall code generation quality.
Quality
806175
๐Ÿงญ
1w ago

Blind Spot Pass

Blind Spot Pass surfaces the unknown unknowns in a domain before you start work, so you can prompt and decide with enough context. Use it when you are new to a field or handing off unfamiliar work and need a brief briefing first.
AI Engineering
806175
๐Ÿ› ๏ธ
1w ago

Build System

Build System detects a project's build tool and runs the right build or test command. Use it when setting up, building, or testing projects with npm, Python, Gradle, Maven, Cargo, Go, or Make.
DevOps
806175
๐Ÿงฉ
1w ago

Cache Components

Cache Components gives guidance for Next.js Cache Components and Partial Prerendering. Use it when `cacheComponents: true` is enabled to decide what to cache, stream, tag, and invalidate in Server Components.
Frontend
806175
๐Ÿค–
1w ago

Claude Code Agent

Claude Code Agent guides Claude Code projects with spec-first planning, context engineering, sub-agents, and post-dev verification. Use it when writing CLAUDE.md or spec.md, dispatching parallel agent work, or running the handoff and docs sync workflow.
AI Engineering
806175
๐Ÿง 
1w ago

Continuous Learning V2

Continuous Learning V2 turns Claude Code sessions into atomic instincts using hooks, confidence scoring, and background pattern detection. Use it to capture repeated behaviors and evolve them into skills, commands, or agents.
AI Engineering
806175