๐Ÿงช
QualityShell

BigCode Evaluation Harness

by sangrokjung

BigCode Evaluation Harness is a Quality skill for Claude Code, published by sangrokjung in claude-forge.

806 stars175 forkson sangrokjung/claude-forgeAdded 2026/08/22Repository updated 2026/08/21
agentsai-assistantai-codingai-frameworkai-pair-programminganthropicautomationclaude-codeclaude-code-agentscli-toolsdeveloper-experiencedeveloper-toolsdotfileshooksllmmacosproductivityshellslash-commandsworkflow
Install in seconds
Install BigCode Evaluation Harness
Copy BigCode Evaluation Harness into your Claude Code skills folder. Run the command in your terminal, or review the source on GitHub before installing.
terminal
npx degit https://github.com/sangrokjung/claude-forge/tree/main/skills/evaluating-code-models ~/.claude/skills/evaluating-code-models

Requires Node.js. Downloads this skill only โ€” not the rest of the repository โ€” into your Claude Code skills folder.

Without Node.js

git clone https://github.com/sangrokjung/claude-forge.git

Clones the whole repository, then copy the skillโ€™s own directory into your skills folder yourself.

In this catalog

Source file
skills/evaluating-code-models/SKILL.md in sangrokjung/claude-forge
Installs to
~/.claude/skills/evaluating-code-models
Collection
One of 25 skills cataloged from this repository
Category
Quality โ€” 1354 skills

What BigCode Evaluation Harness does

BigCode Evaluation Harness evaluates code generation models on HumanEval, MBPP, MultiPL-E, and other benchmarks with pass@k metrics. Use it to compare coding ability, multi-language support, or overall code generation quality.

BigCode Evaluation Harness is cataloged under Quality on DirSkills. BigCode Evaluation Harness comes from a repository tagged agents, ai-assistant, ai-coding, ai-framework and ai-pair-programming.

Documentation

README

BigCode Evaluation Harness - Code Model Benchmarking

Quick Start

BigCode Evaluation Harness evaluates code generation models across 15+ benchmarks including HumanEval, MBPP, and MultiPL-E (18 languages).

Installation:

git clone https://github.com/bigcode-project/bigcode-evaluation-harness.git
cd bigcode-evaluation-harness
pip install -e .
accelerate config

Evaluate on HumanEval:

accelerate launch main.py \
  --model bigcode/starcoder2-7b \
  --tasks humaneval \
  --max_length_generation 512 \
  --temperature 0.2 \
  --n_samples 20 \
  --batch_size 10 \
  --allow_code_execution \
  --save_generations

View available tasks:

python -c "from bigcode_eval.tasks import ALL_TASKS; print(ALL_TASKS)"

Common Workflows

This is the opening of the README. Read the full README on GitHub.

Frequently asked about BigCode Evaluation Harness

  • What else does sangrokjung publish alongside BigCode Evaluation Harness?

    BigCode Evaluation Harness is one of 25 skills that DirSkills catalogs from sangrokjung/claude-forge, the repository it ships in. Its siblings there include Blind Spot Pass, Build System and Cache Components. Each one is a separate skill with its own page in this directory, installs the same way BigCode Evaluation Harness does, and is maintained by sangrokjung in that same repository. The rest of the collection is listed on the sangrokjung/claude-forge page.

  • How does BigCode Evaluation Harness compare to other Quality skills?

    BigCode Evaluation Harness ranks #945 by stars among the 1354 Quality skills in this catalog. The most-starred ones next to it are Benchmark, Benchmark Optimization Loop and API Design Patterns. DirSkills ranks by the star count of the repository each skill ships in, so that order reflects how popular those repositories are rather than any review of BigCode Evaluation Harness against them. Open each page to compare what they document and how they install.

More from sangrokjung/claude-forge

BigCode Evaluation Harness is one of 25 skills cataloged on DirSkills from sangrokjung/claude-forge.

See all 25 skills โ†’
๐Ÿงญ
1w ago

Blind Spot Pass

Blind Spot Pass surfaces the unknown unknowns in a domain before you start work, so you can prompt and decide with enough context. Use it when you are new to a field or handing off unfamiliar work and need a brief briefing first.
AI Engineering
806175
๐Ÿ› ๏ธ
1w ago

Build System

Build System detects a project's build tool and runs the right build or test command. Use it when setting up, building, or testing projects with npm, Python, Gradle, Maven, Cargo, Go, or Make.
DevOps
806175
๐Ÿงฉ
1w ago

Cache Components

Cache Components gives guidance for Next.js Cache Components and Partial Prerendering. Use it when `cacheComponents: true` is enabled to decide what to cache, stream, tag, and invalidate in Server Components.
Frontend
806175
๐Ÿค–
1w ago

Claude Code Agent

Claude Code Agent guides Claude Code projects with spec-first planning, context engineering, sub-agents, and post-dev verification. Use it when writing CLAUDE.md or spec.md, dispatching parallel agent work, or running the handoff and docs sync workflow.
AI Engineering
806175
๐Ÿง 
1w ago

Continuous Learning V2

Continuous Learning V2 turns Claude Code sessions into atomic instincts using hooks, confidence scoring, and background pattern detection. Use it to capture repeated behaviors and evolve them into skills, commands, or agents.
AI Engineering
806175
๐Ÿ›
1w ago

Debugging Strategies

Debugging Strategies helps you reproduce bugs, narrow causes, and verify fixes with systematic techniques and debugging tools. Use it for crashes, performance issues, memory leaks, or unfamiliar codepaths.
Quality
806175