๐Ÿ”
AI EngineeringPython

Karpathy Autoresearch

by gaasher

Karpathy Autoresearch is an AI Engineering skill for Claude Code, published by gaasher in Agent-Loop-Skills.

166 stars19 forkson gaasher/Agent-Loop-SkillsAdded 2026/09/08+2% in starsRepository updated 2026/06/30
agent-skillsagentic-loopsagentic-workflowsai-agentsanthropicautoresearchclaudeclaude-codedata-analysisliterature-reviewllm-agentsmachine-learningml-autoresearchopen-sourceprompt-engineeringred-teamingscientific-writingskillssubagents
Install in seconds
Install Karpathy Autoresearch
Copy Karpathy Autoresearch into your Claude Code skills folder. Run the command in your terminal, or review the source on GitHub before installing.
terminal
npx degit https://github.com/gaasher/Agent-Loop-Skills/tree/main/loops/karpathy ~/.claude/skills/karpathy

Requires Node.js. Downloads this skill only โ€” not the rest of the repository โ€” into your Claude Code skills folder.

Without Node.js

git clone https://github.com/gaasher/Agent-Loop-Skills.git

Clones the whole repository, then copy the skillโ€™s own directory into your skills folder yourself.

In this catalog

Source file
loops/karpathy/SKILL.md in gaasher/Agent-Loop-Skills
Installs to
~/.claude/skills/karpathy
Collection
One of 25 skills cataloged from this repository
Category
AI Engineering โ€” 3670 skills

What Karpathy Autoresearch does

Karpathy Autoresearch runs a fully autonomous loop that edits training code, executes the run, and keeps changes only when a single scalar metric improves. Use it for hands-off ML experimentation where the agent loops until interrupted.

Karpathy Autoresearch is cataloged under AI Engineering on DirSkills. Karpathy Autoresearch comes from a repository tagged agent-skills, agentic-loops, agentic-workflows, ai-agents and anthropic.

Documentation

README

Karpathy Autoresearch

This is an experiment to have the LLM do its own research. You are a completely autonomous researcher: you hack the training code with an idea, run it, keep the change if the metric improves and revert it if it doesn't, advancing a branch as you go โ€” and you repeat forever, until the human interrupts you. The artifact is the <editable_files>; the feedback signal is one scalar <metric> (lower is better, e.g. val_bpb) read from the run. Training runs in the user's own environment via <run_cmd> โ€” this skill installs nothing and imports nothing; it edits code, shells out, and reads the metric from the log.

This is the opening of the README. Read the full README on GitHub.

Frequently asked about Karpathy Autoresearch

  • What else does gaasher publish alongside Karpathy Autoresearch?

    Karpathy Autoresearch is one of 25 skills that DirSkills catalogs from gaasher/Agent-Loop-Skills, the repository it ships in. Its siblings there include Alpha Evolve, Anomaly Investigation and Blue Team. Each one is a separate skill with its own page in this directory, installs the same way Karpathy Autoresearch does, and is maintained by gaasher in that same repository. The rest of the collection is listed on the gaasher/Agent-Loop-Skills page.

  • How does Karpathy Autoresearch compare to other AI Engineering skills?

    Karpathy Autoresearch ranks #3330 by stars among the 3670 AI Engineering skills in this catalog. The most-starred ones next to it are Architecture Decision Records, AI-First Engineering and Agentic OS. DirSkills ranks by the star count of the repository each skill ships in, so that order reflects how popular those repositories are rather than any review of Karpathy Autoresearch against them. Open each page to compare what they document and how they install.

More from gaasher/Agent-Loop-Skills

Karpathy Autoresearch is one of 25 skills cataloged on DirSkills from gaasher/Agent-Loop-Skills.

See all 25 skills โ†’
๐Ÿงฌ
1h ago

Alpha Evolve

Alpha Evolve evolves a model or program with parallel SEARCH/REPLACE mutations, cascade-evaluated runs, and a MAP-Elites archive across islands. Use it for bounded, diversity-preserving search over code or ML experiments rather than a single refine loop.
AI Engineering
16619
๐Ÿ”Ž
1h ago

Anomaly Investigation

Anomaly Investigation diagnoses a known data anomaly by forming candidate causes, testing them against the data, and eliminating those the evidence refutes. Use it when you already have a spike, drop, or outlier and need the confirmed root cause and supporting evidence.
Data
16619
๐Ÿ›ก๏ธ
1h ago

Blue Team

Blue Team patches a target against a concrete set of failing cases, one root-cause class at a time, while checking that previously passing cases still pass. Use it for red-team failure catalogues or CI test failures when you want the fix loop to stop only when regressions are closed.
Quality
16619
๐Ÿงช
1h ago

Claim Verify Loop

Claim Verify Loop checks each discrete claim in a results draft against the underlying dataset, then stress-tests it for outliers, confounds, and subgroup effects. Use it before publishing when a data-backed draft needs adversarial verification and revision.
AI Engineering
16619
๐Ÿ“Š
1h ago

Data Analysis Loop

Data Analysis Loop performs iterative exploratory analysis on a dataset, testing one hypothesis at a time and only keeping findings that reproduce with a meaningful effect size. Use it for open-ended discovery when every claim needs a computed number behind it.
Data
16619
โš”๏ธ
1h ago

Dueling Autoresearch

Dueling Autoresearch runs two different approaches against the same metric in parallel and keeps a shared scoreboard. Use it when you want an analysis-first head-to-head between lanes such as classical versus learned methods.
AI Engineering
16619