๐Ÿš€
DevOpsTeX

Distributed LLM Pretraining with TorchTitan

by Orchestra-Research

Distributed LLM Pretraining with TorchTitan is a DevOps skill for Claude Code, published by Orchestra-Research in AI-Research-SKILLs.

11.6K stars838 forkson Orchestra-Research/AI-Research-SKILLsAdded 2026/07/19+1% in starsRepository updated 2026/06/16
aiai-researchclaudeclaude-codeclaude-skillscodexgeminigpt-5grpohuggingfacemachine-leanringmegatronskillsvllm
Install in seconds
Install Distributed LLM Pretraining with TorchTitan
Copy Distributed LLM Pretraining with TorchTitan into your Claude Code skills folder. Run the command in your terminal, or review the source on GitHub before installing.
terminal
npx degit https://github.com/Orchestra-Research/AI-Research-SKILLs/tree/main/01-model-architecture/torchtitan ~/.claude/skills/torchtitan

Requires Node.js. Downloads this skill only โ€” not the rest of the repository โ€” into your Claude Code skills folder.

Without Node.js

git clone https://github.com/Orchestra-Research/AI-Research-SKILLs.git

Clones the whole repository, then copy the skillโ€™s own directory into your skills folder yourself.

In this catalog

Source file
01-model-architecture/torchtitan/SKILL.md in Orchestra-Research/AI-Research-SKILLs
Installs to
~/.claude/skills/torchtitan
Collection
One of 50 skills cataloged from this repository
Category
DevOps โ€” 1075 skills

What Distributed LLM Pretraining with TorchTitan does

Provides PyTorch-native distributed LLM pretraining using torchtitan with 4D parallelism (FSDP2, TP, PP, CP). Use when pretraining Llama 3.1, DeepSeek V3, or custom models at scale from 8 to 512+ GPUs with Float8, torch.compile, and distributed checkpointing.

Distributed LLM Pretraining with TorchTitan is cataloged under DevOps on DirSkills. Distributed LLM Pretraining with TorchTitan comes from a repository tagged ai, ai-research, claude, claude-code and claude-skills.

Documentation

README

TorchTitan - PyTorch Native Distributed LLM Pretraining

Quick start

TorchTitan is PyTorch's official platform for large-scale LLM pretraining with composable 4D parallelism (FSDP2, TP, PP, CP), achieving 65%+ speedups over baselines on H100 GPUs.

Installation:

# From PyPI (stable)
pip install torchtitan

# From source (latest features, requires PyTorch nightly)
git clone https://github.com/pytorch/torchtitan
cd torchtitan
pip install -r requirements.txt

Download tokenizer:

# Get HF token from https://huggingface.co/settings/tokens
python scripts/download_hf_assets.py --repo_id meta-llama/Llama-3.1-8B --assets tokenizer --hf_token=...

This is the opening of the README. Read the full README on GitHub.

Frequently asked about Distributed LLM Pretraining with TorchTitan

  • What else does Orchestra-Research publish alongside Distributed LLM Pretraining with TorchTitan?

    Distributed LLM Pretraining with TorchTitan is one of 50 skills that DirSkills catalogs from Orchestra-Research/AI-Research-SKILLs, the repository it ships in. Its siblings there include AWQ Quantization, Autoresearch and Axolotl. Each one is a separate skill with its own page in this directory, installs the same way Distributed LLM Pretraining with TorchTitan does, and is maintained by Orchestra-Research in that same repository. The rest of the collection is listed on the Orchestra-Research/AI-Research-SKILLs page.

  • How does Distributed LLM Pretraining with TorchTitan compare to other DevOps skills?

    Distributed LLM Pretraining with TorchTitan ranks #94 by stars among the 1075 DevOps skills in this catalog. The most-starred ones next to it are Backend Patterns, API Connector Builder and Migration. DirSkills ranks by the star count of the repository each skill ships in, so that order reflects how popular those repositories are rather than any review of Distributed LLM Pretraining with TorchTitan against them. Open each page to compare what they document and how they install.

More from Orchestra-Research/AI-Research-SKILLs

Distributed LLM Pretraining with TorchTitan is one of 50 skills cataloged on DirSkills from Orchestra-Research/AI-Research-SKILLs.

See all 50 skills โ†’
โš™๏ธ
2026/08/12

AWQ Quantization

AWQ Quantization compresses large language models to 4-bit with activation-aware weight selection. Use it to deploy 7B-70B models on limited GPU memory while keeping accuracy loss low.
AI Engineering
11.6K844
๐Ÿ”ฌ
2026/07/19

Autoresearch

Orchestrates end-to-end autonomous AI research projects using a two-loop architecture for rapid experiment iteration and synthesis. Routes to domain-specific skills, supports continuous operation, and produces research presentations and papers.
AI Engineering
11.6K838
๐Ÿง 
2026/07/19

Axolotl

Comprehensive guidance for fine-tuning LLMs using Axolotl, including YAML configuration, 100+ model support, LoRA/QLoRA, DPO/KTO/ORPO/GRPO, and multimodal training.
AI Engineering
11.6K838
๐Ÿงช
2026/08/12

BigCode Evaluation Harness

BigCode Evaluation Harness evaluates code generation models on HumanEval, MBPP, MultiPL-E, and other benchmarks with pass@k metrics. Use it to benchmark coding ability, compare models, and measure multi-language code generation quality.
Quality
11.6K844
๐Ÿงฎ
2026/08/12

Bitsandbytes Model Quantization

Bitsandbytes Model Quantization loads LLMs in 8-bit or 4-bit to cut GPU memory use and fit larger models. Use it for Hugging Face Transformers inference, QLoRA fine-tuning, or 8-bit optimizers when VRAM is limited.
AI Engineering
11.6K844
๐Ÿ›ก๏ธ
2026/08/12

Constitutional AI

Constitutional AI trains models with self-critique, revision, and AI feedback to reduce harmful outputs without human labels. Use it when you need safety alignment or a clear set of principles for model behavior.
AI Engineering
11.6K844