DataPython

Spark

by nimadorostkar

Spark is a Data skill for Claude Code, published by nimadorostkar in Claude-Skills-collection.

25 stars3 forkson nimadorostkar/Claude-Skills-collectionAdded 2026/08/12Repository updated 2026/07/26
aiclaudeclaude-skillsskills
Install in seconds
Install Spark
Copy Spark into your Claude Code skills folder. Run the command in your terminal, or review the source on GitHub before installing.
terminal
npx degit https://github.com/nimadorostkar/Claude-Skills-collection/tree/main/skills/data/spark ~/.claude/skills/spark

Requires Node.js. Downloads this skill only — not the rest of the repository — into your Claude Code skills folder.

Without Node.js

git clone https://github.com/nimadorostkar/Claude-Skills-collection.git

Clones the whole repository, then copy the skill’s own directory into your skills folder yourself.

In this catalog

Source file
skills/data/spark/SKILL.md in nimadorostkar/Claude-Skills-collection
Installs to
~/.claude/skills/spark
Collection
One of 50 skills cataloged from this repository
Category
Data668 skills

What Spark does

Spark helps you build and review distributed Spark jobs with a focus on partitioning, shuffles, skew, joins, caching, and Spark UI diagnosis. Use it when a pipeline is slow, spilling, or has a straggling task.

Spark is cataloged under Data on DirSkills. Spark comes from a repository tagged ai, claude, claude-skills and skills.

Documentation

README

Spark

Purpose

Write Spark jobs whose cost is understood. Almost all Spark performance problems are one of three things: too much shuffle, skewed partitions, or reading far more data than the query needs.

When to Use

  • Building or reviewing a Spark pipeline.
  • A job that is slow, failing with out-of-memory errors, or has one straggling task.
  • Tuning partitioning and join strategy.
  • Reading the Spark UI to diagnose a stage.

Capabilities

  • Partitioning strategy and repartitioning.
  • Shuffle minimization and broadcast joins.
  • Skew detection and mitigation.
  • Caching and persistence levels.
  • File-format and predicate-pushdown optimization.
  • Spark UI interpretation.

Inputs

This is the opening of the README. Read the full README on GitHub.

Frequently asked about Spark

  • What else does nimadorostkar publish alongside Spark?

    Spark is one of 50 skills that DirSkills catalogs from nimadorostkar/Claude-Skills-collection, the repository it ships in. Its siblings there include API Design, Agent Design and Agent Instructions. Each one is a separate skill with its own page in this directory, installs the same way Spark does, and is maintained by nimadorostkar in that same repository. The rest of the collection is listed on the nimadorostkar/Claude-Skills-collection page.

  • How does Spark compare to other Data skills?

    Spark ranks #649 by stars among the 668 Data skills in this catalog. The most-starred ones next to it are Benchmark Methodology, Jupyter Notebook and Solana. DirSkills ranks by the star count of the repository each skill ships in, so that order reflects how popular those repositories are rather than any review of Spark against them. Open each page to compare what they document and how they install.

More from nimadorostkar/Claude-Skills-collection

Spark is one of 50 skills cataloged on DirSkills from nimadorostkar/Claude-Skills-collection.

See all 50 skills