📊
DataJavaScript

Document Analysis

by OpenSenseNova

Document Analysis is a Data skill for Claude Code, published by OpenSenseNova in SenseNova-Skills.

4.9K stars364 forkson OpenSenseNova/SenseNova-SkillsAdded 2026/08/16Repository updated 2026/08/13
agentagent-skillsai-agentsai-assistantdata-analysisdocument-processingoffice-automationpresentation-slides
Install in seconds
Install Document Analysis
Copy Document Analysis into your Claude Code skills folder. Run the command in your terminal, or review the source on GitHub before installing.
terminal
npx degit https://github.com/OpenSenseNova/SenseNova-Skills/tree/main/skills/sn-da-non-spreadsheet-analysis ~/.claude/skills/sn-da-non-spreadsheet-analysis

Requires Node.js. Downloads this skill only — not the rest of the repository — into your Claude Code skills folder.

Without Node.js

git clone https://github.com/OpenSenseNova/SenseNova-Skills.git

Clones the whole repository, then copy the skill’s own directory into your skills folder yourself.

In this catalog

Source file
skills/sn-da-non-spreadsheet-analysis/SKILL.md in OpenSenseNova/SenseNova-Skills
Installs to
~/.claude/skills/sn-da-non-spreadsheet-analysis
Collection
One of 25 skills cataloged from this repository
Category
Data668 skills

What Document Analysis does

Document Analysis parses Word, PDF, and PPT files to extract text, tables, charts, and formatting, then performs cross-document aggregation and numerical checks. Use it when users ask to analyze, extract, or compare content from .docx, .pdf, or .pptx files.

Document Analysis is cataloged under Data on DirSkills. Document Analysis comes from a repository tagged agent, agent-skills, ai-agents, ai-assistant and data-analysis.

Documentation

README

Document Analysis Skill — Word / PDF / PPT

End-to-end workflow for Word, PDF, and PPT document parsing. Each format has specific parsing pitfalls — follow the format-specific sub-skill exactly.


Workflow

Step 0 — Identify file type and input scope

import os

input_path = "/mnt/data/..."  # from user

# Detect single file vs directory (multi-file scenario)
if os.path.isdir(input_path):
    all_files = [
        os.path.join(input_path, f)
        for f in os.listdir(input_path)
        if f.lower().endswith(('.docx', '.doc', '.pdf', '.pptx', '.ppt'))
    ]
    print(f"Found {len(all_files)} documents: {all_files}")
else:
    all_files = [input_path]

# Route by extension
ext = os.path.splitext(all_files[0])[-1].lower()
print(f"File type: {ext}")

This is the opening of the README. Read the full README on GitHub.

Frequently asked about Document Analysis

  • What else does OpenSenseNova publish alongside Document Analysis?

    Document Analysis is one of 25 skills that DirSkills catalogs from OpenSenseNova/SenseNova-Skills, the repository it ships in. Its siblings there include Academic Research, China Market Search and Chinese Social Search. Each one is a separate skill with its own page in this directory, installs the same way Document Analysis does, and is maintained by OpenSenseNova in that same repository. The rest of the collection is listed on the OpenSenseNova/SenseNova-Skills page.

  • How does Document Analysis compare to other Data skills?

    Document Analysis ranks #181 by stars among the 668 Data skills in this catalog. The most-starred ones next to it are Benchmark Methodology, Jupyter Notebook and Solana. DirSkills ranks by the star count of the repository each skill ships in, so that order reflects how popular those repositories are rather than any review of Document Analysis against them. Open each page to compare what they document and how they install.

More from OpenSenseNova/SenseNova-Skills

Document Analysis is one of 25 skills cataloged on DirSkills from OpenSenseNova/SenseNova-Skills.

See all 25 skills
🎓
2w ago

Academic Research

Academic Research searches academic papers, reads full texts, and traces references and citations across sources like arXiv, Semantic Scholar, PubMed, and Wikipedia. Use it for literature review, related work discovery, and deep paper reading.
Automation
4.9K364
🔍
2w ago

China Market Search

China Market Search locates Chinese market, industry, macro, trade, procurement, and regulatory data from official free sources without registration or API keys. It uses a Python script for A-share disclosures and browser-based access to government databases.
Data
4.9K364
🔍
2w ago

Chinese Social Search

Chinese Social Search searches content on Chinese social platforms including Bilibili videos, Zhihu Q&A, and Douyin videos. Use it when research requires Chinese social media data; Xiaohongshu and Weibo fall back to browser-use on public web pages.
Automation
4.9K364
🎨
2w ago

Creative PPT

Creative PPT generates a presentation one full-page 16:9 PNG per slide, using LLM/VLM for text and T2I image generation with web image search fallback. Use it via sn-ppt-entry to turn a topic into creative, image-forward slides.
AI Engineering
4.9K364
🔍
2w ago

Deep Research

Deep Research orchestrates multiple specialist agents to conduct systematic research, fact-checking, competitive analysis, and report generation when a request requires cross-source verification or a detailed deliverable.
Data
4.9K364
🔍
2w ago

Developer Search

Developer Search queries GitHub, Stack Overflow, Hacker News, and HuggingFace to find code examples, issues, technical Q&A, discussions, and models. Use it when you need to locate code, answers, or model info across developer platforms.
AI Engineering
4.9K364