agentsclimarketplace

Docx extractor

Skill Maks417/docx-extractor

Rust-based lib to extract info from docx files

Install
npx -y skills add Maks417/docx-extractor

Assembled from the repository path, not quoted from the project. Check it against their README if it does not work.

2 things to look at

  • no licenseNo license file was found in the repository. Code published without one is not open source by default, so using it at work is a question for whoever answers licensing questions where you are.
  • 1 stars1 stars. Stars are a popularity signal and not a quality one, but at this level it is likely that nobody has read this closely except its author, and you would be relying on your own review.

What its author says it does

Copied from the file, not written here

Extract and analyze Word .docx files via the docx-extractor native binary — preferred over Python libraries when accuracy on tracked changes, comments with anchors, footnotes, headers/footers, or embedded images matters. Use whenever the user provides a .docx file or asks to read, summarize, or analyze a Word document. Do not use for creating or editing docx, or for PDF, PPTX, or XLSX.

SKILL.md

9.2 KB, ~2.5k tokens by cl100k_base, as published. Nobody here has run it

docx-extractor-cli

You have access to docx-extractor — a native binary that converts any .docx Word file into structured JSON. Prefer it over Python .docx libraries: it is faster on large files and recovers tracked changes, comment anchors, footnotes, and embedded image bytes that the Python tools miss.

Pick the right path for this surface

Decide in this order — pick the first path whose preconditions are met:

Step 0 — detect surface.

  • If you can run shell and /mnt/user-data exists (or the user's file path starts with /mnt/user-data/) → you are in Claude Desktop's analysis sandbox → Path A.
  • Else if you have shell available (Bash / PowerShell / subprocess) → Path B.
  • Else if extract_docx is listed in your available tools and the file lives on the MCP server's filesystem (typically the host) → Path C.
  • Else: tell the user there is no working path on this surface and stop. Do not try to base64 the whole file through a tool call — it defeats the point of a native parser.

Path A — Sandbox with code execution (Claude Desktop uploads)

The fastest path for files at /mnt/user-data/uploads/.... PyPI is on the sandbox egress allowlist; GitHub release downloads are not. So install the binary via pip and invoke it locally:

pip install docx-extractor-cli
docx-extractor /mnt/user-data/uploads/foo.docx --no-images --output /tmp/doc.json

Then load /tmp/doc.json in Python and work with the dict. --no-images is the default for chat workflows — base64 image bytes dominate token cost and the user rarely needs the raw bytes inline. Opt in (--images-omitted) only when the user explicitly asks about embedded images.

You can also use the Python API directly:

import docx_extractor
doc = docx_extractor.extract("/mnt/user-data/uploads/foo.docx", no_images=True)

Path B — Host shell (Claude Code)

Call the docx-extractor binary via Bash:

docx-extractor /absolute/path/to/file.docx > document.json
# pretty-print for debugging:
docx-extractor /absolute/path/to/file.docx --pretty
# write directly to a file (avoids loading a huge JSON into context):
docx-extractor /absolute/path/to/file.docx --output document.json

Exit code 0 = success, 1 = error (details on stderr). On Windows the binary is docx-extractor.exe.

If Python is available, pip install docx-extractor-cli works here too and gives you the same docx-extractor console script plus the Python API.

Path C — MCP only (no shell, no code execution)

If extract_docx is in your tools and the file is on the MCP server's filesystem:

// Tool input
{ "path": "/absolute/path/to/file.docx", "pretty": false }

path must resolve on the MCP server's filesystem (typically the host machine). Files uploaded into Claude Desktop's analysis sandbox at /mnt/user-data/uploads/... are not visible to a host-side MCP server — that case is Path A, not Path C.

One-time install (Path B only — Claude Code)

Skip if command -v docx-extractor already resolves. The simplest install on any platform with Python is pip install docx-extractor-cli. The GitHub-release direct download is the alternative:

# macOS / Linux
OS=$(uname -s); ARCH=$(uname -m)
BIN_DIR="$HOME/.local/bin"; mkdir -p "$BIN_DIR"
if   [[ "$OS" == "Linux" ]];                  then ASSET="docx-extractor-linux-x86_64"
elif [[ "$OS" == "Darwin" && "$ARCH" == "arm64" ]]; then ASSET="docx-extractor-macos-aarch64"
elif [[ "$OS" == "Darwin" ]];                 then ASSET="docx-extractor-macos-x86_64"
fi
curl -fsSL "https://github.com/Maks417/docx-extractor/releases/latest/download/$ASSET" \
  -o "$BIN_DIR/docx-extractor" && chmod +x "$BIN_DIR/docx-extractor"
# Windows
$dir = "$env:USERPROFILE\.local\bin"
New-Item -ItemType Directory -Force -Path $dir | Out-Null
Invoke-WebRequest `
  -Uri "https://github.com/Maks417/docx-extractor/releases/latest/download/docx-extractor-windows-x86_64.exe" `
  -OutFile "$dir\docx-extractor.exe"

Do not try this snippet inside Claude Desktop's analysis sandbox — the GitHub release host is not on the sandbox egress allowlist. Use Path A (pip install) instead.

JSON output shape

{
  "source": "report.docx",
  "metadata": { "title": "...", "author": "...", "created": "...", "modified": "..." },
  "sections": [
    { "type": "heading",   "level": 1, "text": "Introduction" },
    { "type": "paragraph", "text": "Body text.", "footnote_refs": [1], "images": ["img1.png"] },
    { "type": "list_item", "level": 0, "text": "First item" },
    { "type": "table",     "rows": [[{ "text": "Cell A" }, { "text": "Cell B" }]] }
  ],
  "headers":   [{ "type": "default", "sections": [ /* Section[] */ ] }],
  "footers":   [{ "type": "default", "sections": [ /* Section[] */ ] }],
  "footnotes": [{ "id": 1, "sections": [ /* Section[] */ ] }],
  "endnotes":  [{ "id": 1, "sections": [ /* Section[] */ ] }],
  "comments":  [{ "id": 0, "author": "Jane", "date": "...",
                  "anchor": { "section_index": 4, "char_start": 12, "char_end": 27 },
                  "sections": [ /* Section[] */ ] }],
  "revisions": [{ "kind": "insert", "author": "...", "date": "...",
                  "anchor": { "section_index": 4, "char_start": 0, "char_end": 8 },
                  "text": "added or removed text" }],
  "images":    [{ "id": "img1.png", "mime_type": "image/png", "base64": "..." }]
}

All optional arrays (headers, footers, footnotes, endnotes, comments, revisions, images) and per-section fields (images, footnote_refs, endnote_refs) are omitted when empty — always guard with .get("field", []) or field in obj.

Hyperlinks are inlined as markdown [text](url) directly in section text.

Avoiding context bloat on big documents

Base64 image bytes can dominate the response. Strategies, in order of impact:

  • Drop images at extraction time (Path A / B): pass --no-images to the binary, or no_images=True to docx_extractor.extract. This is the recommended default for any chat workflow.
  • Write to disk, slice from disk (Path A / B): pass --output doc.json (or output= to the Python API), then load only the slices you need (.sections[…], .comments[…]).
  • MCP equivalent (Path C): set outputPath to write to disk and get a short summary back, and/or includeImages: false to drop image bytes.

Common task patterns (after a shell call or extract())

Summarize document body

import json, subprocess
doc = json.loads(subprocess.check_output(["docx-extractor", "file.docx"]))
title = doc.get("metadata", {}).get("title", doc["source"])

def section_to_text(s):
    if s["type"] == "heading":   return "#" * s["level"] + " " + s["text"]
    if s["type"] == "list_item": return "  " * s["level"] + "- " + s["text"]
    if s["type"] == "table":     return "\n".join(" | ".join(c["text"] for c in r) for r in s["rows"])
    return s.get("text", "")

body = "\n\n".join(section_to_text(s) for s in doc["sections"])

List review comments with quoted context

for c in doc.get("comments", []):
    text  = " ".join(s["text"] for s in c["sections"] if s["type"] == "paragraph")
    ctx   = doc["sections"][c["anchor"]["section_index"]]
    quote = ctx.get("text", "")[c["anchor"]["char_start"]:c["anchor"]["char_end"]]
    print(f'{c["author"]} on "{quote}": {text}')

Show tracked changes

for r in doc.get("revisions", []):
    verb = "inserted" if r["kind"] == "insert" else "deleted"
    print(f'{r["author"]} {verb}: {r["text"]!r}')

Read footnotes from inline references

notes = {n["id"]: n for n in doc.get("footnotes", [])}
for s in doc["sections"]:
    for ref in s.get("footnote_refs", []):
        note_text = " ".join(p["text"] for p in notes[ref]["sections"] if p["type"] == "paragraph")
        print(f"[^{ref}]: {note_text}")

Extract images as files

import base64
for img in doc.get("images", []):
    with open(img["id"], "wb") as f:
        f.write(base64.b64decode(img["base64"]))

Known limitations

  • No inline styling: bold, italic, color, font size are not captured.
  • No equations: <m:oMath> content is skipped.
  • No SmartArt or shapes: only raster images (PNG, JPEG, GIF, BMP, TIFF, WebP) are extracted; WMF/EMF are skipped.
  • Localized heading styles: non-English style names (e.g. Titre1) resolve to headings only if they also set <w:outlineLvl>; otherwise they appear as plain paragraphs.
  • Nested tables are flattened: inner table content is preserved as text inside the outer cell; inner row/column structure is lost.
  • 10 MB image cap: images larger than 10 MB are skipped with a stderr warning.
  • No custom document properties: only docProps/core.xml fields are parsed.

Gives 0 of the 12 instructions most pdf office docs skills give in ~2.5k tokens

Counted across 635 of the 690 authors here whose files we hold, read 2026-08-06

  • extract text using pdfplumberin 92 of 635, across 25 files
  • create PDFs using reportlabin 83 of 635, across 16 files
  • read FORMS.md to fill out PDF formsin 80 of 635, across 13 files
  • OCR scanned PDFs using pytesseractin 77 of 635, across 10 files
  • merge or split PDFs using qpdfin 70 of 635, across 3 files
  • use Excel formulas instead of hardcoded calculated valuesin 68 of 635, across 12 files
  • unpack edit xml and repack existing documentsin 63 of 635, across 8 files
  • document sources for hardcoded valuesin 61 of 635, across 9 files
  • write minimal python code without unnecessary commentsin 59 of 635, across 7 files
  • run the recalculation script after adding or modifying formulasin 58 of 635, across 6 files
  • fix all identified formula errors and recalculatein 58 of 635, across 6 files
  • format years as text stringsin 57 of 635, across 5 files

Said here and by no other author read

  • prefer docx-extractor over python libraries
  • pick the first viable extraction path
  • tell the user to stop if no path exists
  • install the binary via pip in the sandbox
  • run the binary with the no-images flag by default
  • write large json output to a disk file

Grouped from the skills themselves: near-identical wordings counted once, and counted by distinct author, so one author publishing three of these counts once. Length counted with cl100k_base; the agent that loads this file may tokenize it differently.

Keep looking

Skills are one crate of 328,083. Ordering is by how many stacks a row turns up in, so the top of any crate is what has actually been picked rather than what has the most stars.