agentsclimarketplace

Markitdown

Skill matematicsolutions/awesome-matematic-skills-pl/dokumenty/skills/markitdown

Konwersja dowolnego dokumentu (PDF, Word, Excel, PowerPoint, HTML, EPUB, audio, obrazy, YouTube) na Markdown dla LLM. Użyj gdy użytkownik mówi "konwertuj PDF", "przerób Word na markdown", "zamień PPT na MD", "markdown z Excela", "wyciągnij tekst z PDF", albo daje plik Office/PDF do analizy. Microsoft MarkItDown (pip) + MCP server.From its SKILL.md

Install
npx -y skills add matematicsolutions/awesome-matematic-skills-pl --skill markitdown

Assembled from the repository path, not quoted from the project. Check it against their README if it does not work.

One thing to look at

  • 5 stars5 stars. Stars are a popularity signal and not a quality one, but at this level it is likely that nobody has read this closely except its author, and you would be relying on your own review.

What its file declares

Copied from the file, not written here

The file declares its own license as MIT. That is the author’s claim about this one file, and it is not the same thing as the license GitHub reports for the repository, which is listed with the other numbers below.

SKILL.md

2.6 KB, 731 tokens by cl100k_base, as published. Nobody here has run it

MarkItDown - konwerter dokumentów do Markdown (PL)

Lightweight utility Microsoftu - zachowuje strukturę (nagłówki, listy, tabele, linki), nie wygląd. Pod LLM, nie pod human.

Instalacja (zrobione 2026-04-21)

python -m pip install --user markitdown markitdown-mcp

Wymaga Python 3.10+ (testowane na 3.14). CLI: python -m markitdown.

Wspierane formaty

  • PDF (preferuj dla krótkich, standardowych PDF; dla złożonych/tabel - OpenDataLoader PDF)
  • Office: Word (.docx), Excel (.xlsx), PowerPoint (.pptx)
  • HTML, EPUB, CSV, JSON, XML
  • Obrazy (EXIF + OCR jeśli zainstalowane [all])
  • Audio (EXIF + transkrypcja jeśli włączone)
  • ZIP (iteruje zawartość)
  • YouTube URL (napisy)

Użycie

CLI (single file)

python -m markitdown input.pdf > output.md
python -m markitdown input.pptx -o output.md

Batch (Obsidian Vault)

for f in "/c/Users/hp/Documents/Obsidian Vault/Konwerter"/*.pdf; do
  python -m markitdown "$f" > "${f%.pdf}.md"
done

Python API

from markitdown import MarkItDown
md = MarkItDown()
result = md.convert("plik.docx")
print(result.text_content)

MCP server

Opcjonalnie - jeśli chcesz udostępnić Claude Code jako MCP tool:

markitdown-mcp

Kiedy użyć MarkItDown vs OpenDataLoader PDF

SytuacjaNarzędzie
Word/Excel/PPTMarkItDown
Prosty PDF, tekst liniowyMarkItDown (szybsze)
Złożony PDF z tabelami, reading order, papers naukoweOpenDataLoader PDF (jakość)
Audio/transkrypcjaWhisper (mamy whisper-asr-pipeline)
Web pageDefuddle (mamy defuddle od Kepano)

Integracja z vault

Output docelowo → Konwerter/ (folder w vault dla source-pdf → MD). Frontmatter: type: source-pdf, tags: [pdf, zrodlo]. Zobacz vault-rules.json → clippings.classification_rules.pdf_source.

Gotcha

  • markitdown[all] na Python 3.14/3.15 failuje (onnxruntime, youtube-transcript-api konflikt) - użyj zainstalowanego markitdown bez extras
  • Dla OCR obrazów / pełnego YouTube - Docker albo starszy Python (3.11/3.12)
  • Nie używaj dla high-fidelity conversion dla human reading - tylko pod LLM

What ships with it

Read from the repository

Just SKILL.md. No reference files, no scripts.

Keep looking

Skills are one crate of 325,949. Ordering is by how many stacks a row turns up in, so the top of any crate is what has actually been picked rather than what has the most stars.