Html to markdown
Lightweight HTML-to-Markdown parser for AI agent workflows — Python CLI and API around a fast Go engine, available on PyPI.
npx -y skills add appautomaton/markmaton --skill html-to-markdownAssembled from the repository path, not quoted from the project. Check it against their README if it does not work.
One thing to look at
- 4 stars4 stars. Stars are a popularity signal and not a quality one, but at this level it is likely that nobody has read this closely except its author, and you would be relying on your own review.
What its author says it does
Copied from the file, not written here
Convert a URL or HTML into clean Markdown with metadata using markmaton. Handles browser capture for JS-heavy pages and deterministic HTML-to-Markdown conversion in one skill.
SKILL.md
2.5 KB, as published. Nobody here has run it
HTML to Markdown
Composes with
- Called by —
RECIPES.mdparse-jdrecipe (step 1). Any capture-a-web-page task in the workspace prefers this over built-inWebFetch. - Wraps — nodriver (CDP-based headless browser capture for JS-heavy pages, with Playwright Chromium discovery) and markmaton (HTML→Markdown with main-content extraction, metadata, and link/image inventory). See
references/integration-patterns.mdfor browser-vs-fetch guidance. - Outputs — JSON envelope by default (markdown body + metadata + links + images + quality signals). Use
--output-format markdownwhen only the raw Markdown body is needed.
Converts a URL or HTML into clean Markdown plus metadata, links, images, and quality signals.
From a URL
Capture the page and convert in one pipeline:
uv run --script scripts/capture_html.py <url> \
| uv run --script scripts/markmaton_convert.py --from-capture --output-format json
The capture script outputs a JSON envelope by default. --from-capture reads it and extracts html, url, final_url, and content_type automatically — no context lost, URL typed once.
- Add
--wait-selector <css>or--wait-text <string>to the capture step for pages that need a readiness signal. - Prefer a simple fetch over browser capture for static articles, wikis, and server-rendered docs.
From HTML
uv run --script scripts/markmaton_convert.py --html-file page.html \
--url <url> --output-format json
Or from stdin:
echo "$html" | uv run --script scripts/markmaton_convert.py --url <url>
Pass --url when available — it improves link resolution and canonical metadata.
Key defaults
- Output:
json. Use--output-format markdownfor raw Markdown only. - Main-content extraction: on. Use
--full-contentto disable. - Capture: always headless. Timeout
10s, override with--timeout. - Browser discovery: user's Chrome → user's Chromium → Playwright's Chromium.
References
Read only when needed:
references/usage.md— full CLI reference for both scriptsreferences/integration-patterns.md— browser vs fetch guidance, contracts, parser defaults
Core documentation
For package internals, release process, and benchmarks: docs/