Paper analyzer
Transform academic papers into in-depth technical articles with multiple writing style options. Use the MinerU Cloud API for high-precision PDF parsing, automatically extracting images, tables, and formulas. Optional formula explanations and GitHub code analysis, generating Markdown and HTML formats.From its SKILL.md
npx -y skills add proyecto26/sherlock-ai-plugin --skill paper-analyzerAssembled from the repository path, not quoted from the project. Check it against their README if it does not work.
SKILL.md
2.9 KB, 592 tokens by cl100k_base, as published. Nobody here has run it
Academic Paper Analyzer – In-Depth Analysis of Academic Papers
Core Capabilities
- MinerU Cloud API for high-precision PDF parsing
- Automatic extraction of images, tables, and LaTeX formulas
- Multiple writing styles: storytelling / academic / concise
- Optional formula explanations: insert formula images with detailed symbol explanations
- Optional code analysis: combine explanations with GitHub open-source code
- Output Markdown + HTML (base64-embedded images)
Prerequisites
MinerU API Token
- Visit https://mineru.net and register an account
- Obtain an API Token
- Set an environment variable (recommended):
export MINERU_TOKEN="your_token_here"
Dependency Installation
pip install requests markdown
Workflow
Step 1: PDF Parsing (Using MinerU API)
python scripts/mineru_api.py <pdf_path> <output_dir>
Or pass the token directly:
python scripts/mineru_api.py paper.pdf ./output YOUR_TOKEN
Output:
output_dir/*.md– Markdown files (including formulas and tables)output_dir/images/– High-quality extracted images
Step 2: Extract Paper Metadata
python scripts/extract_paper_info.py <output_dir>/*.md paper_info.json
Step 3: Style Selection (Ask the User)
Before generating the article, you must ask the user to choose the following options:
1. Writing Style (Required)
| Style | Characteristics | Use Cases |
|---|---|---|
| storytelling | Starts from intuition, uses metaphors and examples, narrative-driven | Blogs, tech columns, popular science |
| academic | Professional terminology, rigorous expression, preserves original concepts | Academic reports, surveys, research group sharing |
| concise | Straight to the point, tables and lists, high information density | Quick reads, paper overviews, technical research |
2. Formula Option (Optional)
| Option | Description |
|---|---|
| with-formulas | Insert formula images and explain symbol meanings in detail |
| no-formulas (default) | Pure text description, no formula images |
3. Code Option (Optional, only if the paper has GitHub)
| Option | Description |
|---|---|
| with-code | Clone the repository, include key source code, and explain it alongside the paper |
| no-code (default) | No code analysis |
Step 4: Intelligent Article Generation
(...)
API Limits
- Maximum file size: 200MB
- Maximum pages per file: 600
- Supports PDF, DOC, PPT, images, and more
What ships with it: 14 files
31.3 KB alongside SKILL.md, 4 of them executable
scripts/
- convert_pdf.pyruns5.2 KB
- extract_paper_info.pyruns3.7 KB
- generate_html.pyruns4.0 KB
- mineru_api.pyruns9.8 KB
styles/
- academic.md1.4 KB
- concise.md1.3 KB
- no-code.md377 B
- no-formulas.md742 B
- storytelling.md1.4 KB
- with-code.md945 B
- with-formulas.md633 B
- cover-prompt.md1.2 KB
- cover-prompt-simple.txt766 B
- requirements.txt32 B
Gives 0 of the 12 instructions most docs writing skills give in 592 tokens
Counted across 1,951 of the 3,904 authors here whose files we hold, read 2026-09-06
- Use third-person for skill descriptionsin 54 of 1951, across 35 files
- Start descriptions with Use whenin 43 of 1951, across 29 files
- Run baseline scenarios before writing any skillin 40 of 1951, across 26 files
- Use active voicein 40 of 1951, across 36 files
- Map file responsibilities before defining tasksin 36 of 1951, across 29 files
- Use checkbox syntax for tracking stepsin 35 of 1951, across 27 files
- Ask one question at a timein 35 of 1951
- Offer execution options after saving the planin 33 of 1951, across 24 files
- Include complete code in every stepin 33 of 1951, across 27 files
- Design units with clear boundaries and interfacesin 31 of 1951, across 23 files
- Announce the skill usage at the startin 30 of 1951
- Verify agent compliance after adding the skillin 29 of 1951, across 17 files
Grouped from the skills themselves: near-identical wordings counted once, and counted by distinct author, so one author publishing three of these counts once. Length counted with cl100k_base; the agent that loads this file may tokenize it differently.