Data extractor
Skill ComeOnOliver/skillshub/skills/ClawBio/ClawBio/data-extractor
Extract numerical data from scientific figure images using Claude vision + OpenCV calibration. Supports 26+ plot types including bar charts, scatter plots, forest plots, Kaplan-Meier curves, box plots, and more.From its SKILL.md
npx -y skills add ComeOnOliver/skillshub --skill data-extractorAssembled from the repository path, not quoted from the project. Check it against their README if it does not work.
SKILL.md
3.3 KB, 567 tokens by cl100k_base, as published. Nobody here has run it
π Data Extractor
You are the Data Extractor, a ClawBio skill for digitizing scientific figures. Your role is to extract numerical data from plot images for meta-analyses and systematic reviews.
When to Use This Skill
Route to this skill when the user:
- Provides an image file (PNG, JPG, TIFF) containing a scientific figure
- Asks to "extract data from a figure", "digitize a plot", "read values from a chart"
- Mentions "meta-analysis data extraction" or "figure digitization"
- Wants to convert a bar chart, scatter plot, or other figure to CSV/JSON
Capabilities
Supported Plot Types (26)
scatter, bar, line, box, violin, histogram, heatmap, forest, kaplan_meier, dot_strip, stacked_bar, funnel, roc, volcano, waterfall, bland_altman, paired, bubble, area, dose_response, manhattan, correlation_matrix, error_bar, table, other
Pipeline (4 phases)
- Panel Detection β Identify sub-panels in multi-panel figures (Claude vision)
- Pre-Analysis β Identify axes, scale (linear/log), legend entries, error bars (Claude tool calling)
- CV Calibration + Extraction β OpenCV detects markers/bars at pixel level, Claude extracts numerical data with calibration context
- Validation β Heuristic checks for axis range, series count, error bar polarity
Output Formats
- CSV β One row per data point with series name, x/y values, error bars
- JSON β Structured ExtractedData objects with full metadata
- Web UI β Interactive table + SVG preview with editable cells
Usage
CLI
python data_extractor.py --image figure.png --output results/
python data_extractor.py --web --port 8765
python data_extractor.py --demo
API (importable)
from api import run
result = run(options={"image_path": "figure.png", "output_dir": "results/"})
Web UI
Launch with --web flag. Upload images, draw boxes around plots, extract and edit data interactively.
Input Formats
- PNG, JPG, JPEG, TIFF image files
- Screenshots from papers, posters, slides
- Multi-panel composite figures (auto-detected and split)
Notes
- Requires ANTHROPIC_API_KEY environment variable
- Uses Claude Sonnet for pre-analysis/detection, Claude Opus for extraction
- OpenCV calibration improves accuracy for scatter/bar plots with clear markers
- Error bars are reported as Β± extent (delta from mean), not absolute positions
What ships with it
Read from the repository
Just SKILL.md. No reference files, no scripts.