Bio stats ml reporting
npx -y skills add fmschulz/omics-skills --skill bio-stats-ml-reportingAssembled from the repository path, not quoted from the project. Check it against their README if it does not work.
One thing to look at
- 7 stars7 stars. Stars are a popularity signal and not a quality one, but at this level it is likely that nobody has read this closely except its author, and you would be relying on your own review.
What its author says it does
Copied from the file, not written here
Analyze biological results with statistics or machine learning and produce validated reports. Use when aggregating features, testing hypotheses, training models, or reporting performance.
SKILL.md
5.1 KB, as published. Nobody here has run it
Bio Stats ML Reporting
Aggregate results, train ML models, and produce reports with validated references.
Instructions
- Join outputs in DuckDB v1.1+ and build feature tables. Arrow / DuckLake integration is the recommended bridge into ML pipelines for large datasets.
- Train baseline models and evaluate with cross-validation.
- CPU baseline: scikit-learn v1.5+ for linear/tree/clustering baselines; XGBoost v2.1.4+ for gradient boosting.
- GPU node available (CUDA): set
device="cuda"on XGBoost (native since v2.0) by default. For sklearn-compatible estimators (random forest, k-means, PCA, UMAP), use RAPIDS cuML as a drop-in replacement and record the device in the run log.
- Generate reports and validate references.
- Validate the prediction table with
scripts/validate_predictions.py. Keep group identifiers and confounder labels in the table so the gate can detect sample/group leakage, class imbalance, calibration failure, batch-outcome imbalance, and performance that does not beat the majority-class null.
- Validate the prediction table with
- For exploratory omics projects, aggregate discovery evidence across the literature-derived analysis playbook, annotation, phylogenomics, viromics, and comparative-genomics outputs.
- Comparative-axes rollup — join the per-axis comparison artifacts produced by upstream skills into a single
comparative_axes_summary.tsv. The rollup must have one row per (query genome, axis) and include:genome-property frontier(size, gene count, etc. — link torelative_genome_metrics.tsvandgenome_size_frontier.tsv)marker-gene census(link tomarker_census.tsv)family copy-number expansions/contractions(link tofamily_copy_number_comparison.tsvandfamily_expansion_candidates.tsv)synteny / conserved neighborhoods(link toconserved_neighborhoods.tsv)non-coding RNA census(link toncRNA_census.tsv) Each row records observation, comparison baseline, literature reference, status (notable / conserved / artifact / negative), and a follow-up test.
- Produce an interesting-findings section that ranks candidate discoveries relative to the literature-derived baseline and separates:
- strong candidates with multiple evidence types
- plausible candidates needing validation
- likely artifacts or conserved lineage features
- explicit negative findings where nothing notable was detected
- Include the comparison baseline, literature context, confidence, and next discriminating analyses for each candidate.
Quick Reference
| Task | Action |
|---|---|
| Run workflow | Follow the steps in this skill and capture outputs. |
| Validate inputs | Confirm required inputs and reference data exist. |
| Review outputs | Inspect reports and QC gates before proceeding. |
| Tool docs | See docs/README.md. |
| Validate predictions | uv run --no-project python scripts/validate_predictions.py predictions.tsv --report validation.json --require-beats-null |
Input Requirements
Prerequisites:
- Tools declared in the project's pinned Pixi environment. See
docs/README.mdfor expected tools. - Results tables and metadata are available. Inputs:
- results/.parquet or results/.tsv
- metadata.tsv
Output
- results/bio-stats-ml-reporting/models/
- results/bio-stats-ml-reporting/metrics.tsv
- results/bio-stats-ml-reporting/comparative_axes_summary.tsv
- results/bio-stats-ml-reporting/discovery_summary.tsv
- results/bio-stats-ml-reporting/report.md
- results/bio-stats-ml-reporting/logs/
- results/bio-stats-ml-reporting/prediction_validation.json
Quality Gates
- Model performance sanity checks pass.
- Reference validation passes.
- On failure: retry with alternative parameters; if still failing, record in report and exit non-zero.
- Verify input tables are readable and schema-consistent.
- Discovery summary joins candidate genes/features to annotation evidence, comparison baseline, literature context, and confidence.
-
comparative_axes_summary.tsvcovers all five mandatory axes (genome-property frontier, marker-gene census, family copy-number, synteny/neighborhoods, ncRNA census) for every query genome, with rows for axes that produced negative findings. - Final report states what is interesting, what is conserved/expected, what is likely artifact, and what should be tested next.
- Grouped splits have no sample or group overlap, calibration is reported, confounding and imbalance are quantified, and the model is compared with a declared null baseline.
Examples
Example 1: Expected input layout
results/*.parquet or results/*.tsv
metadata.tsv
Troubleshooting
Issue: Missing inputs or reference databases Solution: Verify paths and permissions before running the workflow.
Issue: Low-quality results or failed QC gates Solution: Review reports, adjust parameters, and re-run the affected step.