Bio phylogenomics
npx -y skills add fmschulz/omics-skills --skill bio-phylogenomicsAssembled from the repository path, not quoted from the project. Check it against their README if it does not work.
One thing to look at
- 7 stars7 stars. Stars are a popularity signal and not a quality one, but at this level it is likely that nobody has read this closely except its author, and you would be relying on your own review.
What its author says it does
Copied from the file, not written here
Build and validate marker-gene alignments and phylogenetic trees. Use when inferring evolutionary relationships, choosing models, or checking tree support and contamination.
SKILL.md
5.4 KB, as published. Nobody here has run it
Bio Phylogenomics
Build marker gene alignments and phylogenetic trees.
Instructions
-
Validate marker/reference manifests and create a checksum-gated, fixed-seed execution plan:
uv run --no-project python skills/bio-phylogenomics/scripts/run_phylogenomics.py \ markers.tsv --references references.tsv --seed 1729 \ --out results/bio-phylogenomics # Inspect run_manifest.json, then add --execute.The driver restarts only from non-empty stage outputs and normalizes internal support values from either 0–1 or 0–100 notation to
support.tsvon a 0–1 scale. -
Extract marker genes or SSU rRNA sequences.
-
Align with MAFFT v7.5+ and trim with trimAl v1.4 (or ClipKIT when phylogenetically-informed trimming is preferred).
-
Build ML trees with support values. Choose by objective first, then leaf count:
- Exploratory placement, benchmark iterations, reference-set screening, or any time-bounded analysis: use VeryFastTree v4.0 first, even below ~2,000 taxa. Prefer
VeryFastTree -boot 1000 -threads <n> < alignment.faa > tree.nwfor proteins and add-ntfor nucleotide alignments. - Final or publication-quality trees up to ~2,000 taxa: IQ-TREE v3 (v3.1.2+) for comprehensive model selection, MAST/GTRpmix, UFBoot/SH-aLRT, and defensible final inference.
- Above ~2,000 taxa, or when memory/runtime is uncertain: VeryFastTree v4.0 (multi-threaded, SIMD,
-disk-computingfor very large trees). - Use
iqtree3 -fastonly when VeryFastTree is unavailable or a project explicitly requires IQ-TREE-compatible exploratory output; record that fallback in the report.
- Exploratory placement, benchmark iterations, reference-set screening, or any time-bounded analysis: use VeryFastTree v4.0 first, even below ~2,000 taxa. Prefer
-
Post-process trees with ETE v4 (
ete4):- Compute tree statistics (branch lengths, distances, topology metrics).
- Root, prune, or collapse nodes as needed.
- Filter by bootstrap support.
- Add taxonomic or trait annotations.
- Generate publication-quality visualizations.
-
Use the literature-derived analysis playbook to choose markers, reference sampling, rooting, and placement strategy appropriate for the inferred group.
-
Identify nearest neighbors and closest named relatives for each query sequence/genome when the chosen marker/reference set supports that interpretation.
-
Export a closest-relatives table with support values, distances, taxonomy, reference accessions, and uncertainty notes.
-
Fetch and persist the close-relative genomes and proteomes that downstream comparative analyses will use. Save under
results/bio-phylogenomics/relatives/{accession}/genome.fnaandproteins.faa, plusrelatives_manifest.tsvrecording accession, source DB, taxonomy, genome size, gene count, and the reason for inclusion. If a relative cannot be downloaded, record the failure explicitly. Without this artifact, the comparative axes downstream cannot run. -
Use well-supported relatives or a documented broader comparison set to guide downstream comparative analysis with
/bio-protein-clustering-pangenomeand/bio-annotation.
Quick Reference
| Task | Action |
|---|---|
| Run workflow | Follow the steps in this skill and capture outputs. |
| Validate inputs | Confirm required inputs and reference data exist. |
| Review outputs | Inspect reports and QC gates before proceeding. |
| Tool docs | See docs/README.md. |
Input Requirements
Prerequisites:
- Tools declared in the project's pinned Pixi environment. See
docs/README.mdfor expected tools. - Marker gene set or alignments available. Inputs:
- markers.faa (marker genes) or alignments.fasta
Output
- results/bio-phylogenomics/alignments/
- results/bio-phylogenomics/trees/
- results/bio-phylogenomics/closest_relatives.tsv
- results/bio-phylogenomics/relatives/{accession}/genome.fna
- results/bio-phylogenomics/relatives/{accession}/proteins.faa
- results/bio-phylogenomics/relatives_manifest.tsv
- results/bio-phylogenomics/phylo_report.md
- results/bio-phylogenomics/logs/
Quality Gates
- Alignment length and missingness meet project thresholds.
- Every reference checksum matches before alignment, and the run manifest records a positive fixed seed.
- Internal supports are exported on a documented 0–1 scale without mixing raw IQ-TREE and VeryFastTree conventions.
- Bootstrap support summary meets project thresholds.
- On failure: retry with alternative parameters; if still failing, record in report and exit non-zero.
- Verify markers.faa is non-empty and aligned sequences are consistent.
- Marker and reference choices are justified against the literature-derived analysis playbook.
- Closest relatives are reported with support/distance metrics or uncertainty is stated.
- Tree interpretation distinguishes well-supported nearest relatives from weakly supported placements.
-
relatives_manifest.tsvis populated and the matching genome/proteome files are present on disk (or each failure is recorded with a reason).
Examples
Example 1: Expected input layout
markers.faa (marker genes) or alignments.fasta
Troubleshooting
Issue: Missing inputs or reference databases Solution: Verify paths and permissions before running the workflow.
Issue: Low-quality results or failed QC gates Solution: Review reports, adjust parameters, and re-run the affected step.