Bio reads qc mapping
npx -y skills add fmschulz/omics-skills --skill bio-reads-qc-mappingAssembled from the repository path, not quoted from the project. Check it against their README if it does not work.
One thing to look at
- 7 stars7 stars. Stars are a popularity signal and not a quality one, but at this level it is likely that nobody has read this closely except its author, and you would be relying on your own review.
What its author says it does
Copied from the file, not written here
Ingest, quality-control, and map sequencing reads with reproducible outputs. Use when processing raw reads, removing contaminants, or calculating mapping and coverage statistics.
SKILL.md
5.9 KB, as published. Nobody here has run it
Bio Reads QC Mapping
Ingest, QC, and map reads with reproducible outputs. Use for raw read processing and coverage stats.
Instructions
-
Parse and validate
sample_sheet.tsvagainstschemas/sample-sheet.schema.json. Use the executable driver for both planning and restartable execution:uv run --no-project python skills/bio-reads-qc-mapping/scripts/run_reads_qc_mapping.py \ sample_sheet.tsv --out results/bio-reads-qc-mapping # Inspect run_manifest.json, then execute the same plan: uv run --no-project python skills/bio-reads-qc-mapping/scripts/run_reads_qc_mapping.py \ sample_sheet.tsv --out results/bio-reads-qc-mapping --executeread_typemust bepaired_short,single_short, orlong. Mapping is scheduled only for rows with a non-emptyreference; a missing reference is not a mapping failure. -
For short reads: run QC and adapter/quality trimming with
bbdukorfastpv1.3.3+. -
For long reads: use current basecaller-aware QC first. For ONT, prefer Dorado summaries/trimming during basecalling or demultiplexing when starting from signal/BAM; for FASTQ-only filtering use
chopperfor quality/length/end trimming orfiltlongv0.2.1 when selecting reads for assembly. UsePychopperfor full-length cDNA. TreatPorechop_ABIas a targeted legacy/fallback adapter-discovery tool, and record why it is needed.- For very large ONT FASTQ inputs, do not burn the first full read pass on raw
gzip -tor rawseqkit statspreflight unless the user explicitly asks for it. Record rawstatmetadata and, if needed, a small sampled sanity check; let the first full pass be the actual filtering/orientation step, then runseqkit statson produced outputs. - For ONT cDNA with
Pychopper, write outputs with plain.fastqsuffixes unless you explicitly pipe/compress them yourself.Pychoppercan write plain FASTQ even when the output path ends in.gz; avoidgzip -tonPychopperoutputs unless magic bytes confirm gzip. If legacy outputs have.fastq.gznames but plain FASTQ content, rename them to.fastqbefore resuming. Pychopperreport plotting can fail after the reads are already processed, for example from a pandas/statistics type-conversion error. On that failure, inspect whether the classified/unclassified/rescued/read-stats outputs exist and are non-empty. If they do, resume downstream from those outputs rather than rerunning the fullPychopperpass.
- For very large ONT FASTQ inputs, do not burn the first full read pass on raw
-
Map reads and produce coverage tables:
- Short reads, CPU:
bbmaporbwa-mem2v2.2.1+. Short reads, GPU node available: NVIDIA Parabricksfq2bam(wrapsbwa-mem2+ GATK markdup; typically 3–4× faster thanbwa-mem2on 8 cores and up to ~80× over a 96-core CPU pipeline). - Long reads, CPU:
minimap2v2.30+. AVX-512 hardware:mm2-fastas a drop-in replacement (~1.8× speedup). GPU node available:mm2-gbormm2-axfor CUDA-accelerated long-read alignment.
- Short reads, CPU:
-
Record the tool, version, and any GPU device used in the run log.
Quick Reference
| Task | Action |
|---|---|
| Run workflow | Follow the steps in this skill and capture outputs. |
| Validate inputs | Confirm required inputs and reference data exist. |
| Review outputs | Inspect reports and QC gates before proceeding. |
| Tool docs | See docs/README.md. |
Input Requirements
Prerequisites:
- Tools declared in the project's pinned Pixi environment. See
docs/README.mdfor expected tools. - Sample sheet and reads are available. Inputs:
- sample_sheet.tsv
- reads/*.fastq.gz
- reference.fasta (optional)
Output
- results/bio-reads-qc-mapping/trimmed_reads/
- results/bio-reads-qc-mapping/qc_reports/
- results/bio-reads-qc-mapping/mapping_stats.tsv
- results/bio-reads-qc-mapping/coverage.tsv
- results/bio-reads-qc-mapping/logs/
Quality Gates
- Post-QC read count sanity checks pass.
- Mapping rate meets project thresholds.
- On failure: retry with alternative parameters; if still failing, record in report and exit non-zero.
- Validate sample sheet schema and FASTQ integrity.
- The plan covers every sheet row exactly once, and mapping gates are applied only to rows that supplied a reference.
- For long-read QC, record whether trimming happened in the basecaller/demultiplexer,
chopper,filtlong,Pychopper, or a documented Porechop_ABI fallback. - For huge ONT inputs, avoid redundant full-file raw preflights; document raw file size/mtime and make the first full pass productive.
- For
Pychopperoutputs, verify actual file type by content, not suffix. Plain FASTQ with a.gzsuffix must be renamed or explicitly compressed before downstream tools that expect gzip. - Resume guards should skip expensive completed steps only after confirming the expected output exists, is non-empty, and passes a lightweight content sanity check (
seqkit stats, FASTQ header sniff, or gzip magic as appropriate).
Examples
Example 1: Expected input layout
sample_sheet.tsv
reads/*.fastq.gz
reference.fasta (optional)
The runnable fixture at fixtures/sample_sheet.tsv covers paired-end, single-end, and long reads.
Troubleshooting
Issue: Missing inputs or reference databases Solution: Verify paths and permissions before running the workflow.
Issue: Low-quality results or failed QC gates Solution: Review reports, adjust parameters, and re-run the affected step.
Issue: Pychopper failed during report/stat plotting but output FASTQs exist
Solution: Treat this as a recoverable post-processing failure. Confirm the classified FASTQ is non-empty and readable, fix any misleading .gz suffix, run seqkit stats, and resume downstream steps from the existing Pychopper outputs.