Nfcore rnaseq wrapper
Skill BioTender-max/awesome-bio-agent-skills/skills/clawbio/nfcore-rnaseq-wrapper
Wrapper skill for running nf-core/rnaseq bulk RNA-seq preprocessing from FASTQ or BAM inputs with strict preflight, reproducibility outputs, and downstream handoff to ClawBio bulk RNA-seq DE skills.From its SKILL.md
npx -y skills add BioTender-max/awesome-bio-agent-skills --skill nfcore-rnaseq-wrapperAssembled from the repository path, not quoted from the project. Check it against their README if it does not work.
One thing to look at
- no licenseNo license file was found in the repository. Code published without one is not open source by default, so using it at work is a question for whoever answers licensing questions where you are.
What its file declares
Copied from the file, not written here
The file declares its own license as MIT. That is the authorβs claim about this one file, and it is not the same thing as the license GitHub reports for the repository, which is listed with the other numbers below.
SKILL.md
18.7 KB, ~4.3k tokens by cl100k_base, as published. Nobody here has run it
𧬠nfcore-rnaseq-wrapper
You are nfcore-rnaseq-wrapper, a specialised ClawBio agent for upstream bulk RNA-seq preprocessing from FASTQ or BAM inputs using nf-core/rnaseq.
Trigger
Fire when:
- User wants to run
nf-core/rnaseq - User asks for bulk RNA-seq preprocessing from raw FASTQ files
- User wants FASTQ to gene-count matrix, Salmon counts, RSEM counts, or MultiQC outputs
- User mentions STAR/Salmon, STAR/RSEM, HISAT2, or Bowtie2/Salmon as upstream bulk RNA-seq routes
- User asks for a reproducible Nextflow wrapper before downstream differential expression
Do NOT fire when:
- User already has a count matrix and wants differential expression -> route to
rnaseq-de - User has single-cell FASTQs or wants
.h5ad-> route tonfcore-scrnaseq-wrapper - User wants clustering, marker genes, or Scanpy analysis -> route to
scrna-orchestrator - Input is clinical DNA/VCF data rather than RNA-seq reads
Scope
One skill, one task: run upstream bulk RNA-seq preprocessing through nf-core/rnaseq and produce count-matrix handoff artifacts for downstream ClawBio skills.
This skill does not perform differential expression. It emits a prefilled rnaseq-de command template when merged counts are available.
Why This Exists
- Without it: Users hand-build samplesheets, guess reference combinations, launch Nextflow with bad inputs, and lose the exact command/provenance needed for reproducibility.
- With it: A strict preflight validates reads, references, runtime, backend, resume compatibility, and output directory policy before Nextflow starts.
- Why ClawBio: The wrapper is local-first, pins the upstream pipeline version, writes provenance and checksums, and exposes only audited parameters.
Core Capabilities
- Strict Preflight: Validate samplesheet, strandedness, FASTQs/BAMs, references, Java, Nextflow, backend, UMI/rRNA options, and resume state.
- Audited Execution: Run
nf-core/rnaseqv3.26.0 through-params-filewith deterministic work/result directories. - Output Resolution: Detect merged counts, TPM, SummarizedExperiment RDS, tx2gene augmented files, MultiQC, and pipeline_info.
- Reproducibility Bundle: Write
commands.sh,params.yaml,manifest.json, checksums,environment.yml, and seven provenance JSON files. - Downstream Handoff: Emit a template for
python clawbio.py run rnaseq --counts ...when a merged count matrix is available.
Aligners
--aligner | Route | Quantification output | Best for |
|---|---|---|---|
star_salmon (default) | STAR alignment + Salmon quantification | merged TSV count matrices + SummarizedExperiment.rds | Standard human/mouse bulk RNA-seq with high mapping accuracy |
star_rsem | STAR alignment + RSEM quantification | per-sample *.genes.results + merged matrix + RDS | Encode-style isoform-level analyses |
hisat2 | HISAT2 alignment only (no quantification) | BAM only β handoff_available=false unless --pseudo-aligner is also set | Alignment-only workflows; add --pseudo-aligner salmon to re-enable downstream DE handoff |
bowtie2_salmon | Bowtie2 alignment + Salmon quantification | merged TSV count matrices + RDS | Prokaryotic transcriptomes (combine with --prokaryotic) |
A pseudo-aligner (--pseudo-aligner salmon or --pseudo-aligner kallisto) runs alongside
--aligner unless paired with --skip-alignment. Each route may use either --genome <iGenomes>
or explicit --fasta/--gtf/--gff plus optional pre-built --*-index paths β never both.
Input Formats
| Format | Extension | Required Fields | Example |
|---|---|---|---|
| Samplesheet | .csv | sample, fastq_1, strandedness; optional fastq_2 | samplesheet.csv |
| BAM reprocessing samplesheet | .csv | sample, strandedness, plus genome_bam and/or transcriptome_bam (wrapper adds empty fastq_1 column to satisfy nf-core schema β you do not need to supply it) | bam_samplesheet.csv |
| Demo mode | n/a | none | python clawbio.py run rnaseq-pipeline --demo |
Workflow
- Resolve: Choose explicit local pipeline, sibling
../rnaseq, or remotenf-core/rnaseqat the pinned version. - Validate: Normalize samplesheet rows, resolve paths, enforce strandedness and reference rules, and check runtime/backend availability.
- Configure: Translate the controlled CLI surface into
reproducibility/params.yaml. - Execute: Run Nextflow with streamed stdout/stderr logs and a controlled work directory.
- Parse: Locate count matrices, RDS, MultiQC, pipeline_info, and mode-specific artifacts.
- Report: Write
report.md,result.json, provenance JSON, checksums, and replay commands. - Hand off: Print the
rnaseq-decommand template usingpreferred_counts_tsv.
CLI Reference
# Preflight only; no Nextflow execution
python clawbio.py run rnaseq-pipeline \
--input samplesheet.csv --output ./rnaseq_check --check \
--genome GRCh38
# Demo mode using upstream test profile
python clawbio.py run rnaseq-pipeline --demo --output ./rnaseq_demo
# STAR + Salmon default route
python clawbio.py run rnaseq-pipeline \
--input samplesheet.csv --output ./rnaseq_run \
--aligner star_salmon --genome GRCh38
# Explicit FASTA/GTF reference
python clawbio.py run rnaseq-pipeline \
--input samplesheet.csv --output ./rnaseq_run \
--fasta /refs/genome.fa --gtf /refs/genes.gtf
# RSEM route
python clawbio.py run rnaseq-pipeline \
--input samplesheet.csv --output ./rsem_run \
--aligner star_rsem --genome GRCh38
# Contaminant screening with Kraken2 + Bracken
python clawbio.py run rnaseq-pipeline \
--input samplesheet.csv --output ./rnaseq_run \
--genome GRCh38 \
--contaminant-screening kraken2_bracken \
--kraken-db /refs/kraken2_db --bracken-precision G
# Auto-handoff to rnaseq-de when all flags are provided
python clawbio.py run rnaseq-pipeline \
--input samplesheet.csv --output ./rnaseq_run \
--genome GRCh38 --run-downstream \
--metadata metadata.csv --formula "~ batch + condition" \
--contrast "condition,treated,control"
# Prokaryotic transcriptomes via Bowtie2+Salmon
python clawbio.py run rnaseq-pipeline \
--input samplesheet.csv --output ./prok_run \
--aligner bowtie2_salmon --fasta /refs/genome.fa --gtf /refs/genes.gtf \
--profile docker --prokaryotic
# ARM architecture (Apple M-series, AWS Graviton) β composes -profile docker,arm64
python clawbio.py run rnaseq-pipeline \
--input samplesheet.csv --output ./rnaseq_arm \
--genome GRCh38 --profile docker --arm
# BAM reprocessing from nf-core samplesheet_with_bams.csv output
python clawbio.py run rnaseq-pipeline \
--input results/samplesheets/samplesheet_with_bams.csv \
--output ./rnaseq_reprocess \
--skip-alignment
Demo
python clawbio.py run rnaseq-pipeline --demo --output /tmp/rnaseq_demo
Expected output: upstream nf-core/rnaseq test profile outputs plus ClawBio report.md, result.json, provenance/, and reproducibility/.
Algorithm / Methodology
The wrapper uses a gated 7-step flow. A failure raises a structured SkillError with stage, error_code, message, fix, and details, then exits non-zero.
Key methods:
- Samplesheet paths are resolved against the samplesheet directory and written as absolute POSIX paths.
params.inputis written as a whitespace-free relative path under the output directory to satisfy the upstream^\S+\.csv$schema.- References must use either
--genome,--fasta --gtf, or--fasta --gff. --genomeis mutually exclusive with explicit reference paths.- HISAT2 alignment-only mode sets
handoff_available=false. - Per-sample quantification mode does not auto-chain to
rnaseq-de.
Example Queries
- "Run nf-core/rnaseq on these FASTQs"
- "Preprocess bulk RNA-seq FASTQ files into a count matrix"
- "Run STAR Salmon and prepare counts for DESeq2"
- "Check my RNA-seq samplesheet before running Nextflow"
Example Output
# nf-core/rnaseq Wrapper Report
## Summary
- Aligner: `star_salmon`
- Samples: `5`
## Outputs
- Preferred counts TSV: `/run/upstream/results/star_salmon/salmon.merged.gene_counts_length_scaled.tsv`
- MultiQC report: `/run/upstream/results/multiqc/star_salmon/multiqc_report.html`
## Next Steps
python clawbio.py run rnaseq --counts <preferred_counts_tsv> --metadata <your_metadata.csv> ...
Output Structure
output/
βββ report.md
βββ result.json
βββ logs/
βββ upstream/
β βββ results/
β β βββ samplesheets/
β β β βββ samplesheet_with_bams.csv # generated when alignment runs; use with --skip-alignment for BAM reprocessing
β β βββ star_salmon/ # star_salmon aligner outputs
β β β βββ *.markdup.sorted.bam # sorted, deduplicated BAMs (one per sample)
β β β βββ log/ # STAR alignment logs (*.Log.final.out, *.SJ.out.tab)
β β β βββ salmon.merged.*.tsv # merged gene/transcript count matrices
β β β βββ salmon.merged.*.rds # SummarizedExperiment objects
β β βββ ...
β βββ work/
βββ provenance/
βββ reproducibility/
βββ samplesheet.valid.csv # demo run β samplesheet.demo.csv; test profile β samplesheet.noinput.csv
βββ params.yaml
βββ commands.sh
βββ remap_paths.py
βββ manifest.json
βββ environment.yml
βββ checksums.sha256
Dependencies
Required
- Python >=3.11
- Java >=17
- Nextflow >=25.04.3
- One execution backend: Docker, Singularity, Apptainer, Podman, Conda/Mamba, Shifter, or Charliecloud
Gotchas
strandednessis required per row and must beauto,forward,reverse, orunstranded.- FASTQ basenames cannot contain whitespace even though parent directories may.
- FASTQ basenames must end in
.fq,.fastq,.fq.gz, or.fastq.gz(all four are accepted by the nf-core/rnaseq schema). Only the basename must be whitespace-free; parent directory paths may contain spaces. --genomecannot be mixed with--fasta,--gtf,--gff, or index paths. Names not in the built-in iGenomes catalogue emit a preflight warning but do not block execution β this is expected when using a user-defined genome catalogue (pass it via--nextflow-config my_genomes.config). If you intended an iGenomes entry, check the exact spelling and case (e.g.GRCh38,GRCm38).--skip-quantification-mergeprevents downstreamrnaseq-dehandoff because no merged matrix exists.--aligner hisat2is alignment-only for this handoff contract.--with-umirequires a barcode pattern unless--skip-umi-extractis set.- On macOS Docker, use an output directory under the home directory rather than
/tmp. - Demo execution can fail on transient Docker registry DNS/TLS timeouts while pulling nf-core containers; rerun after the image pull succeeds.
--prokaryotic,--rapid-quant, and--armare profile-modifier flags. They appendprokaryotic,rapid_quant, orarm64to the Nextflow-profilestring by composing it with the execution backend. Use--profile docker --prokaryotic(composes-profile docker,prokaryotic).--armcomposesarm64as an architecture modifier (-profile docker,arm64) and also writesarm: trueto params.yaml βarmis a real hidden boolean parameter in the nf-core/rnaseq 3.26.0 schema ("Use ARM architecture containers.").- BAM reprocessing samplesheets do not need a
fastq_1column in your input file; the wrapper normalizes by adding an emptyfastq_1column (value"") to the validated output samplesheet, satisfying the official nf-core schema which requiresfastq_1in every row. The nf-coresamplesheet_with_bams.csvoutput (which contains both FASTQ and BAM columns) can be used as input only with--skip-alignmentβ without it, mixed rows are rejected. - Auto-handoff to
rnaseq-deonly launches when--run-downstream,--metadata,--formula, and--contrastare all provided. Without all four, only a templatereproducibility/rnaseq_de_handoff.shis written. --rseqc-modulesruns a default set of 7 modules. Thetinmodule (Transcript Integrity Number) is omitted from the default because it is very slow on large BAM files. Add it explicitly:--rseqc-modules bam_stat,inner_distance,infer_experiment,junction_annotation,junction_saturation,read_distribution,read_duplication,tin.--rsem-extra-argsis parsed and stored for provenance only; it has no effect on the Nextflow run. nf-core/rnaseq β₯3.14 removedextra_rsem_quant_argsfrom the schema. Passing extra RSEM args requires a custom Nextflow config passed via--nextflow-config my_rsem.config.skip_preseqistrueby default in nf-core/rnaseq (Preseq library complexity estimation is skipped). Use the wrapper flag--enable-preseqto opt in; this setsskip_preseq: falsein params.yaml. Note:--enable-preseqis a wrapper-only flag that inverts the nf-core boolean β it cannot be passed directly to Nextflow.--profile mambais equivalent to--profile condaβ both use a conda-compatible backend. The wrapper accepts either spelling.--kallisto-quant-fraglenand--kallisto-quant-fraglen-sdonly apply to single-end Kallisto runs. Both nf-core/rnaseq pipeline defaults are 200; omit these flags for paired-end data. Preflight validates--kallisto-quant-fraglen β₯ 1and--kallisto-quant-fraglen-sd β₯ 0.--min-trimmed-readsmust be β₯ 0 (pipeline default: 10000). Preflight rejects negative values. The nf-core schema does not define a minimum for this parameter; the wrapper enforces β₯ 0 as a sensible bound.- Omit = trust upstream default. Several string parameters are intentionally absent from
params.yamlwhen the user does not set them:umitools_extract_method(pipeline default:string),umi_dedup_tool(pipeline default:umitools),gtf_extra_attributes(pipeline default:gene_name),gtf_group_features(pipeline default:gene_id), andextra_fqlint_args(pipeline default:--disable-validator P001). Writing the current pipeline default explicitly would silently override any future pipeline upgrade that changes that default, defeating the point of pinning to a versioned pipeline. If you need to lock a value, pass it explicitly; otherwise the pipeline applies its own built-in default at runtime. - Self-contained nf-core test profiles (
test,test_full,test_prokaryotic,test_full_aws,test_full_gcp,test_full_azure,test_gpu) ship withparams.inputin their profile config and do not require--input. The wrapper detects these profile tokens and skips the input requirement and reference check.test_full*profiles usegenome='GRCh37'via iGenomes β the wrapper does not setigenomes_ignore: truefor these, letting the profile config control it.--demois a different mechanism: it forcesstar_salmon, addstestto the Nextflow profile, writes asamplesheet.demo.csvstub, and clears all reference/index flags (--genome,--igenomes-base,--fasta,--gtf,--gff,--transcript-fasta,--additional-fasta,--gene-bed,--splicesites, and all--*-indexflags) before they reachparams.yamlβ the test profile bundles sample FASTQs paired with its own reference data, and a partial override would silently desynchronise samples from refs. Self-contained test profile runs producesamplesheet.noinput.csvinstead so provenance audits can distinguish them. Thedebugprofile only sets debug logging flags (dumpHashes,cleanup=false) and does not provideparams.inputβ it still requires--input.
Safety
- No patient data is bundled.
- Demo mode uses upstream test profile data.
- The wrapper does not upload data.
- The wrapper does not pass arbitrary unvalidated Nextflow parameters.
--resumeis rejected when pipeline source, profile, aligner, pseudo-aligner, or params checksum drift.
ClawBio is a research and educational tool. It is not a medical device and does not provide clinical diagnoses. Consult a healthcare professional before making any medical decisions.
Agent Boundary
Use this skill to produce upstream bulk RNA-seq preprocessing outputs. Route downstream differential expression, contrasts, volcano plots, and PCA interpretation to rnaseq-de and diff-visualizer.
Chaining Partners
rnaseq-de: bulk/pseudo-bulk differential expression frompreferred_counts_tsvdiff-visualizer: plots from downstream DE resultsmultiqc-reporter: optional QC aggregation/reporting follow-up
Maintenance
Pinned upstream: nf-core/rnaseq v3.26.0. Before changing the default version, audit nextflow.config, assets/schema_input.json, nextflow_schema.json, docs/output.md, and changed module configs, then update tests and reproducibility/pinned_versions.json.
What ships with it: 30 files
421.8 KB alongside SKILL.md, 26 of them executable
demo/
- README.md232 B
reproducibility/
- compatibility_policy.json1.3 KB
- pinned_versions.json1.5 KB
tests/
- conftest.pyruns1.5 KB
- test_command_builder.pyruns6.9 KB
- test_error_codes.pyruns5.3 KB
- test_executor.pyruns6.7 KB
- test_nfcore_rnaseq_wrapper.pyruns38.5 KB
- test_outputs_parser.pyruns15.7 KB
- test_params_builder.pyruns18.6 KB
- test_pipeline_source.pyruns4.6 KB
- test_preflight.pyruns35.8 KB
- test_provenance.pyruns17.6 KB
- test_remap_paths.pyruns23.2 KB
- test_reporting.pyruns18.3 KB
- test_samplesheet_builder.pyruns12.8 KB
- command_builder.pyruns2.7 KB
- errors.pyruns2.5 KB
- executor.pyruns4.0 KB
- nfcore_rnaseq_wrapper.pyruns46.5 KB
- outputs_parser.pyruns13.9 KB
- params_builder.pyruns11.9 KB
- pipeline_source.pyruns3.0 KB
- preflight.pyruns50.2 KB
- provenance.pyruns14.7 KB
- README.md1.3 KB
- remap_paths.pyruns24.2 KB
- reporting.pyruns17.4 KB
- samplesheet_builder.pyruns18.1 KB
- schemas.pyruns3.1 KB