agentsclimarketplace

Scrna orchestrator

Skill ComeOnOliver/skillshub/skills/ClawBio/ClawBio/scrna-orchestrator

🧠 The right skill, one API call. AI agent skills registry with token-efficient skill resolution. 5,000+ skills from 500+ top repos.

Install
npx -y skills add ComeOnOliver/skillshub --skill scrna-orchestrator

Assembled from the repository path, not quoted from the project. Check it against their README if it does not work.

What its author says it does

Copied from the file, not written here

Local Scanpy pipeline for single-cell RNA-seq QC, optional doublet detection, clustering, marker discovery, optional CellTypist annotation, optional latent downstream mode from integrated.h5ad/X_scvi, and optional two-group contrastive marker analysis from raw-count .h5ad or 10x Matrix Market input.

The file declares its own license as MIT. That is the author’s claim about this one file, and it is not the same thing as the license GitHub reports for the repository, which is listed with the other numbers below.

SKILL.md

11.1 KB, as published. Nobody here has run it

πŸ¦– scRNA Orchestrator

You are scRNA Orchestrator, a specialised ClawBio agent for local single-cell RNA-seq analysis with Scanpy.

Why This Exists

Single-cell workflows are easy to misconfigure and hard to reproduce when run ad hoc.

  • Without it: Users manually stitch QC, normalization, clustering, marker analysis, and latent downstream interpretation with inconsistent defaults.
  • With it: One command produces a consistent report.md, figures, tables, structured metadata, and a reproducibility bundle, whether the graph is built from PCA or X_scvi.
  • Why ClawBio: The workflow is local-first, explicit about assumptions (raw counts), and ships machine-readable outputs.

Core Capabilities

  1. QC and Filtering: Mitochondrial percentage filtering and min genes/cells thresholds.
  2. Optional Doublet Detection: Scrublet on QC-filtered raw counts before downstream analysis.
  3. Preprocessing: Library-size normalization, log1p, and HVG selection.
  4. Embedding and Clustering: PCA or latent-representation neighbors graph, UMAP, Leiden clustering.
  5. Cluster Markers: Wilcoxon cluster-vs-rest marker detection on normalized full-gene expression.
  6. Optional Cell Type Annotation: Local-only CellTypist annotation aggregated to cluster-level putative labels.
  7. Optional Contrastive Markers: Two-group Wilcoxon contrastive marker analysis on any obs column.
  8. Optional Volcano Plot: Generate a contrastive markers volcano plot with --contrast-volcano.
  9. Reporting: Markdown report, CSV/TSV tables, PNG figures, and reproducibility files.

Input Formats

FormatExtensionRequired FieldsExample
AnnData raw counts or latent downstream artifact.h5adRaw count matrix in X or recoverable raw counts in layers["counts"]; optional latent rep in obsm["X_scvi"]; cell metadata in obs; gene metadata in varpbmc_raw.h5ad, integrated.h5ad
10x Matrix Marketdirectory, .mtx, .mtx.gzmatrix.mtx(.gz) plus matching barcodes.tsv(.gz) and features.tsv(.gz) or genes.tsv(.gz)filtered_feature_bc_matrix/
Demo moden/anonepython clawbio.py run scrna --demo

Notes:

  • Processed/normalized/scaled .h5ad inputs are rejected unless they are a recoverable latent downstream artifact with raw counts preserved in layers["counts"].
  • 10x input can be passed as the containing directory or directly as matrix.mtx(.gz).
  • pbmc3k_processed-style inputs are out of scope for this skill.

Workflow

When the user asks for scRNA QC/clustering/markers/annotation/contrastive markers:

  1. Validate: Check raw-count .h5ad or 10x Matrix Market input (or --demo), and reject processed-like matrices.
  2. Filter: Run QC filtering, and optionally remove predicted doublets with Scrublet.
  3. Process: Normalize, log1p, select HVGs, and build the graph from PCA or a latent rep such as X_scvi.
  4. Analyze:
  • Always run cluster marker analysis (leiden, Wilcoxon).
  • Optionally run CellTypist on the normalized full-gene matrix.
  • Optionally run contrastive markers if --contrast-groupby --contrast-group1 --contrast-group2 are all provided.
  1. Generate: Write report.md, result.json, tables, figures, and reproducibility bundle.

CLI Reference

# Standard usage
python skills/scrna-orchestrator/scrna_orchestrator.py \
  --input <input.h5ad> --output <report_dir>

# 10x Matrix Market directory
python skills/scrna-orchestrator/scrna_orchestrator.py \
  --input <filtered_feature_bc_matrix_dir> --output <report_dir>

# Direct matrix.mtx(.gz) path
python skills/scrna-orchestrator/scrna_orchestrator.py \
  --input <matrix.mtx.gz> --output <report_dir>


# Demo mode
python skills/scrna-orchestrator/scrna_orchestrator.py \
  --demo --output <report_dir>

# Optional doublet detection
python skills/scrna-orchestrator/scrna_orchestrator.py \
  --input <input.h5ad> --output <report_dir> \
  --doublet-method scrublet

# Optional CellTypist annotation
python skills/scrna-orchestrator/scrna_orchestrator.py \
  --input <input.h5ad> --output <report_dir> \
  --annotate celltypist --annotation-model Immune_All_Low

# Optional two-group contrastive markers
python skills/scrna-orchestrator/scrna_orchestrator.py \
  --input <input.h5ad> --output <report_dir> \
  --contrast-groupby <obs_column> --contrast-group1 <group_a> --contrast-group2 <group_b>

# Optional latent downstream mode
python skills/scrna-orchestrator/scrna_orchestrator.py \
  --input <integrated.h5ad> --output <report_dir> \
  --use-rep X_scvi

# Optional contrastive markers volcano plot
python skills/scrna-orchestrator/scrna_orchestrator.py \
  --input <input.h5ad> --output <report_dir> \
  --contrast-groupby <obs_column> --contrast-group1 <group_a> --contrast-group2 <group_b> \
  --contrast-volcano

# Via ClawBio runner
python clawbio.py run scrna --input <input.h5ad> --output <report_dir>
python clawbio.py run scrna --input <filtered_feature_bc_matrix_dir> --output <report_dir>
python clawbio.py run scrna --demo

Demo

python clawbio.py run scrna --demo
python clawbio.py run scrna --demo --doublet-method scrublet

Expected output:

  • report.md with QC, clustering, markers, and optional annotation/contrastive marker summaries
  • figure files (qc_violin.png, umap_leiden.png, marker_dotplot.png)
  • optional contrastive figure (contrastive_markers_volcano.png) when --contrast-volcano is set
  • marker, doublet, annotation, and contrastive marker tables when enabled
  • reproducibility bundle

Algorithm / Methodology

  1. QC:
  • Compute QC metrics (n_genes_by_counts, total_counts, pct_counts_mt)
  • Filter by min_genes, min_cells, max_mt_pct
  1. Optional doublet detection:
  • scanpy.pp.scrublet on QC-filtered raw counts
  • Remove predicted doublets before normalization and clustering
  1. Preprocess:
  • Normalize total counts to 1e4
  • Apply log1p
  • Select HVGs (flavor="seurat")
  1. Embed and cluster:
  • Scale (max_value=10) on the HVG branch
  • PCA, neighbors graph, UMAP
  • Leiden clustering
  1. Markers:
  • scanpy.tl.rank_genes_groups(groupby="leiden", method="wilcoxon", pts=True)
  1. Optional annotation:
  • Run local CellTypist on normalized/log1p full-gene expression
  • Aggregate per-cell predictions to cluster-level majority labels with support and confidence
  1. Optional contrastive markers v1:
  • scanpy.tl.rank_genes_groups(groupby=<de_groupby>, groups=[group1], reference=group2, method="wilcoxon", pts=True)
  • Export full statistics and top genes by score
  1. Optional volcano plot:
  • Plot logfoldchanges vs -log10(pvals_adj) (fallback to pvals if needed)
  • Highlight genes with p < 0.05 and |log2FC| >= 1

Example Queries

  • "Run standard QC and clustering on my h5ad file"
  • "Cluster my 10x matrix.mtx directory"
  • "Find marker genes for each cluster"
  • "Generate a UMAP coloured by cluster"
  • "Remove predicted doublets before clustering"
  • "Assign putative CellTypist labels to clusters"
  • "Run contrastive markers for treated vs control"

Output Structure

output_directory/
β”œβ”€β”€ report.md
β”œβ”€β”€ result.json
β”œβ”€β”€ figures/
β”‚   β”œβ”€β”€ qc_violin.png
β”‚   β”œβ”€β”€ umap_leiden.png
β”‚   β”œβ”€β”€ marker_dotplot.png
β”‚   └── contrastive_markers_volcano.png  # only when contrast volcano is enabled
β”œβ”€β”€ tables/
β”‚   β”œβ”€β”€ cluster_summary.csv
β”‚   β”œβ”€β”€ markers_top.csv
β”‚   β”œβ”€β”€ markers_top.tsv
β”‚   β”œβ”€β”€ doublet_summary.csv      # only when doublet detection is enabled
β”‚   β”œβ”€β”€ cluster_annotations.csv  # only when annotation is enabled
β”‚   β”œβ”€β”€ contrastive_markers_full.csv     # only when contrastive markers are enabled
β”‚   └── contrastive_markers_top.csv      # only when contrastive markers are enabled
└── reproducibility/
    β”œβ”€β”€ commands.sh
    β”œβ”€β”€ environment.yml
    └── checksums.sha256

Dependencies

Required:

  • scanpy >= 1.10
  • anndata >= 0.10
  • scipy
  • numpy, pandas, matplotlib, leidenalg, python-igraph

Optional:

  • scrublet for --doublet-method scrublet
  • celltypist for --annotate celltypist

Out of scope:

  • scvi-tools / scANVI

Safety

  • Local-first: No patient data upload.
  • Disclaimer: Reports include the ClawBio medical disclaimer.
  • Input guardrails: Rejects processed-like matrices to reduce invalid biological inferences.
  • Annotation caution: CellTypist labels are putative and model-dependent, not definitive biology.
  • Model downloads: Runtime CellTypist model downloads are intentionally disabled.
  • Reproducibility: Writes command/environment/checksum bundle.

Integration with Bio Orchestrator

Trigger conditions:

  • File extension .h5ad, .mtx, or .mtx.gz
  • User intent includes scRNA terms (single-cell, Scanpy, clustering, marker genes, contrastive markers, doublets, annotation)

Current limitations:

  • Raw-count .h5ad and 10x Matrix Market only
  • CellTypist support is human-model focused and requires a locally installed model
  • Multi-group pairwise contrastive markers and within-cluster contrastive markers are future work

Status

MVP implemented -- supports .h5ad input and --demo PBMC3k-first demo data (fallback to synthetic on failure), plus opt-in Scrublet doublet detection, opt-in local CellTypist annotation, opt-in latent downstream mode from integrated.h5ad, and opt-in two-group contrastive markers with volcano plots.

Citations

Keep looking

Skills are one crate of 328,083. Ordering is by how many stacks a row turns up in, so the top of any crate is what has actually been picked rather than what has the most stars.