agentsclimarketplace

Bio applied gene regulatory networks

Skill Pavel-Kravchenko/Bioinformatics/Skills/bio-applied-gene-regulatory-networks

208 bioinformatics skills for Claude Code — NGS, single-cell, metagenomics, structural biology, algorithms, AI for science

Install
npx -y skills add Pavel-Kravchenko/Bioinformatics --skill bio-applied-gene-regulatory-networks

Assembled from the repository path, not quoted from the project. Check it against their README if it does not work.

2 things to look at

  • no licenseNo license file was found in the repository. Code published without one is not open source by default, so using it at work is a question for whoever answers licensing questions where you are.
  • 3 stars3 stars. Stars are a popularity signal and not a quality one, but at this level it is likely that nobody has read this closely except its author, and you would be relying on your own review.

What its author says it does

Copied from the file, not written here

Infer TF-target regulatory networks via correlation, ARACNE mutual information, and GENIE3 random-forest importance; find feed-forward loops; validate against TRRUST. Use when building a GRN or asked about GENIE3, ARACNE, or regulons.

SKILL.md

10.3 KB, as published. Nobody here has run it

Gene Regulatory Network Inference

When to Use

  • Inferring which transcription factors (TFs) regulate which target genes from bulk or single-cell expression data.
  • Comparing correlation, mutual-information (ARACNE), and random-forest (GENIE3/GRNBoost2) GRN inference methods.
  • Detecting network motifs (feed-forward loops, autoregulation) in an inferred or curated regulatory network.
  • Validating predicted TF-target edges against curated databases (TRRUST, DoRothEA, ChEA3).
  • Deciding between a fast correlation screen and a slower, more accurate GENIE3/SCENIC run.

Version Compatibility

scikit-learn >= 1.3, scipy >= 1.11, NetworkX >= 3.2, pandas >= 2.0, Python >= 3.10. For production-scale GRN inference use arboreto (GRNBoost2/GENIE3, Python) or Bioconductor GENIE3/SCENIC (R) — this skill's genie3() is a from-scratch reference implementation for understanding the algorithm, not a substitute for those.

Prerequisites

  • pip install scikit-learn scipy pandas networkx matplotlib
  • Expression matrix already normalized/log-transformed (see bio-differential-expression-deseq2-basics or bio-single-cell-preprocessing).
  • A defined list of candidate TFs (e.g. from AnimalTFDB or a curated TF panel) — GRN inference only makes sense TF-vs-target, not all-vs-all.

GRN vs PPI Networks

AspectPPI NetworkGRN
EdgesPhysical protein bindingTranscriptional regulation
DirectionUndirectedDirected (TF -> target)
Data sourceY2H, co-IPExpression + ChIP-seq
DatabasesSTRING, BioGRIDTRRUST, ChEA, DoRothEA

Method Comparison

MethodCapturesProsCons
Pearson/Spearman correlationLinear co-variationFast, interpretableSymmetric; many indirect edges
ARACNE (mutual information + DPI)Non-linear relationshipsPrunes indirect edgesStill symmetric (no direction)
GENIE3 / GRNBoost2 (random forest / gradient boosting)Complex, combinatorial dependenciesDirectional importance scores; best BEELINE benchmark AUPRCSlower; no sign (activation vs repression)
SCENIC(+)Cell-type-specific regulationIntegrates chromatin accessibility + expressionComplex setup, needs scATAC-seq

Choosing a method: start with correlation for a quick screen. Move to GENIE3/GRNBoost2 (via arboreto) when you need directionality and non-linear signal. Use SCENIC+ for single-cell multiome data with matched ATAC.

Correlation-Based Inference

Goal: get a fast first-pass list of candidate TF-target edges. Approach: Pearson correlation between each TF and each target, thresholded on |r| and p-value.

import numpy as np
import pandas as pd
from scipy.stats import pearsonr

def infer_grn_correlation(tf_expr, target_expr, tf_names, target_names,
                           r_thresh=0.25, p_thresh=0.01):
    """Infer TF-target edges from Pearson correlation.

    tf_expr, target_expr: (n_samples, n_tfs)/(n_samples, n_targets) arrays.
    Returns a DataFrame of edges with |r| > r_thresh and p < p_thresh.
    """
    edges = []
    for i, tf in enumerate(tf_names):
        for j, tgt in enumerate(target_names):
            r, p = pearsonr(tf_expr[:, i], target_expr[:, j])
            if abs(r) > r_thresh and p < p_thresh:
                edges.append({"TF": tf, "Target": tgt, "r": r, "p": p})
    return pd.DataFrame(edges)

Mutual Information / ARACNE

Goal: capture non-linear TF-target relationships and prune indirect edges. Approach: compute mutual information (MI) per TF-target pair, then apply the Data Processing Inequality (DPI): for a triangle A-B-C, drop the weakest edge.

import numpy as np
from sklearn.feature_selection import mutual_info_regression

def compute_mi_matrix(tf_expr, target_expr, random_state=42):
    """Mutual information between each TF (columns of tf_expr) and every target."""
    n_tfs = tf_expr.shape[1]
    mi = np.zeros((n_tfs, target_expr.shape[1]))
    for i in range(n_tfs):
        mi[i] = mutual_info_regression(
            tf_expr[:, i:i + 1], target_expr,
            discrete_features=False, random_state=random_state,
        )
    return mi

def aracne_dpi_prune(mi_tf_tf, tolerance=0.0):
    """Data Processing Inequality pruning over a square TF-TF MI matrix.

    For every triangle (i, j, k), drop the smallest edge if it is
    smaller than both others by more than `tolerance` (indirect edge).
    Returns a boolean keep-mask matching mi_tf_tf's shape.
    """
    n = mi_tf_tf.shape[0]
    keep = np.ones_like(mi_tf_tf, dtype=bool)
    for i in range(n):
        for j in range(n):
            if i == j:
                continue
            for k in range(n):
                if k in (i, j):
                    continue
                triangle = sorted([mi_tf_tf[i, j], mi_tf_tf[i, k], mi_tf_tf[j, k]])
                weakest = triangle[0]
                if weakest == mi_tf_tf[i, j] and weakest < triangle[1] - tolerance:
                    keep[i, j] = False
    return keep

GENIE3: Random Forest Importance

Goal: rank TF-target edges by how well each TF predicts each target, capturing combinatorial and non-linear regulation. Approach: for every target gene, fit a random forest on all TFs as predictors; the feature importances are the TF -> target edge weights.

import numpy as np
from sklearn.ensemble import RandomForestRegressor

def genie3(expr_matrix, tf_indices, target_indices, n_estimators=100, random_state=42):
    """Simplified GENIE3: rank TF importance for each target via random forest.

    expr_matrix: (n_samples, n_genes) full expression matrix.
    tf_indices, target_indices: column indices into expr_matrix for TFs/targets.
    Returns an (n_tfs, n_targets) importance matrix (TF rows, target columns).
    """
    X = expr_matrix[:, tf_indices]
    importance = np.zeros((len(tf_indices), len(target_indices)))
    for j, tgt_idx in enumerate(target_indices):
        y = expr_matrix[:, tgt_idx]
        rf = RandomForestRegressor(
            n_estimators=n_estimators, random_state=random_state,
            n_jobs=-1, max_features="sqrt",
        )
        rf.fit(X, y)
        importance[:, j] = rf.feature_importances_
    return importance

For genome-scale data use arboreto.algo.grnboost2(expression_data=df, tf_names=tfs) (gradient boosting, much faster than this reference loop).

Network Motifs: Feed-Forward Loops

Goal: find A -> B -> C, A -> C triads (the most common GRN motif) in a directed edge list. Approach: build a networkx.DiGraph from thresholded edges and enumerate triangles that satisfy the FFL pattern.

import networkx as nx

def find_feed_forward_loops(G, tf_names):
    """Enumerate feed-forward loops A->B->C, A->C where A,B are TFs.

    G: networkx.DiGraph of TF/target edges.
    Returns a list of (A, B, C) triples.
    """
    ffls = []
    for a in tf_names:
        if a not in G:
            continue
        for b in G.successors(a):
            if b not in tf_names or b not in G:
                continue
            for c in G.successors(b):
                if c not in tf_names and G.has_edge(a, c):
                    ffls.append((a, b, c))
    return ffls

Evaluation Against a Reference Network

def evaluate_grn(predicted_binary, true_binary):
    """Precision/recall/F1 of a predicted binary TF-target adjacency vs. ground truth."""
    tp = (predicted_binary & true_binary).sum()
    fp = (predicted_binary & ~true_binary).sum()
    fn = (~predicted_binary & true_binary).sum()
    precision = tp / (tp + fp) if (tp + fp) > 0 else 0
    recall = tp / (tp + fn) if (tp + fn) > 0 else 0
    f1 = 2 * precision * recall / (precision + recall) if (precision + recall) > 0 else 0
    return {"precision": precision, "recall": recall, "f1": f1}

Use TRRUST (8,444 curated human TF-target pairs, grnpedia.org/trrust), DoRothEA, or ChEA3 as the true_binary reference. Published BEELINE benchmark numbers (Pratapa et al. 2020, Nat Methods): GENIE3 AUPRC ~0.15, GRNBoost2 ~0.16, ARACNE ~0.09, Pearson ~0.07 — all methods perform modestly; use them to rank candidates, not as ground truth.

Key Concepts

  • Regulon: the set of genes regulated by one TF.
  • Feed-forward loop (FFL): A->B->C and A->C — most common 3-node motif; coherent FFLs filter noise/delay response, incoherent FFLs produce pulses/adaptation.
  • VIPER: infers TF activity (not just mRNA expression) from a regulon's target expression pattern — more informative than raw TF expression, especially when the TF itself is post-translationally regulated.

Pitfalls

  • Correlation != regulation: two genes co-expressed due to a shared upstream regulator will appear as a direct edge.
  • Symmetric methods (correlation, MI) cannot infer direction: need orthogonal evidence (ChIP-seq, perturbation, known TF/target roles) to resolve A->B vs B->A.
  • High false-positive rate: GRN inference on bulk RNA-seq is noisy; always validate top edges against TRRUST/DoRothEA/ChIP-seq before acting on them.
  • Sample size: reliable inference needs n >> number of TFs; partial correlation and Graphical Lasso fail outright when n < p.
  • Batch effects: check for batch confounding before interpreting co-expression as regulation — batch is a classic shared upstream "regulator."
  • Multiple testing: apply FDR correction when testing thousands of TF-target pairs (correlation/MI approaches).
  • GENIE3 has no sign: importance scores don't say activation vs. repression — check correlation sign or use signed resources (CollecTRI) if direction of effect matters.

See Also

  • bio-applied-ppi-networks — undirected protein-protein interaction networks (STRING/BioGRID), centrality, community detection.
  • bio-applied-network-modules — Louvain/Leiden module detection and WGCNA co-expression modules, complementary to directed GRN edges.
  • bio-applied-regulatory-analysis — promoter-level analysis (PWM/PFM, TFBS scanning) to corroborate inferred TF-target edges.
  • arboreto — production GRNBoost2/GENIE3 implementation for genome-scale inference.

Keep looking

Skills are one crate of 328,083. Ordering is by how many stacks a row turns up in, so the top of any crate is what has actually been picked rather than what has the most stars.