agentsclimarketplace

Genomic foundation models

Skill Pavel-Kravchenko/Bioinformatics/Skills/genomic-foundation-models

208 bioinformatics skills for Claude Code — NGS, single-cell, metagenomics, structural biology, algorithms, AI for science

Install
npx -y skills add Pavel-Kravchenko/Bioinformatics --skill genomic-foundation-models

Assembled from the repository path, not quoted from the project. Check it against their README if it does not work.

2 things to look at

  • no licenseNo license file was found in the repository. Code published without one is not open source by default, so using it at work is a question for whoever answers licensing questions where you are.
  • 3 stars3 stars. Stars are a popularity signal and not a quality one, but at this level it is likely that nobody has read this closely except its author, and you would be relying on your own review.

What its author says it does

Copied from the file, not written here

Choose/run DNA foundation models (Nucleotide Transformer, HyenaDNA, Evo, Enformer, Borzoi) via transformers: embed sequences, fine-tune, score variants with Enformer ISM. Use for genomic LLM choice or variant scoring.

SKILL.md

8.5 KB, as published. Nobody here has run it

genomic-foundation-models

When to Use

  • Deciding which DNA/RNA foundation model fits a task: embeddings vs. regulatory-track prediction vs. splice scoring vs. sequence design.
  • Extracting sequence embeddings (Nucleotide Transformer, DNABERT-2, HyenaDNA) for downstream classifiers (promoter/enhancer/splice-site).
  • Fine-tuning a genomic LLM for a binary/multi-class sequence classification task.
  • Running Enformer for multi-track regulatory prediction or in-silico mutagenesis (ISM) variant-effect scoring.
  • Routing a coding variant onward to a protein structure model (AlphaFold2/3, RoseTTAFold) once sequence-level scoring is done.

Version Compatibility

transformers >=4.40, torch >=2.1, enformer-pytorch >=0.8, Python >=3.10. Nucleotide Transformer checkpoints live under InstaDeepAI/ on the HF Hub; Enformer via the community PyTorch port EleutherAI/enformer-official-rough (DeepMind's original is TensorFlow/Sonnet, on GitHub). HyenaDNA and Evo ship their own HF-compatible repos (LongSafari/hyenadna-*, togethercomputer/evo-1-*).

Prerequisites

pip install transformers torch enformer-pytorch accelerate

GPU guidance: T4 (16 GB) is enough for NT-500M or DNABERT-2; NT-2.5B and Evo-7B need an A100-class GPU (~40 GB) or model.half() + gradient checkpointing. Enformer/SpliceAI inference on short windows runs fine on CPU. Familiarity with the HF AutoModel/Trainer API is assumed.

Model Landscape

ModelContextBest For
DNABERT-2~3 kbp (BPE tokens)Short-window classification (promoter/enhancer/splice)
Nucleotide Transformerkb-scale (6-mer tokens)Transfer learning across species
HyenaDNAup to 1M bpDistal regulatory context
Evo~100 kb+Prokaryotic sequence design (autoregressive)
Enformer196,608 bpMulti-track regulatory signal (5,313 tracks, 128 bp bins)
Borzoi524 kbRNA-seq coverage at 32 bp bins
SpliceAIlocal windowClinical splice variant delta scores (DS_AG/AL/DG/DL)
AlphaGenomeup to 1M bpUnified variant effect (expression + splicing + chromatin + contacts)

Routing: general embeddings -> DNABERT-2 / NT / HyenaDNA. Regulatory track prediction -> Enformer / Borzoi / AlphaGenome. Splice-specific scoring -> SpliceAI. Sequence design -> Evo. Coding variants with a structural hypothesis -> hand off to AlphaFold2/3 or RoseTTAFold (see bio-structural-biology-modern-structure-prediction).

1. Extract Sequence Embeddings

Goal: turn raw DNA strings into fixed-length vectors for a downstream classifier or similarity search. Approach: tokenize with the model's native scheme (NT uses overlapping 6-mers), forward-pass with no grad, mean-pool over the token dimension.

from transformers import AutoTokenizer, AutoModel
import torch


def embed_sequences(sequences: list[str], model_name: str = "InstaDeepAI/nucleotide-transformer-v2-500m-multi-species",
                     max_length: int = 512) -> "torch.Tensor":
    """Mean-pooled per-sequence embeddings from a genomic foundation model.

    Returns a (n_sequences, hidden_dim) tensor, e.g. (n, 1024) for NT-500M.
    """
    tokenizer = AutoTokenizer.from_pretrained(model_name)
    model = AutoModel.from_pretrained(model_name).eval()

    inputs = tokenizer(sequences, return_tensors="pt", padding=True,
                       truncation=True, max_length=max_length)
    with torch.no_grad():
        outputs = model(**inputs)

    # Per-residue (token-level) embeddings are outputs.last_hidden_state
    # (n_sequences, seq_len_tokens, hidden_dim) -- use that instead for per-base tasks.
    return outputs.last_hidden_state.mean(dim=1)

2. Fine-Tune for Sequence Classification

Goal: adapt a pretrained genomic LLM to a labeled task (e.g. promoter vs. non-promoter, splice-site usage). Approach: swap in a classification head (AutoModelForSequenceClassification) and fine-tune with the HF Trainer; freeze the backbone first if the label set is small (<1k examples) to avoid overfitting.

from transformers import AutoModelForSequenceClassification, AutoTokenizer, TrainingArguments, Trainer


def build_classifier_trainer(model_name: str, tokenized_dataset, num_labels: int = 2,
                              output_dir: str = "./finetuned") -> Trainer:
    """Configure a Trainer for fine-tuning a genomic LLM classification head.

    tokenized_dataset must already contain input_ids/attention_mask/labels
    (produced by AutoTokenizer(...) over your sequences).
    """
    model = AutoModelForSequenceClassification.from_pretrained(model_name, num_labels=num_labels)
    args = TrainingArguments(
        output_dir=output_dir,
        per_device_train_batch_size=8,
        num_train_epochs=3,
        learning_rate=1e-4,
        eval_strategy="epoch",
        fp16=True,
    )
    return Trainer(model=model, args=args, train_dataset=tokenized_dataset)

3. Enformer: Regulatory Tracks + In-Silico Mutagenesis

Goal: predict 5,313 regulatory tracks (CAGE, DNase, ChIP-seq) from a 196,608 bp window and score a single-base variant's effect on one track. Approach: one-hot encode the exact-length window, run reference vs. mutated sequence, take the delta at the track/bin of interest (bin 448 is the window center at 128 bp resolution).

from enformer_pytorch import from_pretrained
import torch


def one_hot_encode(seq: str) -> torch.Tensor:
    """One-hot encode an ACGT string to shape (1, len(seq), 4); ambiguous bases -> 0.25 each."""
    mapping = {"A": [1, 0, 0, 0], "C": [0, 1, 0, 0], "G": [0, 0, 1, 0], "T": [0, 0, 0, 1]}
    return torch.tensor([[mapping.get(b, [0.25] * 4) for b in seq.upper()]], dtype=torch.float32)


def ism_score(model, ref_one_hot: torch.Tensor, center_pos: int = 196608 // 2,
              target_track: int = 4799) -> list[float]:
    """In-silico mutagenesis: delta of a target track for each possible base at center_pos.

    Returns [delta_A, delta_C, delta_G, delta_T] relative to the reference prediction.
    """
    with torch.no_grad():
        baseline = model(ref_one_hot)["human"][0, 448, target_track].item()
        deltas = []
        for alt_idx in range(4):
            mut = ref_one_hot.clone()
            mut[0, center_pos, :] = 0
            mut[0, center_pos, alt_idx] = 1
            pred = model(mut)["human"][0, 448, target_track].item()
            deltas.append(pred - baseline)
    return deltas


model = from_pretrained("EleutherAI/enformer-official-rough", target_length=-1).eval()
#NB: window must be exactly 196,608 bp; pad/extract from a reference genome with pysam.FastaFile.

Pitfalls

  • Context length: NT 6-mers with 512 tokens covers only ~3 kbp; use HyenaDNA/Borzoi/AlphaGenome for distal regulatory interactions beyond that.
  • Tokenization mismatch: DNABERT-2 uses BPE, NT uses overlapping 6-mers — embeddings from different models are not directly comparable.
  • Enformer input: requires exactly 196,608 bp; a mismatched window silently reshapes or errors depending on the port.
  • GPU memory: NT-2.5B/Evo-7B need ~40 GB; use model.half() or a smaller checkpoint before assuming OOM means "add more data."
  • Species/checkpoint fit: NT-multi-species (850 species) vs. NT-human — pick the checkpoint matching your organism.
  • Gain vs. loss splice deltas: SpliceAI's DS_AG/DS_DG (new site) are harder to interpret than DS_AL/DS_DL (weakened known site) — don't treat all four scores as equivalent.
  • Modality confusion: Enformer/Borzoi/AlphaGenome are regulatory-genomics models; AlphaFold/RoseTTAFold predict protein structure — a high ISM delta does not imply anything about protein folding.
  • eQTL validation: correlating model output with GTEx-style eQTLs only tests population-level directionality, not per-individual expression prediction (Sasse et al. 2023).

See Also

  • ai-science-genomic-llms — deeper NT/HyenaDNA/Evo embedding and k-mer tokenization walkthrough.
  • ai-science-epigenomic-sequence-models — Borzoi vs. Epiformer vs. AlphaGenome selection for RNA-seq/accessibility.
  • ai-science-variant-to-structure-models — triage variant scores into AlphaFold2/3 or RoseTTAFold structural follow-up.
  • bio-structural-biology-modern-structure-prediction — run AlphaFold2/3/RoseTTAFold once a coding variant needs structural interpretation.

Keep looking

Skills are one crate of 328,083. Ordering is by how many stacks a row turns up in, so the top of any crate is what has actually been picked rather than what has the most stars.