agentsclimarketplace

Chip atlas target genes

Skill BioTender-max/awesome-bio-agent-skills/skills/omicsclaw/chip-atlas-target-genes

A curated collection of AI agent skills for biomedical research, covering genomics, proteomics, single-cell analysis, clinical AI, and protein design.From the repository description

Install
npx -y skills add BioTender-max/awesome-bio-agent-skills --skill chip-atlas-target-genes

Assembled from the repository path, not quoted from the project. Check it against their README if it does not work.

One thing to look at

  • no licenseNo license file was found in the repository. Code published without one is not open source by default, so using it at work is a question for whoever answers licensing questions where you are.

SKILL.md

13.5 KB, ~3.3k tokens by cl100k_base, as published. Nobody here has run it

ChIP-Atlas Target Genes

Find target genes for any transcription factor using pre-computed ChIP-Atlas public ChIP-seq data.

When to Use This Skill

Use ChIP-Atlas target genes when you need to:

  • Identify target genes of a specific TF from all public ChIP-seq experiments
  • Rank potential targets by MACS2 binding score across hundreds of experiments
  • Compare TF binding across cell types using per-experiment binding scores
  • Validate known TF-target relationships with independent ChIP-seq evidence
  • Cross-reference with STRING protein interaction data for high-confidence targets

Don't use for:

  • Finding which TFs bind near your genes (use chip-atlas-peak-enrichment instead)
  • Histone mark targets (only non-histone antigens/TFs available)
  • Offline analysis (requires internet for data download)
  • Raw ChIP-seq analysis from FASTQ/BAM files

Key Concept: Downloads pre-computed TSV files containing MACS2 binding scores for every gene, across all public ChIP-seq experiments for the specified protein. Genes are ranked by average binding score. No API job submission needed β€” data is served as static files. STRING protein interaction scores are pre-embedded columns in the ChIP-Atlas TSV β€” no separate STRING API query is performed.

Installation

SoftwareVersionLicenseCommercial UseInstallation
pandas>=1.3BSD-3-ClausePermittedpip install pandas
requests>=2.25Apache-2.0Permittedpip install requests
numpy>=1.20BSD-3-ClausePermittedpip install numpy
plotnine>=0.12MITPermittedpip install plotnine
plotnine_prism>=0.1MITPermittedpip install plotnine_prism
matplotlib>=3.4PSF-basedPermittedpip install matplotlib
seaborn>=0.12BSD-3-ClausePermittedpip install seaborn
pip install pandas requests numpy plotnine plotnine_prism matplotlib seaborn

System requirements: Internet connection (downloads from ChIP-Atlas data server)

Inputs

Query parameters:

  • Protein/TF name: Case-sensitive gene symbol (e.g., "TP53", "CTCF", "MYC")
  • Genome: hg38 (default), hg19, mm10, mm9, rn6, dm6, dm3, ce11, ce10, sacCer3
  • Distance from TSS: 1kb, 5kb (default), or 10kb

Optional filters:

  • min_score: Minimum average MACS2 binding score (default: 0)
  • top_n: Keep top N genes (default: 500)
  • cell_types: List of cell types to subset (recalculates averages)
  • min_string_score: Minimum STRING interaction score
  • min_binding_rate: Minimum fraction of experiments with binding

Outputs

Analysis objects (Pickle):

  • analysis_object.pkl - Complete results for downstream use
    • Load with: import pickle; obj = pickle.load(open('analysis_object.pkl', 'rb'))
    • Contains: target_genes, experiment_data, cell_types, protein, parameters, metadata

Results (CSV):

  • target_genes_all.csv - All target genes (gene, avg_score, string_score, binding_rate, num_bound, max_score, colocated_group)
  • target_genes_top50.csv - Top 50 by average binding score
  • target_genes_with_string.csv - Genes with STRING interaction evidence
  • experiment_scores_top50.csv - Wide-format per-experiment scores for top 50

Visualizations (PNG + SVG, plotnine with Prism theme):

  • target_genes_top_targets.png/.svg - Top target genes barplot
  • target_genes_score_distribution.png/.svg - Binding score distribution histogram
  • target_genes_heatmap.png/.svg - Binding heatmap (top genes Γ— experiments)
  • target_genes_string_vs_binding.png/.svg - STRING vs binding scatter

Reports:

  • summary_report.md - Human-readable analysis summary

Clarification Questions

🚨 ALWAYS ask Question 1 FIRST. Do not ask about species, genome, or analysis parameters before the user has answered Question 1.

1. Query (ASK THIS FIRST):

  • Which transcription factor / protein do you want to find target genes for?
  • Protein name is case-sensitive (e.g., "TP53" not "tp53")
  • Or use example data? tp53 (large, ~16K genes), e2f1 (cell cycle), myc (moderate)

🚨 IF EXAMPLE DATA SELECTED: All parameters are pre-defined (human hg38, ±5kb TSS, all cell types, no score threshold). DO NOT ask question 2. Proceed directly to Step 1.

Question 2 is ONLY for users providing their own query:

2. Analysis parameters:

  • Species/genome? Human hg38 (default), hg19, mouse mm10/mm9, rat rn6, fly, worm, yeast
  • Distance from TSS? Β±5kb (default), Β±1kb (proximal only), Β±10kb (distal included)
  • Cell type filter? All cell types (default), or specific types to get cell-type-specific rankings
  • Score threshold? No minimum (default), or set min_score to focus on strong targets

Standard Workflow

Note: Run from the OmicsClaw root directory and add the workflow scripts to sys.path:

import sys; import os; sys.path.insert(0, os.path.abspath('knowledge_base/scripts/chip-atlas-target-genes'))

🚨 MANDATORY: USE SCRIPTS EXACTLY AS SHOWN - DO NOT WRITE INLINE CODE 🚨

Step 1 - Load query:

# Option 1: Example query
from load_example_query import load_example_query
query = load_example_query("tp53")

# Option 2: Your own protein
# from load_user_query import load_user_query
# query = load_user_query("TP53", genome="hg38", distance=5)

VERIFICATION: "βœ“ Query loaded: TP53 target genes (hg38, Β±5kb)"

Step 2 - Run target genes analysis:

from run_target_genes_workflow import run_target_genes_workflow

results = run_target_genes_workflow(
    protein=query['protein'],
    genome=query['genome'],
    distance=query['distance'],
    top_n=500,
    output_dir="target_genes_results"
)

DO NOT write inline download/parsing code. Just use the script.

VERIFICATION: "βœ“ Target genes analysis completed successfully!"

Step 3 - Generate visualizations:

from generate_all_plots import generate_all_plots
generate_all_plots(results, output_dir="target_genes_results", top_n=25)

DO NOT write inline plotting code. The script handles PNG + SVG with graceful fallback.

VERIFICATION: "βœ“ All visualizations generated successfully!"

Step 4 - Export results:

from export_all import export_all
export_all(results, output_dir="target_genes_results")

DO NOT write custom export code. Use export_all().

VERIFICATION: "=== Export Complete ==="

⚠️ CRITICAL - DO NOT:

  • ❌ Write inline download/parsing code β†’ STOP: Use run_target_genes_workflow()
  • ❌ Write inline plotting code β†’ STOP: Use generate_all_plots()
  • ❌ Write custom export code β†’ STOP: Use export_all()

⚠️ IF SCRIPTS FAIL - Script Failure Hierarchy:

  1. Fix and Retry (90%) - Install missing package, check internet, re-run
  2. Modify Script (5%) - Edit the script file itself, document changes
  3. Use as Reference (4%) - Read script, adapt approach, cite source
  4. Write from Scratch (1%) - Only if genuinely impossible, explain why

NEVER skip directly to writing inline code without trying the script first.

Common Issues

ErrorCauseSolution
HTTP 404 for proteinInvalid or unavailable antigenCheck case sensitivity ("TP53" not "tp53"). Histone marks not available. See references/target_genes_data_format.md.
Download timeoutLarge file or slow connectionTP53 is ~13MB; allow up to 2 minutes. Try smaller TF first (e.g., MYC).
Memory error on large fileVery wide TSV (100s of columns)Use top_n parameter to limit genes. Cell-type filter reduces columns.
No STRING data (all zeros)Protein not in STRING databaseNormal for less-studied TFs. Binding scores still valid without STRING.
Empty results after filteringFilters too strictLower min_score, remove cell_type filter, increase top_n.
SVG export errorMissing optional dependencyNormal - generate_all_plots() handles fallback. PNG always created.

Interpretation Guidelines

Average Binding Score (MACS2): βˆ’10 Γ— log10(Q-value). Higher = stronger binding evidence.

  • β‰₯500: Very strong binding (Q ≀ 1e-50) β€” high-confidence direct target
  • 100-500: Strong binding β€” likely direct target
  • 50-100: Moderate binding β€” possible target, may be cell-type-specific
  • <50: Weak binding β€” marginal evidence

Note: The Q-value thresholds above apply to individual experiment scores. Average scores include zeros from non-binding experiments, so an average of 500 reflects a consensus level β€” not that every experiment shows Q ≀ 1e-50.

Binding Rate: Fraction of experiments with any binding, shown as % with n/N count (e.g., "66.3% (260/392)"). >50% = consistent across cell types; <10% = cell-type-specific.

STRING Score: Independent evidence of regulatory interaction. >400 = medium confidence; >700 = high confidence. Genes with BOTH high binding + high STRING = highest-confidence targets. STRING score of 0 does NOT mean "not a target" β€” even well-characterized targets (e.g., BBC3/PUMA for TP53) can have STRING score 0 due to gaps in STRING coverage.

Caveats:

  • Averaging includes zeros: Average scores are computed across ALL experiments (including those with score 0). Use max_score and binding_rate for complementary views.
  • Cell-type bias: Experiments are unevenly distributed β€” a few well-studied cell lines dominate. See the "Experiment Composition" section in summary_report.md for exact distribution.
  • Co-located genes: Some genes share identical binding scores because they sit at the same genomic locus within the TSS window. The colocated_group column in the CSV flags these genes. For pathway enrichment, consider collapsing co-located groups to avoid double-counting loci. See summary_report.md Caveats for exact counts.
  • External annotations: ChIP-Atlas provides binding data only. Any biological role descriptions in the agent's summary are from general knowledge, not from this analysis output. Always cite the actual data columns (avg_score, binding_rate, string_score) when reporting results.

Reporting Results

🚨 CRITICAL: Follow these rules when presenting results to the user.

  1. Rankings MUST come from the data files. Read summary_report.md or target_genes_all.csv for exact gene names, ranks, and scores. NEVER construct ranking tables from general biological knowledge β€” even if a gene is a well-known target, its rank must match the data.
  2. Do NOT substitute biologically famous genes into top-N lists where the data ranks them lower. If a well-known target is not in the top 10, say so explicitly (e.g., "BAX, a well-characterized target, ranks #17 with avg_score 368.8").
  3. Use exact values from the CSV/report (1 decimal place for scores, 1 decimal for percentages). Do not round to integers.
  4. Cite the data source as: ChIP-Atlas (Zou et al., 2024) with the DOI from the References section.
  5. Mention co-located gene groups if the summary report flags them β€” they affect the effective number of independent targets.
  6. Label biological annotations: When describing gene functions or pathway roles, explicitly note these come from general knowledge (e.g., "CDKN1A β€” known cell cycle arrest effector β€” ranks #1"), not from ChIP-Atlas output.

Suggested Next Steps

  1. Run peak enrichment with top target genes to find co-regulatory factors (chip-atlas-peak-enrichment)
  2. Cell-type-specific analysis β€” re-run with cell_types filter matching your experimental system
  3. Gene regulatory network construction using top targets as nodes
  4. Functional enrichment of top target genes (GO, pathway analysis)

Related Skills

References

What ships with it

Read from the repository

Just SKILL.md. No reference files, no scripts.

Keep looking

Skills are one crate of 325,949. Ordering is by how many stacks a row turns up in, so the top of any crate is what has actually been picked rather than what has the most stars.