agentsclimarketplace

Mendelian randomization twosamplemr

Skill BioTender-max/awesome-bio-agent-skills/skills/omicsclaw/mendelian-randomization-twosamplemr

A curated collection of AI agent skills for biomedical research, covering genomics, proteomics, single-cell analysis, clinical AI, and protein design.From the repository description

Install
npx -y skills add BioTender-max/awesome-bio-agent-skills --skill mendelian-randomization-twosamplemr

Assembled from the repository path, not quoted from the project. Check it against their README if it does not work.

One thing to look at

  • no licenseNo license file was found in the repository. Code published without one is not open source by default, so using it at work is a question for whoever answers licensing questions where you are.

SKILL.md

10.6 KB, ~2.6k tokens by cl100k_base, as published. Nobody here has run it

Two-Sample Mendelian Randomization

When to Use This Skill

  • You have GWAS summary statistics for an exposure and outcome trait
  • You want to test causal direction between two traits (not just correlation)
  • You need to assess whether an observed association is likely causal or confounded
  • You want to use genetic variants as instrumental variables (natural experiment)
  • You have OpenGWAS trait IDs or your own GWAS summary statistics files

Not suitable for: One-sample MR (individual-level data), non-linear MR, multivariable MR with >2 exposures

Installation

install.packages(c("remotes", "ggplot2", "ggprism", "dplyr", "rmarkdown"))
remotes::install_github("MRCIEU/TwoSampleMR")
# For PDF report generation (optional but recommended):
# install.packages("tinytex"); tinytex::install_tinytex()
# For MR-PRESSO outlier detection (optional but recommended):
# remotes::install_github("rondolab/MR-PRESSO")
SoftwareVersionLicenseCommercial Use
TwoSampleMR≥0.5.6GPL-3✅ Permitted
ieugwasr≥0.2.1MIT✅ Permitted
ggplot2≥3.4.0MIT✅ Permitted
ggprism≥1.0.3GPL (≥3)✅ Permitted
dplyr≥1.1.0MIT✅ Permitted
rmarkdown≥2.20GPL-3✅ Permitted

Inputs

Option A — OpenGWAS IDs (recommended):

  • Exposure ID (e.g., "ieu-a-300" for LDL cholesterol)
  • Outcome ID (e.g., "ieu-a-7" for coronary heart disease)
  • Browse available traits at: https://gwas.mrcieu.ac.uk/

Option B — User-provided files (CSV/TSV):

  • Exposure GWAS summary statistics
  • Outcome GWAS summary statistics
Required ColumnDescriptionExample
SNPrsIDrs1234567
betaEffect estimate0.05
seStandard error0.01
pvalP-value5e-10
effect_alleleEffect alleleA
other_alleleOther alleleG
eafEffect allele frequency (optional)0.3

Outputs

Results (CSV):

  • mr_results.csv — MR estimates from all 4 methods (beta, SE, p-value, nSNP, F-statistics)
  • heterogeneity_results.csv — Cochran's Q test for instrument heterogeneity
  • pleiotropy_results.csv — MR-Egger intercept test for directional pleiotropy
  • directionality_results.csv — Steiger test confirming causal direction
  • harmonized_data.csv — SNP-level harmonized exposure-outcome data
  • single_snp_results.csv — Per-SNP Wald ratio estimates
  • leaveoneout_results.csv — Leave-one-out robustness estimates
  • MR-PRESSO outlier results (if heterogeneity significant and MRPRESSO installed)

Plots (PNG + SVG):

  • mr_scatter_plot — SNP-exposure vs SNP-outcome with method regression lines
  • mr_forest_plot — Individual + combined SNP effect estimates
  • mr_funnel_plot — Precision vs effect size (asymmetry = pleiotropy)
  • mr_leaveoneout_plot — Effect stability when removing each SNP

Report:

  • mr_report.pdf — Structured analysis report (Introduction, Methods, Results, Figures, Conclusions)

Analysis objects (RDS):

  • mr_object.rds — Complete analysis (results, sensitivity, harmonized data)
    • Load with: mr_obj <- readRDS("mr_results/mr_object.rds")

Clarification Questions

  1. Input Data (ASK THIS FIRST):

    • Do you have specific GWAS summary statistics or OpenGWAS trait IDs?
    • If files uploaded: Are these the exposure and outcome GWAS files?
    • Expected formats: CSV/TSV with SNP, beta, se, pval, effect_allele, other_allele
    • Or use example data? LDL Cholesterol → Coronary Heart Disease demo (drug-target validation example)
  2. Exposure and Outcome:

    • (If using example data) Pre-set: Exposure = LDL Cholesterol, Outcome = Coronary Heart Disease — no need to specify
    • (If using your own data) What is the exposure (potential cause)? What is the outcome (potential effect)?
  3. Parameters (defaults usually fine):

    • P-value threshold for instruments? (default: 5×10⁻⁸)
    • LD clumping r²? (default: 0.001)

Standard Workflow

Note: Run from the OmicsClaw root directory and add the workflow scripts to sys.path:

import sys; import os; sys.path.insert(0, os.path.abspath('knowledge_base/scripts/mendelian-randomization-twosamplemr'))

🚨 MANDATORY: USE SCRIPTS EXACTLY AS SHOWN — DO NOT WRITE INLINE CODE 🚨

Step 1 — Load and harmonize data:

source("scripts/load_data.R")
dat <- load_example_data()
# OR: dat <- load_from_opengwas("ieu-a-300", "ieu-a-7")
# OR: dat <- load_from_files("exposure.csv", "outcome.csv")

DO NOT write inline data loading or harmonization code. Use the functions above.

✅ VERIFICATION: You MUST see "✓ Data loaded and harmonized successfully!"

Step 2 — Run MR analysis:

source("scripts/run_mr_analysis.R")
mr_results <- run_mr(dat)
sensitivity <- run_sensitivity(dat, mr_results)

DO NOT write inline MR code. Just source the script and call the functions.

✅ VERIFICATION: You MUST see "✓ MR analysis completed successfully!" AND "✓ Sensitivity analyses completed successfully!"

Step 3 — Generate visualizations:

source("scripts/mr_plots.R")
generate_all_plots(mr_results, dat, sensitivity$singlesnp, sensitivity$leaveoneout, output_dir = "mr_results")

🚨 DO NOT write inline plotting code (ggsave, ggplot, etc.). Just use the function. 🚨

✅ VERIFICATION: You MUST see "✓ All MR plots generated successfully!"

Step 4 — Export results and generate report:

source("scripts/export_results.R")
export_all(mr_results, sensitivity, dat, output_dir = "mr_results")

DO NOT write custom export code. Use export_all(). It automatically generates the PDF report.

✅ VERIFICATION: You MUST see "✓ Report generated successfully!" AND "=== Export Complete ==="

❌ IF YOU DON'T SEE VERIFICATION MESSAGES: You wrote inline code. Stop and use the scripts.

⚠️ CRITICAL — DO NOT:

  • ❌ Write inline MR analysis code → STOP: Use run_mr() and run_sensitivity()
  • ❌ Write inline plotting code → STOP: Use generate_all_plots()
  • ❌ Write custom export code → STOP: Use export_all()
  • ❌ Write custom report code → STOP: Use generate_report()
  • ❌ Try to install system dependencies → Scripts handle package installation

⚠️ IF SCRIPTS FAIL — Script Failure Hierarchy:

  1. Fix and Retry (90%) — Install missing R package, re-run script
  2. Modify Script (5%) — Edit the script file, document changes
  3. Use as Reference (4%) — Read script, adapt approach, cite source
  4. Write from Scratch (1%) — Only if genuinely impossible, explain why

Common Issues

IssueCauseSolution
"No instruments found"No SNPs below p-value thresholdTry a less stringent threshold or check trait ID
LD clumping API failsOpenGWAS/IEU API temporarily downScript falls back to no clumping with warning; results may be affected by LD
"Only N SNPs retained"Allele harmonization removed most SNPsCheck if exposure/outcome are from same genome build
Steiger test failsSample sizes unavailable in metadataNormal for some datasets; other sensitivity tests still valid
SVG export errorMissing optional dependencyNormal — generate_all_plots() falls back to base R svg() automatically
OpenGWAS rate limitingToo many API requestsWait a few minutes and retry
PDF report failsLaTeX/tinytex not installedInstall with tinytex::install_tinytex() — report auto-falls back to HTML or base R PDF
Steiger R² warning for binary outcomeOutcome is case-control, not quantitativeUse get_r_from_lor() with prevalence to compute liability-scale R² before directionality test
MR-PRESSO not availableMRPRESSO package not installedremotes::install_github('rondolab/MR-PRESSO') — optional but recommended when heterogeneity is significant
"Cannot find function"Script not sourcedRun source("scripts/load_data.R") before calling functions

Interpreting Results

See references/interpretation-guide.md for detailed guidance.

Quick interpretation:

  • Concordant methods (IVW, Egger, WM, WMode agree on direction + significance) → stronger evidence
  • Any method non-significant or discordant → must be discussed explicitly, not dismissed
  • IVW significant + no heterogeneity + no pleiotropy → strongest evidence
  • Egger intercept p < 0.05 → directional pleiotropy may bias IVW
  • High heterogeneity (Q p < 0.05) → run MR-PRESSO, flag outlier instruments
  • Steiger direction incorrect → reverse causation concern (check binary outcome R² correction)
  • F-statistic < 10 → weak instrument bias toward the null

Suggested Next Steps

  • Multiple exposures? → Run bidirectional MR (swap exposure/outcome)
  • Pleiotropy detected? → Consider MR-PRESSO or multivariable MR
  • Significant result? → Replicate with independent GWAS datasets
  • Drug target validation? → Use cis-MR with variants near gene of interest
  • Pathway analysis? → Combine with functional enrichment skills

Related Skills

  • polygenic-risk-score — Polygenic risk score computation (LDpred2)
  • polygenic-risk-score-prs-catalog — PRS from pre-computed PGS Catalog weights
  • eqtl-colocalization-coloc — eQTL colocalization analysis (MR follow-up)

References

What ships with it

Read from the repository

Just SKILL.md. No reference files, no scripts.

Keep looking

Skills are one crate of 325,949. Ordering is by how many stacks a row turns up in, so the top of any crate is what has actually been picked rather than what has the most stars.