agentsclimarketplace

Wikipathways

Skill BioTender-max/awesome-bio-agent-skills/skills/bioskills/wikipathways

WikiPathways enrichment using clusterProfiler and rWikiPathways. Use when analyzing gene lists against community-curated open-source pathways. Performs over-representation analysis and GSEA for 30+ species.From its SKILL.md

Install
npx -y skills add BioTender-max/awesome-bio-agent-skills --skill wikipathways

Assembled from the repository path, not quoted from the project. Check it against their README if it does not work.

One thing to look at

  • no licenseNo license file was found in the repository. Code published without one is not open source by default, so using it at work is a question for whoever answers licensing questions where you are.

SKILL.md

7.5 KB, ~1.9k tokens by cl100k_base, as published. Nobody here has run it

Version Compatibility

Reference examples tested with: ReactomePA 1.46+, clusterProfiler 4.10+, rWikiPathways 1.24+

Before using code patterns, verify installed versions match. If versions differ:

  • R: packageVersion('<pkg>') then ?function_name to verify parameters

If code throws ImportError, AttributeError, or TypeError, introspect the installed package and adapt the example to match the actual API rather than retrying.

WikiPathways Enrichment

When to Use WikiPathways

WikiPathways is community-curated (wiki model), not expert or peer-reviewed like KEGG/Reactome. This means:

  • Strengths: disease-specific and drug-related pathways not found in KEGG/Reactome; fully open (CC0 license); newer pathways contributed by active researchers; 30+ species
  • Limitations: quality varies by pathway -- some are meticulously curated by domain experts, others may be incomplete or contributed by non-specialists
  • Best use: complement to KEGG/Reactome, not a standalone primary database. Run WikiPathways alongside KEGG or Reactome to catch pathways unique to the WikiPathways collection.

Check the "Last edited" date and contributor for specific pathways before relying on them for key conclusions.

Core Pattern - Over-Representation Analysis

Goal: Identify WikiPathways that are over-represented in a gene list.

Approach: Test for enrichment using enrichWP against community-curated open-source pathway definitions.

"Run pathway enrichment against WikiPathways" → Test whether genes from community-curated WikiPathways are over-represented among significant genes.

library(clusterProfiler)
library(org.Hs.eg.db)

wp_result <- enrichWP(
    gene = entrez_ids,         # Character vector of Entrez IDs
    organism = 'Homo sapiens', # Full species name
    pvalueCutoff = 0.05,
    pAdjustMethod = 'BH'
)

head(as.data.frame(wp_result))

Prepare Gene List

Goal: Extract significant Entrez gene IDs from DE results for WikiPathways enrichment.

Approach: Filter by significance thresholds and convert gene symbols to Entrez IDs with bitr.

de_results <- read.csv('de_results.csv')
sig_genes <- de_results[de_results$padj < 0.05 & abs(de_results$log2FoldChange) > 1, 'gene_symbol']

gene_ids <- bitr(sig_genes, fromType = 'SYMBOL', toType = 'ENTREZID', OrgDb = org.Hs.eg.db)
entrez_ids <- gene_ids$ENTREZID

GSEA on WikiPathways

Goal: Detect coordinated expression changes in WikiPathways using a ranked gene list.

Approach: Sort genes by fold change and run gseWP for rank-based enrichment testing.

# Create ranked gene list
gene_list <- de_results$log2FoldChange
names(gene_list) <- de_results$entrez_id
gene_list <- sort(gene_list, decreasing = TRUE)

gsea_wp <- gseWP(
    geneList = gene_list,
    organism = 'Homo sapiens',
    pvalueCutoff = 0.05,
    pAdjustMethod = 'BH'
)

head(as.data.frame(gsea_wp))

With Background Universe

all_genes <- de_results$entrez_id

wp_result <- enrichWP(
    gene = entrez_ids,
    universe = all_genes,
    organism = 'Homo sapiens',
    pvalueCutoff = 0.05
)

Make Results Readable

# Convert Entrez IDs to gene symbols
wp_readable <- setReadable(wp_result, OrgDb = org.Hs.eg.db, keyType = 'ENTREZID')

Visualization

Goal: Create summary plots of WikiPathways enrichment results.

Approach: Use enrichplot functions (dotplot, barplot, cnetplot, emapplot) on the enrichment result object.

library(enrichplot)

# Dot plot
dotplot(wp_result, showCategory = 15)

# Bar plot
barplot(wp_result, showCategory = 15)

# Gene-concept network
cnetplot(wp_readable, categorySize = 'pvalue')

# Enrichment map
wp_result <- pairwise_termsim(wp_result)
emapplot(wp_result)

Using rWikiPathways Directly

Goal: Query the WikiPathways database directly for pathway metadata, gene lists, and GMT files.

Approach: Use rWikiPathways API functions to list organisms, retrieve pathway info, and download gene set definitions.

library(rWikiPathways)

# List available organisms
listOrganisms()

# Get all pathways for an organism
human_pathways <- listPathways('Homo sapiens')

# Get pathway info
pathway_info <- getPathwayInfo('WP554')  # ACE Inhibitor Pathway

# Get genes in a pathway
pathway_genes <- getXrefList('WP554', 'H')  # HGNC symbols
pathway_entrez <- getXrefList('WP554', 'L')  # Entrez IDs

# Download pathway as GMT for custom analysis
downloadPathwayArchive(organism = 'Homo sapiens', format = 'gmt')

Custom GMT-Based Analysis

Goal: Run enrichment using a downloaded WikiPathways GMT file for offline or custom analysis.

Approach: Download the GMT archive via rWikiPathways, read it with read.gmt, and run enricher.

# Download WikiPathways GMT
library(rWikiPathways)
downloadPathwayArchive(organism = 'Homo sapiens', format = 'gmt', destpath = '.')

# Read GMT and run enrichment
wp_gmt <- read.gmt('wikipathways-Homo_sapiens.gmt')

wp_custom <- enricher(
    gene = entrez_ids,
    TERM2GENE = wp_gmt,
    pvalueCutoff = 0.05
)

Different Organisms

# Mouse
wp_mouse <- enrichWP(gene = mouse_entrez, organism = 'Mus musculus')

# Rat
wp_rat <- enrichWP(gene = rat_entrez, organism = 'Rattus norvegicus')

# Zebrafish
wp_zfish <- enrichWP(gene = zfish_entrez, organism = 'Danio rerio')

# List all available organisms
library(rWikiPathways)
listOrganisms()

Compare Clusters

Goal: Compare WikiPathways enrichment across multiple gene lists (e.g., upregulated vs downregulated).

Approach: Use compareCluster with enrichWP to run enrichment per group and visualize with dotplot.

gene_clusters <- list(
    upregulated = up_genes,
    downregulated = down_genes
)

compare_wp <- compareCluster(
    geneClusters = gene_clusters,
    fun = 'enrichWP',
    organism = 'Homo sapiens',
    pvalueCutoff = 0.05
)

dotplot(compare_wp)

Export Results

results_df <- as.data.frame(wp_result)
write.csv(results_df, 'wikipathways_enrichment.csv', row.names = FALSE)

Key Parameters

ParameterDefaultDescription
generequiredVector of Entrez IDs
organismrequiredFull species name
pvalueCutoff0.05P-value threshold
pAdjustMethodBHAdjustment method
universeNULLBackground genes
minGSSize10Min genes per pathway
maxGSSize500Max genes per pathway

Common Organisms

Common NameScientific Name
HumanHomo sapiens
MouseMus musculus
RatRattus norvegicus
ZebrafishDanio rerio
Fruit flyDrosophila melanogaster
C. elegansCaenorhabditis elegans
ArabidopsisArabidopsis thaliana
YeastSaccharomyces cerevisiae

WikiPathways vs Other Databases

FeatureWikiPathwaysKEGGReactome
CurationCommunityExpertPeer-reviewed
LicenseOpen (CC0)CommercialOpen
Species30+4000+7
FocusDisease, drugMetabolicSignaling
UpdatesContinuousOngoingQuarterly

Related Skills

  • go-enrichment - Gene Ontology enrichment
  • kegg-pathways - KEGG pathway enrichment
  • reactome-pathways - Reactome pathway enrichment
  • gsea - Gene Set Enrichment Analysis
  • enrichment-visualization - Visualization functions

What ships with it: 3 files

4.6 KB alongside SKILL.md

Keep looking

Skills are one crate of 325,949. Ordering is by how many stacks a row turns up in, so the top of any crate is what has actually been picked rather than what has the most stars.