agentsclimarketplace

Molecular targets data

Skill BioTender-max/awesome-bio-agent-skills/skills/drugclaw/molecular_targets_data

Query the NCI-60 Molecular Target (Protein) database from the Developmental Therapeutics Program. Use when the user asks about protein expression of drug targets across the NCI-60 cancer cell line panel, or wants to look up a gene, cell line, or cancer panel in the NCI DTP molecular target dataset.From its SKILL.md

Install
npx -y skills add BioTender-max/awesome-bio-agent-skills --skill molecular_targets_data

Assembled from the repository path, not quoted from the project. Check it against their README if it does not work.

One thing to look at

  • no licenseNo license file was found in the repository. Code published without one is not open source by default, so using it at work is a question for whoever answers licensing questions where you are.

SKILL.md

3.9 KB, ~1.0k tokens by cl100k_base, as published. Nobody here has run it

NCI DTP Molecular Target Query Skill

Search NCI-60 protein-level molecular target data by any entity. Auto-detects type:

Input PatternDetected AsMatch Logic
EGFR, TP53gene / proteinexact on GENE or substring on ENTITY_MEASURED, TITLE
MCF7, NCI-H460cell linesubstring on cellname
Breast, Leukemiacancer panelsubstring on pname
12345 (pure digits)MOLTIDexact on MOLTID (NCI pattern #)

Drug-Relevance Guide

Only protein expression data is downloaded — proteins are the direct molecular targets of drugs (kinase inhibitors → kinase expression, antibodies → receptor levels).

DatasetDownload?Reason
WEB_DATA_PROTEIN.ZIPYESDrug targets: protein expression across NCI-60
WEB_DATA_ALL_MT.ZIPOptionalSuperset incl. enzyme activity (also drug-relevant)
WEB_DATA_DNA.ZIPNoGenomic characterisation, not drug targets
WEB_DATA_SEQUENOM_METHYLATION.ZIPNoEpigenetic; indirect
WEB_DATA_*_MIR.ZIPNomicroRNA regulation; indirect
WEB_DATA_METABOLON*.ZIPNoMetabolomics; downstream, indirect
Microarray / SNP / CopyNum / KaryotypeNoCell-line genomic profiling; indirect

API

FunctionInputReturns
load_data(path)file or directory pathlist[dict]
search(data, entity)data + single entity stringlist[dict]
search_batch(data, entities)data + list of entity stringsdict[str, list[dict]]
summarize(hits, entity)hit list + labelcompact LLM-readable text
to_json(hits)hit listlist[dict] (JSON-serialisable)
query(data, entities, top_n)data + str or listtext block

Record Fields

Each record contains:

FieldDescription
MOLTIDNCI pattern number (molecular target ID)
GENEGene symbol
TITLEGene / protein full name
MOLTNBRNCI experiment ID
PANELNBR / CELLNBRInternal panel / cell identifiers
pnameCancer panel name (e.g., Breast, Leukemia)
cellnameCell line name (e.g., MCF7, A549/ATCC)
ENTITY_MEASUREDWhat was measured (e.g., protein name)
GeneIDNCBI Gene ID
UNITSMeasurement units
METHODAssay method
VALUENumeric measurement value
TEXTAdditional notes

Usage

See if __name__ == "__main__" block in 16_NCI_DTP_MolTarget.py for runnable examples covering: single gene query, cell line profile, batch multi-target search, cancer panel query, and JSON output.

Quick Examples

from importlib.machinery import SourceFileLoader
mt = SourceFileLoader("mt", "16_NCI_DTP_MolTarget.py").load_module()

data = mt.load_data()

# Single drug target
print(mt.query(data, "EGFR"))

# Multiple targets
print(mt.query(data, ["TP53", "BRAF", "HER2"]))

# Cell line molecular profile
print(mt.query(data, "MCF7"))

# JSON for downstream pipeline
import json
hits = mt.search(data, "EGFR")
print(json.dumps(mt.to_json(hits[:5]), indent=2))

Data Source & Download

Download Commands

# File already at:
# resources_metadata/dti/Molecular Target Data/WEB_DATA_PROTEIN.TXT

What ships with it: 4 files

15.0 KB alongside SKILL.md, 4 of them executable

Keep looking

Skills are one crate of 325,949. Ordering is by how many stacks a row turns up in, so the top of any crate is what has actually been picked rather than what has the most stars.