agentsclimarketplace

Drug database record extraction

Skill HolobiomicsLab/asb-skill-collections/collections/metabolomics/v2/skills/drug-database-record-extraction

Curated, evidence-grounded skill and software-tool collections for scientific AI agents, generated by the AgenticScienceBuilder

Install
npx -y skills add HolobiomicsLab/asb-skill-collections --skill drug-database-record-extraction

Assembled from the repository path, not quoted from the project. Check it against their README if it does not work.

One thing to look at

  • 14 stars14 stars. Stars are a popularity signal and not a quality one, but at this level it is likely that nobody has read this closely except its author, and you would be relying on your own review.

What its author says it does

Copied from the file, not written here

Use when when you have obtained a DrugBank release file (requiring access credentials) and need to integrate drug chemical structure, name, and identifier information into a metadata cleanup or chemical enrichment pipeline.

The file declares its own license as CC-BY-4.0. That is the author’s claim about this one file, and it is not the same thing as the license GitHub reports for the repository, which is listed with the other numbers below.

SKILL.md

6.6 KB, ~1.1k tokens by cl100k_base, as published. Nobody here has run it

drug-database-record-extraction

Summary

Extract and parse structured drug information from DrugBank release files using the drugbank_extraction.py script to populate a standardized drug metadata table for downstream chemical and biomedical queries. This skill transforms raw DrugBank XML or tabular releases into validated, queryable drug records with consistent schema.

When to use

When you have obtained a DrugBank release file (requiring access credentials) and need to integrate drug chemical structure, name, and identifier information into a metadata cleanup or chemical enrichment pipeline. Trigger: raw DrugBank file is available locally and downstream workflows expect normalized drug records with consistent column names and minimum required fields (name, structure, identifiers).

When NOT to use

  • DrugBank access credentials are unavailable or expired; use alternative drug sources (Broad Institute Drug Repurposing Hub, DrugCentral) instead.
  • Input file is already a validated, normalized drug table conforming to the metadata template; skip extraction and proceed directly to metadata cleanup.
  • Downstream workflow does not require DrugBank-specific fields (e.g., only PubChem or LOTUS natural product data is needed).

Inputs

Outputs

  • Parsed drug records in standardized tabular format (CSV/TSV or structured object)
  • Extracted drug information including: drug name, chemical structure (SMILES or InChI), DrugBank accession ID, synonyms, and cross-references
  • Validated artifact conforming to metadata template schema

How to apply

Download or obtain the DrugBank release file from the official DrugBank portal (https://go.drugbank.com/releases/latest), which requires account access. Execute drugbank_extraction.py on the downloaded file to parse and extract drug records into a structured format (matching the metadata template with standardized column names). Validate the output artifact by confirming that (1) all expected drug records are present, (2) mandatory fields (drug name, structure/SMILES, accession IDs) are populated, and (3) the schema matches downstream workflow expectations. The script should produce a tabular or structured output suitable for import into the metadata cleanup pipeline.

Related tools

Examples

python drugbank_extraction.py --input drugbank_release.xml --output drug_records.csv

Evaluation signals

  • Output file exists and contains non-zero rows of drug records with expected column names (drug name, structure, accession ID, synonyms).
  • All mandatory fields are populated for ≥95% of extracted records; document any missing-value patterns or records with incomplete structures.
  • Schema validation: output column names and data types match the metadata template (https://docs.google.com/spreadsheets/d/1v6_IlGS3VgycGc-mSSdNeocY-CFXpONVZbuh3XNLX2E/edit?usp=sharing).
  • Spot-check: sample 10–20 extracted records against the original DrugBank release file to verify name, structure, and identifier accuracy.
  • Downstream integration: extracted records can be successfully loaded into the metadata cleanup pipeline without schema or type errors.

Limitations

  • DrugBank requires active account access and periodic credential renewal; extraction will fail if credentials are expired or invalid.
  • Script is tightly coupled to the specific DrugBank file format and release version; updates to DrugBank schema may require script modifications.
  • DrugBank may contain incomplete or outdated structure information for some compounds; users should cross-validate against PubChem if needed.
  • No changelog is documented in the repository, making it difficult to track upstream changes to the extraction script.

Evidence

  • [intro] run drugbank_extraction.py on that file: "DrugBank (access needed): Download and run drugbank_extraction.py on that file"
  • [other] parse and extract drug information into a structured format: "The drugbank_extraction.py script is executed on a downloaded DrugBank release file to extract and parse drug information for use in the metadata cleanup pipeline."
  • [other] Validate the output artifact contains the expected drug records and fields required by downstream metadata cleanup workflows: "Validate the output artifact contains the expected drug records and fields required by downstream metadata cleanup workflows."
  • [readme] same column names and minimum needed information for the query: "Please use the [template] for your metadata for having same column names and minimum needed information for the query"
  • [readme] access needed: "DrugBank (access needed): Download"

Keep looking

Skills are one crate of 328,083. Ordering is by how many stacks a row turns up in, so the top of any crate is what has actually been picked rather than what has the most stars.