agentsclimarketplace

Compound database matching

Skill HolobiomicsLab/asb-skill-collections/packs/metabolomics/lc-ms/skills/compound-database-matching

Curated, evidence-grounded skill and software-tool collections for scientific AI agents, generated by the AgenticScienceBuilder

Install
npx -y skills add HolobiomicsLab/asb-skill-collections --skill compound-database-matching

Assembled from the repository path, not quoted from the project. Check it against their README if it does not work.

One thing to look at

  • 14 stars14 stars. Stars are a popularity signal and not a quality one, but at this level it is likely that nobody has read this closely except its author, and you would be relying on your own review.

What its author says it does

Copied from the file, not written here

Use when you have MS2 .mzML spectral data from untargeted metabolomics and need to assign chemical identities to detected precursor ions.

The file declares its own license as CC-BY-4.0. That is the author’s claim about this one file, and it is not the same thing as the license GitHub reports for the repository, which is listed with the other numbers below.

SKILL.md

8.9 KB, ~1.7k tokens by cl100k_base, as published. Nobody here has run it

compound-database-matching

Summary

Perform compound database dereplication on MS2 spectral data by dispatching spectra through SIRIUS or MetFrag to generate per-spectrum candidate annotation lists ranked by match score. This skill reduces false-positive identifications and prioritizes putative metabolite assignments for downstream validation.

When to use

Apply this skill when you have MS2 .mzML spectral data from untargeted metabolomics and need to assign chemical identities to detected precursor ions. Use it after spectral database dereplication (library matching) has been performed, when you want to expand beyond known reference spectra to include in silico predictions from compound structure databases like PubChem or COCONUT.

When NOT to use

  • Input is only MS1 precursor mass data without MS/MS fragmentation spectra — SIRIUS and MetFrag require MS/MS fragment information to score and rank candidates.
  • Workflow has already performed final candidate selection (e.g., output is a single assigned compound per spectrum) — this skill is for candidate generation and ranking, not validation.
  • Spectral data are in formats other than .mzML (e.g., raw vendor formats, NetCDF, or already-processed peak lists without intensity calibration) — requires conversion upstream.

Inputs

  • MS2 spectral data in .mzML format
  • Per-spectrum precursor m/z and intensity
  • MS/MS fragment peak list (m/z, intensity) per spectrum
  • Compound structure database (CSV: Identifier, InChI, SMILES, molecular_weight) if using MetFrag

Outputs

  • Per-spectrum candidate annotation list (CSV/TSV table)
  • Columns: spectrum_id, candidate_identifier, SMILES, InChI, molecular_weight, match_score, rank
  • Ranked candidate compounds with structural metadata and scoring metrics

How to apply

Load processed MS2 spectral data (e.g., from the Spectra R package after spectral library dereplication). For each spectrum with a precursor m/z, dispatch it to either SIRIUS or MetFrag. SIRIUS performs structure elucidation and database searching using MS/MS fragmentation patterns; MetFrag requires a local compound database (CSV with Identifier, InChI, SMILES, molecular_weight columns) and scores candidate structures by fragment peak matching. Collect the ranked candidate matches with association scores for each spectrum. Output the results as a structured table (CSV or TSV) with columns for spectrum ID, candidate compound identifier, structure metadata (SMILES, InChI), calculated/library molecular weight, match score, and rank. Use a score threshold (e.g., 0.75 for MetFrag) in downstream filtering to control false-discovery rate.

Related tools

Examples

Rscript Workflow_R_Script_all_MetFrag.r sample.mzML gnps.rda hmdb.rda mbankNIST.rda 15 TRUE coconut COCONUT_Jan2022.csv sample/insilico/metparam_list.txt MetFragCommandLine-2.5.0.jar

Evaluation signals

  • Output table contains one or more ranked candidate rows per input spectrum; no spectrum should return zero candidates unless precursor m/z or fragment data are invalid.
  • Match scores are within the expected range for the tool used (e.g., MetFrag scores typically 0–1); scores should correlate with chemical plausibility (higher score = better fragment peak overlap).
  • Candidate SMILES and InChI strings are valid and parse without error in RDKit; calculated molecular weights match the candidate's structure within ±0.01 Da.
  • Spectrum with higher MS/MS spectral quality (more fragment peaks, higher intensity) should yield candidates with higher match scores than low-quality spectra.
  • Candidates ranked #1 should be chemically reasonable for the experimental context (e.g., polar metabolites in aqueous extract should not be rank-1 lipids).

Limitations

  • SIRIUS support in CWL workflows is not yet fully implemented; SIRIUS is available only via Docker containers at present, limiting parallelization options.
  • MetFrag performance depends heavily on the completeness and accuracy of the input compound database; missing or incorrect SMILES/InChI will cause low scores or false negatives.
  • Both tools assume high-resolution MS/MS data with accurate m/z and intensity calibration; low-resolution or uncalibrated spectra may yield false candidates.
  • Computational cost scales with database size and number of spectra; processing time per precursor mass is ~2 minutes on 64 GB RAM Ubuntu system; batching or HPC submission (SLURM) recommended for >10 precursor masses.
  • No internal validation that a top-ranked candidate is correct; downstream filtering (e.g., score threshold of 0.75) and orthogonal confirmation (e.g., retention time, biological plausibility) are necessary.

Evidence

  • [other] The workflow performs compound database dereplication by dispatching MS2 spectral data through either SIRIUS or MetFrag tools, which generate per-spectrum candidate annotation lists.: "performs compound database dereplication by dispatching MS2 spectral data through either SIRIUS or MetFrag tools, which generate per-spectrum candidate annotation lists"
  • [other] For each spectrum, dispatch to either SIRIUS or MetFrag for compound database dereplication against PubChem. Collect per-spectrum candidate matches with scores and structural metadata. Output ranked candidate annotation list as a structured table (CSV or TSV format).: "dispatch to either SIRIUS or MetFrag for compound database dereplication against PubChem. Collect per-spectrum candidate matches with scores and structural metadata. Output ranked candidate"
  • [readme] The workflow takes MS2 .mzML format data files as an input in R. It performs spectral database dereplication using R Package Spectra and compound database dereplication using SIRIUS OR MetFrag.: "workflow takes MS2 .mzML format data files as an input in R. It performs spectral database dereplication using R Package Spectra and compound database dereplication using SIRIUS OR MetFrag"
  • [readme] We recommend that the local file should be a csv file with atleast the following columns: 'Identifier' 'InChI' 'SMILES' 'molecular_weight'.: "local file should be a csv file with atleast the following columns: 'Identifier' 'InChI' 'SMILES' 'molecular_weight'"
  • [readme] one precursor mass takes 2 minutes on an Ubuntu system with 64GB RAM to run Workflow_R_Script_all_MetFrag.r: "one precursor mass takes 2 minutes on an Ubuntu system with 64GB RAM to run Workflow_R_Script_all_MetFrag.r"
  • [readme] SIRIUS is only accomodated with the docker containers, (the workflow will be completely operable in CWL in the future). At the moment, the CWL version can be used with MetFrag.: "SIRIUS is only accomodated with the docker containers. At the moment, the CWL version can be used with MetFrag"

Keep looking

Skills are one crate of 328,083. Ordering is by how many stacks a row turns up in, so the top of any crate is what has actually been picked rather than what has the most stars.