agentsclimarketplace

Annotation accuracy and coverage metrics computation

Skill HolobiomicsLab/asb-skill-collections/collections/metabolomics/v2/skills/annotation-accuracy-and-coverage-metrics-computation

Curated, evidence-grounded skill and software-tool collections for scientific AI agents, generated by the AgenticScienceBuilder

Install
npx -y skills add HolobiomicsLab/asb-skill-collections --skill annotation-accuracy-and-coverage-metrics-computation

Assembled from the repository path, not quoted from the project. Check it against their README if it does not work.

One thing to look at

  • 14 stars14 stars. Stars are a popularity signal and not a quality one, but at this level it is likely that nobody has read this closely except its author, and you would be relying on your own review.

What its author says it does

Copied from the file, not written here

Use when after executing an end-to-end structure annotation pipeline (such as BAM) on a validation dataset with known reference annotations.

The file declares its own license as CC-BY-4.0. That is the author’s claim about this one file, and it is not the same thing as the license GitHub reports for the repository, which is listed with the other numbers below.

SKILL.md

6.2 KB, 972 tokens by cl100k_base, as published. Nobody here has run it

annotation-accuracy-and-coverage-metrics-computation

License: restricted — no clear open-source license detected for the underlying tool; verify licensing before commercial use or redistribution. <!-- asb-license-banner -->

Summary

Compute validation metrics (accuracy and coverage) for molecular structure annotations by comparing pipeline predictions against reference annotations. This skill quantifies the quality of biotransformation-based annotation methods on untargeted metabolomics data.

When to use

Apply this skill after executing an end-to-end structure annotation pipeline (such as BAM) on a validation dataset with known reference annotations. Use it when you need to benchmark annotation performance, verify reproducibility of reported metrics, or assess the quality of predicted molecular structures against ground truth.

When NOT to use

  • No reference annotations are available or validation dataset is absent
  • Pipeline has not yet been executed on the validation data (metrics require both predictions and ground truth)
  • Annotations are not directly comparable (e.g., predictions in different chemical representation formats without conversion)

Inputs

  • Pipeline output annotations (predicted structure identifiers, SMILES, or InChI strings)
  • Reference annotation dataset (ground-truth molecular structures with SMILES or InChI)
  • Query molecules list (suspects/anchors with identifiers and masses)

Outputs

  • Annotation accuracy metric (fraction of correct predictions)
  • Annotation coverage metric (fraction of queries receiving predictions)
  • Metrics report documenting accuracy, coverage, and performance benchmarks

How to apply

Obtain both the predicted structure annotations generated by the annotation pipeline (e.g., BAM output) and the reference annotations from your validation dataset. For each predicted annotation, compare it against the corresponding reference annotation using a metric of structural equivalence (e.g., SMILES matching or InChI comparison). Compute accuracy as the fraction of predicted annotations that match reference annotations; compute coverage as the fraction of query molecules that received a prediction. Generate a metrics report documenting achieved accuracy, coverage, and any additional performance benchmarks (e.g., per-molecule or per-class breakdowns). Document the comparison method and any tie-breaking rules used when multiple candidate structures are ranked.

Related tools

Evaluation signals

  • Accuracy is computed as the fraction of predicted annotations matching reference annotations (range 0–1)
  • Coverage is computed as the fraction of query molecules that received at least one prediction (range 0–1)
  • Metrics report includes per-dataset breakdowns (e.g., KEGG vs. RetroRules reaction data) when applicable
  • Comparison method (e.g., SMILES or InChI matching) is explicitly documented to ensure reproducibility
  • Results are stratified by molecular class, anchor type, or suspect mass range where data permits

Limitations

  • Accuracy depends on the quality and completeness of the reference annotation dataset; incomplete or incorrect ground truth will bias results
  • Coverage may be artificially low if the pipeline fails to generate predictions for molecules outside the scope of the reaction rule dataset (e.g., rare biotransformations)
  • Structural equivalence comparison requires careful handling of stereochemistry, tautomerism, and chemical representation canonicalization to avoid false mismatches
  • Metrics do not capture ranking quality: a correct structure ranked low is counted the same as one ranked high

Evidence

  • [other] Compute validation metrics (annotation accuracy, coverage) by comparing pipeline predictions against reference annotations.: "Compute validation metrics (annotation accuracy, coverage) by comparing pipeline predictions against reference annotations."
  • [readme] All data necessary to run the evaluation of BAM described in our paper is included in the data folder.: "All data necessary to run the evaluation of BAM described in our paper is included in the data folder."
  • [readme] BAM checks if the suspect molecule is known by checking whether the SMILES or InChI is specified in the molecules_of_interest csv file.: "BAM checks if the suspect molecule is known by checking whether the SMILES or InChI is specified in the molecules_of_interest csv file."
  • [other] Generate a metrics report documenting achieved accuracy, coverage, and any performance benchmarks.: "Generate a metrics report documenting achieved accuracy, coverage, and any performance benchmarks."

What ships with it

Read from the repository

Just SKILL.md. No reference files, no scripts.

Keep looking

Skills are one crate of 326,984. Ordering is by how many stacks a row turns up in, so the top of any crate is what has actually been picked rather than what has the most stars.