agentsclimarketplace

Mibig metadata extraction

Skill HolobiomicsLab/asb-skill-collections/collections/metabolomics/v2/skills/mibig-metadata-extraction

Curated, evidence-grounded skill and software-tool collections for scientific AI agents, generated by the AgenticScienceBuilder

Install
npx -y skills add HolobiomicsLab/asb-skill-collections --skill mibig-metadata-extraction

Assembled from the repository path, not quoted from the project. Check it against their README if it does not work.

One thing to look at

  • 14 stars14 stars. Stars are a popularity signal and not a quality one, but at this level it is likely that nobody has read this closely except its author, and you would be relying on your own review.

What its author says it does

Copied from the file, not written here

Use when when you need to audit, inventory, or report on the curation state of MIBiG entries; when cluster.status values must be validated or aggregated for quality control; when building a status index to support data governance or release workflows.

The file declares its own license as CC-BY-4.0. That is the author’s claim about this one file, and it is not the same thing as the license GitHub reports for the repository, which is listed with the other numbers below.

SKILL.md

5.3 KB, 856 tokens by cl100k_base, as published. Nobody here has run it

mibig-metadata-extraction

License: restricted — no clear open-source license detected for the underlying tool; verify licensing before commercial use or redistribution. <!-- asb-license-banner -->

Summary

Extract and aggregate entry-level metadata from MIBiG JSON files, particularly the cluster.status field, to construct a comprehensive index of curation status across the repository. This skill enables systematic tracking and validation of annotation completeness in the MIBiG secondary metabolite database.

When to use

When you need to audit, inventory, or report on the curation state of MIBiG entries; when cluster.status values must be validated or aggregated for quality control; when building a status index to support data governance or release workflows.

When NOT to use

  • When only sequence data or GenBank files are needed — use the genbanks directory instead
  • When a pre-built, curated status index is already available from MIBiG's public API or web interface
  • When cluster.status field semantics or controlled vocabulary are not documented — context required first

Inputs

  • MIBiG JSON file collection (from mibig-json/data directory)
  • Entry identifier mappings (filename or internal ID field)

Outputs

  • Structured entry-status index (CSV or JSON format)
  • Validation report (parsing completeness, missing status fields)

How to apply

Clone or download the mibig-json repository from github.com/mibig-secmet/mibig-json and locate the data directory. Iterate over all JSON files in the directory, parsing each file to extract the entry identifier (from filename or internal ID field) and the cluster.status field value. Aggregate the extracted tuples into a structured index using CSV or JSON format with columns for entry ID and cluster status. Validate that all JSON files have been parsed, that every entry has a status value present, and that status values conform to the expected controlled vocabulary. Manual expert review is recommended to spot-check a sample of entries against the original JSON structure.

Related tools

Evaluation signals

  • All JSON files in the data directory have been successfully parsed with no read errors
  • Entry count in output index matches the total number of JSON files in data directory
  • Every entry in the index has a non-null cluster.status value
  • Status values conform to the expected controlled vocabulary (manually verified against sample entries)
  • Index is sortable, searchable, and can be linked back to original JSON files by entry ID

Limitations

  • No changelog is available in the repository to document version history or changes to the cluster.status field semantics
  • Manual expert review is required to validate status field values — automated validation alone cannot ensure semantic correctness
  • The index will only be as current as the last repository clone/download; updates require re-running the extraction workflow

Evidence

  • [intro] MIBiG curation data is maintained in JSON format with entry status tracked through the cluster.status field: "MIBiG curation data in JSON format... entry status is now tracked via the cluster.status field"
  • [readme] The data directory contains the current MIBiG datasets: "The current datasets for [MIBiG] live in the data directory, entry status is now tracked via the cluster.status field"
  • [other] Extraction workflow steps including cloning, JSON parsing, field extraction, and aggregation: "Clone or download the mibig-json repository from github.com/mibig-secmet/mibig-json. 2. Locate and read all JSON files in the data directory. 3. For each JSON file, extract the entry identifier"
  • [other] Validation of extraction completeness and status value presence: "Validate that all entries have been parsed and status values are present"

What ships with it

Read from the repository

Just SKILL.md. No reference files, no scripts.

Keep looking

Skills are one crate of 326,984. Ordering is by how many stacks a row turns up in, so the top of any crate is what has actually been picked rather than what has the most stars.