agentsclimarketplace

Metadata validation rule specification

Skill HolobiomicsLab/asb-skill-collections/collections/metabolomics/v2/skills/metadata-validation-rule-specification

Curated, evidence-grounded skill and software-tool collections for scientific AI agents, generated by the AgenticScienceBuilder

Install
npx -y skills add HolobiomicsLab/asb-skill-collections --skill metadata-validation-rule-specification

Assembled from the repository path, not quoted from the project. Check it against their README if it does not work.

One thing to look at

  • 14 stars14 stars. Stars are a popularity signal and not a quality one, but at this level it is likely that nobody has read this closely except its author, and you would be relying on your own review.

What its author says it does

Copied from the file, not written here

Use when mSMetaEnhancer retrieves metadata attributes (SMILES, InChI, CAS numbers, IUPAC names, formulas) from external services (CIR, CTS, PubChem, IDSM, BridgeDb) and you need to guarantee that only correctly formatted values are written back to .msp output files.

The file declares its own license as CC-BY-4.0. That is the author’s claim about this one file, and it is not the same thing as the license GitHub reports for the repository, which is listed with the other numbers below.

SKILL.md

8.8 KB, ~1.5k tokens by cl100k_base, as published. Nobody here has run it

metadata-validation-rule-specification

Summary

Define and implement attribute-specific validation schemas that enforce format and content rules on metadata fetched from external services before writing to mass spectra files. This skill ensures data quality and consistency in MSMetaEnhancer's metadata enrichment pipeline by intercepting and validating converter output against predefined acceptance criteria.

When to use

Apply this skill when MSMetaEnhancer retrieves metadata attributes (SMILES, InChI, CAS numbers, IUPAC names, formulas) from external services (CIR, CTS, PubChem, IDSM, BridgeDb) and you need to guarantee that only correctly formatted values are written back to .msp output files. Use it whenever converter output requires quality assurance before committing to spectra objects.

When NOT to use

  • Input metadata is already validated by upstream quality control or is sourced from trusted internal databases rather than external web services
  • Use case requires accepting any fetched value regardless of format (e.g., exploratory or permissive annotation modes)
  • Validation rules or service specifications are undefined or not documented for the metadata attributes being enriched

Inputs

  • .msp spectra files with unvalidated metadata annotations
  • converter output dictionaries from web services (CIR, CTS, PubChem, IDSM, BridgeDb)
  • attribute-type specifications and format requirements from service documentation

Outputs

  • .msp files with validated metadata annotations
  • structured validation report log (attribute name, value, rule, pass/fail status, service source)
  • validation statistics and reject/accept counts per attribute type

How to apply

Design a validation schema that defines acceptance criteria for each metadata attribute type—for example, SMILES format validation, InChI prefix checking, CAS number format verification—based on service-specific specifications. Implement a Validator class that intercepts converter output dictionaries between response parsing and spectra write operations, applying attribute-specific validation rules using the existing Job and Converter architecture. Log all validation decisions (attribute name, fetched value, validation rule applied, pass/fail status, service source) to a structured report. Unit test each attribute type with pytest to verify correct formats are accepted and malformed data rejected. End-to-end validation with sample .msp files confirms only validated metadata reaches output and all decisions are captured in validation logs.

Related tools

Examples

from MSMetaEnhancer.libs.validators import Validator; validator = Validator(schema={'SMILES': {'pattern': r'^[A-Za-z0-9()\[\]{}\\%=+#@-]*$'}, 'InChI': {'prefix': 'InChI='}, 'CAS': {'pattern': r'^\d{1,6}-\d{2}-\d'}}); result = validator.validate({'SMILES': 'CCO', 'InChI': 'InChI=1S/C2H6O/c1-2-3/h3H,2H2,1H3', 'CAS': '64-17-5'}); print(result)

Evaluation signals

  • Validation logs capture all metadata attributes with rule applied, pass/fail status, and source service; 100% traceability of decisions
  • Unit tests confirm SMILES, InChI, CAS number, and other attribute formats are correctly accepted when well-formed and rejected when malformed
  • End-to-end .msp output files contain only metadata that passed validation; no rejected values are written to spectra
  • Validation report shows accept/reject counts and distribution by attribute type and service, revealing data quality patterns
  • Structured validation schema schema file documents acceptance criteria for each attribute type with reference to service specifications

Limitations

  • Validation rules must be manually specified based on service documentation; incomplete or ambiguous service specs may result in overly permissive or overly strict schemas
  • External service response formats may vary or change unexpectedly, requiring schema updates and test maintenance
  • Validated-only approach may discard valid but non-standard metadata values if schema is too narrow; balance between strictness and coverage must be tuned empirically

Evidence

  • [other] Design a validation schema defining acceptance criteria for each metadata attribute type (e.g., SMILES format, InChI prefix, CAS number format) based on service specifications: "Design a validation schema defining acceptance criteria for each metadata attribute type (e.g., SMILES format, InChI prefix, CAS number format) based on service specifications."
  • [other] Implement a Validator class that intercepts converter output dictionaries and applies attribute-specific validation rules before metadata is committed to the spectra object: "Implement a Validator class that intercepts converter output dictionaries and applies attribute-specific validation rules before metadata is committed to the spectra object."
  • [other] Integrate the Validator into the conversion pipeline (between converter response parsing and spectra write) using the existing Job and Converter architecture: "Integrate the Validator into the conversion pipeline (between converter response parsing and spectra write) using the existing Job and Converter architecture."
  • [other] Log all validation results (attribute name, fetched value, validation rule applied, pass/fail status, service source) to a structured validation report file: "Log all validation results (attribute name, fetched value, validation rule applied, pass/fail status, service source) to a structured validation report file."
  • [other] MSMetaEnhancer implements attribute validation tracked in logs to check fetched annotation values before they are written back to spectra: "MSMetaEnhancer implements attribute validation tracked in logs to check fetched annotation values before they are written back to spectra, ensuring data quality during the metadata enrichment process."
  • [readme] It adds metadata like SMILES, InChI, and CAS number fetched from the following services: [CIR], [CTS], [PubChem], [IDSM], and [BridgeDb]: "It adds metadata like SMILES, InChI, and CAS number fetched from the following services: [CIR], [CTS], [PubChem], [IDSM], and [BridgeDb]."

What ships with it

Read from the repository

Just SKILL.md. No reference files, no scripts.

Keep looking

Skills are one crate of 328,083. Ordering is by how many stacks a row turns up in, so the top of any crate is what has actually been picked rather than what has the most stars.