agentsclimarketplace

Model weight loading and deployment

Skill HolobiomicsLab/asb-skill-collections/collections/metabolomics/v2/skills/model-weight-loading-and-deployment

Use when you have a pre-trained MSGO model checkpoint (PFAS or lipid variant) and need to evaluate it against a real mass spectrometry dataset (300+ real spectra, LC–QTOF, or custom CSV) to generate predicted molecular structures and compare against ground truth or baseline results.From its SKILL.md

Install
npx -y skills add HolobiomicsLab/asb-skill-collections --skill model-weight-loading-and-deployment

Assembled from the repository path, not quoted from the project. Check it against their README if it does not work.

One thing to look at

  • 15 stars15 stars. Stars are a popularity signal and not a quality one, but at this level it is likely that nobody has read this closely except its author, and you would be relying on your own review.

What its file declares

Copied from the file, not written here

The file declares its own license as CC-BY-4.0. That is the author’s claim about this one file, and it is not the same thing as the license GitHub reports for the repository, which is listed with the other numbers below.

SKILL.md

6.4 KB, ~1.1k tokens by cl100k_base, as published. Nobody here has run it

model-weight-loading-and-deployment

Summary

Load pre-trained deep learning model weights from a repository and deploy them for inference on new spectral data to generate molecular structure predictions. This skill enables reproducibility of published structure-generation performance without retraining.

When to use

You have a pre-trained MSGO model checkpoint (PFAS or lipid variant) and need to evaluate it against a real mass spectrometry dataset (300+ real spectra, LC–QTOF, or custom CSV) to generate predicted molecular structures and compare against ground truth or baseline results.

When NOT to use

  • Model weights have not been downloaded or checkpoint path is invalid—training must be performed first using tools/train.py
  • Input spectra are in non-standard formats (not CSV or incompatible with eval_standard.py schema)
  • Evaluating on pseudo SMILES-spectrum pairs used for training—use the 300+ real spectrum validation set instead to assess generalization

Inputs

  • Pre-trained model checkpoint directory (ckpts/pfas or ckpts/lipid)
  • Real mass spectrometry spectrum dataset (CSV format with spectrum features)
  • Polarization mode specification (pos or neg)
  • Beam search size parameter (integer, 300–500)

Outputs

  • Results CSV file with predicted SMILES structures
  • Top-10 ranked predictions per spectrum with scores
  • Inference time and structure-generation performance metrics

How to apply

Load the released MSGO model weights from github.com/aaronma2020/MSGO using Python 3.7 and Torch 1.7.1 by specifying the checkpoint path (ckpts/pfas or ckpts/lipid). Prepare your input spectrum dataset as a CSV file compatible with the eval_standard.py evaluation script, specifying polarization mode (pos or neg) and beam search size (300–500 depending on model variant). Execute inference by calling tools/eval.py or tools/eval_standard.py with the model path and input CSV, collecting predicted SMILES structures ranked by beam search score. Verify correctness by confirming the output CSV includes top-10 predictions with associated confidence scores for each spectrum.

Related tools

  • Python (Runtime environment for model loading and inference)
  • Torch (Deep learning framework for checkpoint deserialization and GPU-accelerated inference)
  • MSGO repository (Source of pre-trained model weights and evaluation scripts (eval.py, eval_standard.py)) — github.com/aaronma2020/MSGO

Examples

python tools/eval_standard.py --log_path ckpts/pfas --real_csv ./data/example/pfas.csv --out_csv ./pfas_results.csv --beam_size 500 --polar neg

Evaluation signals

  • Model checkpoint successfully loads without Torch deserialization errors
  • Output CSV contains exactly 10 ranked predictions per spectrum with monotonically decreasing beam search scores
  • Predicted SMILES are valid and canonicalizable (no malformed SMILES strings)
  • Structure-generation metrics (e.g., exact match rate on 300+ real spectra) match or closely reproduce the reported paper results
  • Inference completes without CUDA out-of-memory or framework compatibility errors for the specified Python 3.7 and Torch 1.7.1 versions

Limitations

  • Model weights are specialized to PFAS or lipid chemical classes—deployment on spectra from other compound classes may yield poor predictions
  • Evaluation requires exact Python 3.7 and Torch 1.7.1 versions; newer PyTorch releases may break checkpoint compatibility
  • Pseudo SMILES-spectrum pairs used during training may introduce systematic bias in predicted structures; real-world validation datasets are recommended
  • Beam search size (300–500) trades inference speed against coverage of candidate structures—smaller beam sizes may miss correct predictions

Evidence

  • [other] Load the pre-trained MSGO model weights from the released github.com/aaronma2020/MSGO repository using Python 3.7 and Torch 1.7.1.: "Load the pre-trained MSGO model weights from the released github.com/aaronma2020/MSGO repository using Python 3.7 and Torch 1.7.1."
  • [readme] Download the model weights in ckpts/pfas or ckpts/lipid, run python tools/eval.py --log_path [ckpts/pfas or ckpts/lipid]: "Download the model weights in ckpts/pfas or ckpts/lipid, run
python tools/eval.py --log_path [ckpts/pfas or ckpts/lipid]
```"
- [other] Execute inference on each spectrum using the MSGO model to generate predicted molecular structures. Collect and format the structure-generation predictions and performance metrics into a results file.: "Execute inference on each spectrum using the MSGO model to generate predicted molecular structures. Collect and format the structure-generation predictions and performance metrics into a results file."
- [readme] Then you can obatin a results csv file inluding top 10 predicts.: "Then you can obatin a results csv file inluding top 10 predicts."
- [readme] For Training, we use 30k+ pseudo smiles-specturm pairs generated by cfmid. For evaluation, we use 300+ real specturm to verify our method: "For Training, we use 30k+ pseudo smiles-specturm pairs generated by cfmid. For evaluation, we use 300+ real specturm to verify our method"

What ships with it

Read from the repository

Just SKILL.md. No reference files, no scripts.

Keep looking

Skills are one crate of 326,546. Ordering is by how many stacks a row turns up in, so the top of any crate is what has actually been picked rather than what has the most stars.