agentsclimarketplace

Computational efficiency single pass vs repeated detection

Skill HolobiomicsLab/asb-skill-collections/collections/metabolomics/v2/skills/computational-efficiency-single-pass-vs-repeated-detection

Curated, evidence-grounded skill and software-tool collections for scientific AI agents, generated by the AgenticScienceBuilder

Install
npx -y skills add HolobiomicsLab/asb-skill-collections --skill computational-efficiency-single-pass-vs-repeated-detection

Assembled from the repository path, not quoted from the project. Check it against their README if it does not work.

One thing to look at

  • 14 stars14 stars. Stars are a popularity signal and not a quality one, but at this level it is likely that nobody has read this closely except its author, and you would be relying on your own review.

What its author says it does

Copied from the file, not written here

Use when when processing aligned LC-MS data across multiple samples where the computational bottleneck is repeated peak-detection algorithm calls (one per sample per m/z value). Typical scenario: >10 samples with >1000 m/z values each, where N individual find_peaks invocations dominate runtime.

The file declares its own license as CC-BY-4.0. That is the author’s claim about this one file, and it is not the same thing as the license GitHub reports for the repository, which is listed with the other numbers below.

SKILL.md

11.8 KB, ~2.2k tokens by cl100k_base, as published. Nobody here has run it

computational-efficiency-single-pass-vs-repeated-detection

License: restricted — no clear open-source license detected for the underlying tool; verify licensing before commercial use or redistribution. <!-- asb-license-banner -->

Summary

Replace per-sample peak detection with single-pass detection on a composite mass track (summed intensities across all samples) to reduce scipy.signal.find_peaks invocations from N to 1, dramatically improving computational scalability in untargeted LC-MS metabolomics without sacrificing peak quality metrics.

When to use

When processing aligned LC-MS data across multiple samples where the computational bottleneck is repeated peak-detection algorithm calls (one per sample per m/z value). Typical scenario: >10 samples with >1000 m/z values each, where N individual find_peaks invocations dominate runtime. Particularly valuable when samples are aligned to a common retention-time reference via rt_cal_dict calibration dictionaries and when composite signal is strong enough to enable robust peak detection (median intensity >1e3).

When NOT to use

  • Input is already a per-sample feature table or pre-aggregated peak list; skip composite construction.
  • Samples are poorly aligned or rt_cal_dict calibration fails; composite will be corrupted by misaligned noise.
  • Composite signal is too weak (median intensity <1e3) to reliably detect peaks; per-sample detection may recover low-abundance features.
  • Peak detection is not a computational bottleneck (e.g., small projects with <5 samples or <100 m/z values per sample).

Inputs

  • aligned mass tracks (intensity arrays indexed by m/z) from all N samples
  • retention-time calibration dictionaries (rt_cal_dict) mapping per-sample scan numbers to reference coordinates
  • mass grid (m/z values common across aligned samples)
  • detection parameters: min_peak_height, min_timepoints (default 25 scans), intensity ceiling (1E8), noise/smoothing thresholds

Outputs

  • composite mass tracks (summed-intensity arrays per m/z)
  • detected peaks with m/z, retention time, intensity, cSelectivity, SNR, gaussian fit metrics
  • audit metadata: baseline, noise, smoothing flags, detrend status per composite track
  • mapping of detected peaks back to constituent sample contributions (via cmap, mass_grid_mapping)

How to apply

After aligning and calibrating retention times across all samples via rt_cal_dict, synchronize scan-number indices to a reference sample coordinate system. Sum intensity values element-wise across all N samples for each m/z value to construct composite mass tracks (full-length intensity arrays). Audit each composite track: check intensity ceiling (1E8), detect and detrend if median >10× min_peak_height and >50% of points exceed threshold, compute baseline/noise from bottom-signal quartiles, apply moving-average smoothing if noise >1% of max intensity AND max intensity <10× min_peak_height. Subtract baseline+noise to isolate positive regions, split into segments, and invoke scipy.signal.find_peaks exactly once per segment (not N times) with dynamic prominence (max of 1/3 min_peak_height, noise level, 5% segment max). Filter detected peaks by gaussian fit goodness, chromatographic selectivity (cSelectivity), and SNR >2. This approach leverages increased signal-to-noise in the composite to enable single-pass detection while maintaining per-peak quality metrics traceable back to individual samples.

Related tools

Examples

```python
from asari.peaks import audit_mass_track, stats_detect_elution_peaks
from scipy.signal import find_peaks
import numpy as np

# Composite mass track from summed intensities across N aligned samples
composite_track = np.sum([sample_track for sample_track in aligned_samples], axis=0)

# Audit and preprocess
audit_result = audit_mass_track(composite_track, min_peak_height=10000)
processed_track = audit_result['processed_track']

# Single-pass peak detection on composite (N=1, not N invocations)
peaks, properties = find_peaks(
    processed_track,
    height=10000,
    prominence=max(3333, audit_result['noise_level'], 0.05*processed_track.max())
)

# Evaluate peak quality
for peak_idx in peaks:
    cSelectivity = compute_cSelectivity(processed_track, peak_idx)
    snr = compute_SNR(processed_track[peak_idx], audit_result['noise_level'])
    if snr > 2 and cSelectivity > 0:
        print(f'Peak at scan {peak_idx}: SNR={snr:.2f}, cSelectivity={cSelectivity:.3f}')

## Evaluation signals

- Peak detection invokes find_peaks exactly 1 time per composite mass track (audit log or code profiling confirms N→1 reduction).
- All detected peaks include non-null cSelectivity, SNR, and gaussian fit goodness metrics, verifying per-peak quality evaluation was applied.
- Composite track construction is verified: sum of per-sample intensities at each m/z matches element-wise sum of input mass tracks.
- Retention-time synchronization is correct: rt_cal_dict was applied to map all sample scan numbers to reference coordinates before summation (check alignment audit).
- Runtime is measured: wall-clock time for peak detection on composite is <1/N of wall-clock time for per-sample detection on the same m/z values (where N = number of samples).
- Output feature table reports peaks with cSelectivity >0 (indicating selectivity was evaluated) and SNR ≥2 threshold applied (per documented filter).

## Limitations

- Composite detection relies on strong summed signal; low-abundance features present in few samples may be lost if individual sample SNR is <2 but composite SNR >2 threshold is marginal.
- Retention-time calibration (rt_cal_dict) must succeed for all samples; failure in any sample corrupts the alignment and composite construction.
- Intensity ceiling (1E8) and noise-floor thresholds (median >1e3) are tuned for typical LC-MS data; extreme dynamic range or very noisy samples may require parameter retuning.
- Per-sample abundance variation is invisible in composite detection; peaks present in only 1–2 samples will not be highlighted during composite detection and must be recovered post-hoc if needed.
- Composite approach assumes additive intensity model; ion-suppression effects or sample-dependent chromatographic shifts are averaged out, potentially masking sample-specific phenomena.

## Evidence

- [other] Asari implements peak detection using scipy.signal.find_peaks on a composite map constructed from summed mass tracks across all samples, reducing the number of peak-detection algorithm calls from N (one per sample) to one (composite), thereby improving scalability and computational efficiency.: "peak detection using scipy.signal.find_peaks on a composite map constructed from summed mass tracks across all samples, reducing the number of peak-detection algorithm calls from N (one per sample)"
- [other] Sum intensity values element-wise across all samples for each unique m/z value in the mass grid, producing composite mass tracks as full-length intensity arrays.: "Sum intensity values element-wise across all samples for each unique m/z value in the mass grid, producing composite mass tracks as full-length intensity arrays."
- [other] Audit each composite mass track (peaks.audit_mass_track): check max intensity ceiling (1E8), apply rescaling if needed, detect low-intensity tracks (median < 1e3), perform detrend if median > 10× min_peak_height and >50% of points exceed threshold: "check max intensity ceiling (1E8), detect low-intensity tracks (median < 1e3), perform detrend if median > 10× min_peak_height and >50% of points exceed threshold"
- [other] Apply scipy.signal.find_peaks on each segment with dynamic prominence (max of 1/3 min_peak_height, noise level, and 5% of segment max intensity), using min_peak_height and min_timepoints (default 25 scans) for window control.: "Apply scipy.signal.find_peaks on each segment with dynamic prominence (max of 1/3 min_peak_height, noise level, and 5% of segment max intensity)"
- [other] Evaluate detected peaks for gaussian fit (goodness_fitting), chromatographic selectivity (cSelectivity), and signal-to-noise ratio (SNR > 2), retaining only peaks meeting thresholds.: "Evaluate detected peaks for gaussian fit (goodness_fitting), chromatographic selectivity (cSelectivity), and signal-to-noise ratio (SNR > 2), retaining only peaks meeting thresholds."
- [intro] Peak detection on a composite map instead of repeated on individual samples: "Peak detection on a composite map instead of repeated on individual samples"
- [other] Synchronize scan-number indices across samples by applying rt_cal_dict to map each sample's scan numbers to reference sample coordinates.: "Synchronize scan-number indices across samples by applying rt_cal_dict to map each sample's scan numbers to reference sample coordinates."

What ships with it

Read from the repository

Just SKILL.md. No reference files, no scripts.

Keep looking

Skills are one crate of 328,083. Ordering is by how many stacks a row turns up in, so the top of any crate is what has actually been picked rather than what has the most stars.