agentsclimarketplace

Qsar modeling

Skill BioTender-max/awesome-bio-agent-skills/skills/bioskills/qsar-modeling

Builds QSAR / QSPR models using chemprop D-MPNN, MolFormer, Uni-Mol, ChemBERTa, random forest baselines, and Gaussian processes with explicit handling of OECD 5 principles, applicability domain (kNN, leverage, conformal prediction, Mahalanobis), scaffold-balanced splits, ensemble uncertainty, calibration (Platt, isotonic), feature importance (SHAP, atomic attribution), and prospective validation. Use when building target-specific predictive models from in-house bioassay data, ADMET endpoints, or selectivity profiles.From its SKILL.md

Install
npx -y skills add BioTender-max/awesome-bio-agent-skills --skill qsar-modeling

Assembled from the repository path, not quoted from the project. Check it against their README if it does not work.

One thing to look at

  • no licenseNo license file was found in the repository. Code published without one is not open source by default, so using it at work is a question for whoever answers licensing questions where you are.

SKILL.md

17.3 KB, ~4.2k tokens by cl100k_base, as published. Nobody here has run it

Version Compatibility

Reference examples tested with: chemprop 2.0+ (major API change from 1.x), RDKit 2024.09+, scikit-learn 1.4+, MAPIE 0.8+ (conformal prediction), shap 0.44+, pytorch 2.1+.

Before using code patterns, verify installed versions match. If versions differ:

  • Python: pip show <package> then help(module.function) to check signatures
  • CLI: chemprop train --help (chemprop 2.x); chemprop_train --help (1.x legacy)

If code throws ImportError, AttributeError, or TypeError, introspect the installed package and adapt the example to match the actual API rather than retrying.

QSAR Modeling

Build quantitative structure-activity relationship models from molecular structure inputs. The choice of model + featurization + split strategy determines whether the model captures real chemical signal or memorizes the training data. chemprop D-MPNN with optional Morgan / RDKit descriptors is the modern open-source standard; transformer-based methods (MolFormer, Uni-Mol, ChemBERTa) compete on benchmarks but offer minimal practical gain when domain-specific data is sparse. The OECD 5 principles structure the model for regulatory acceptance: defined endpoint, unambiguous algorithm, defined applicability domain, validation, and mechanistic interpretation.

For descriptor/fingerprint choices, see chemoinformatics/molecular-descriptors. For ADMET-specific QSAR, see chemoinformatics/admet-prediction. For molecular standardization (critical upstream), see chemoinformatics/molecular-standardization.

Model Taxonomy

ModelArchitectureUse caseFails when
Random Forest + ECFP4Classical baselineSmall data (<200 compounds), interpretabilitySaturates at ~AUC 0.85 for complex endpoints
chemprop D-MPNNDirected message passingModern default; 100-10k compoundsVery small datasets (<100) overfits
chemprop D-MPNN + RDKit 2DHybrid graph + descriptorshERG SOTA (Cai 2024 AUC 0.956)Diminishing returns at large data
MolFormerTransformer (87M params)Large public training data benefitCompute overhead; OOD risk
Uni-Mol3D-aware transformer3D-relevant endpoints (binding)Requires 3D conformers
ChemBERTa-2Transformer (77M params)SMILES language modelLarge training data needed
Gaussian Process + ECFP4ProbabilisticActive learning; uncertaintyO(N^3) scaling
MultiTask DNNJoint trainingMultiple endpointsData must overlap
AttentiveFPGNN with attentionPre-chemprop SOTALess actively maintained
KPGT / Graphormer / GROVERPretrained GNNPretraining liftDiminishing returns

Decision: For 200-10k compounds, chemprop 2.0 D-MPNN + RDKit 2D descriptors is the modern standard. For <200 compounds, Random Forest + ECFP4 is competitive and more interpretable.

Decision Tree by Scenario

Dataset sizeEndpoint typeModel
<50AnythingDon't model; use as test set for literature models
50-200Regression / classificationRF + ECFP4 + cross-validation
200-1kRegressionchemprop D-MPNN + RDKit 2D, 5-fold
1k-10kAnythingchemprop D-MPNN + RDKit 2D ensemble
10k-100kClassificationchemprop ensemble OR MolFormer fine-tuning
>100kAnythingMolFormer / Uni-Mol pretrained + LoRA fine-tune
Multi-taskRelated endpoints (CYP3A4, CYP2D6, etc.)chemprop MultiTask
3D-relevantBinding, conformer-dependentUni-Mol with conformer ensemble

OECD 5 Principles

For regulatory use (REACH, ECHA, FDA submissions):

  1. Defined endpoint: specific bioassay, units, threshold definitions
  2. Unambiguous algorithm: reproducible code, fixed random seeds, version-pinned dependencies
  3. Defined applicability domain (AD): where the model is valid
  4. Appropriate statistical validation: external test set, cross-validation
  5. Mechanistic interpretation: biological/chemical rationale where possible

For non-regulatory QSAR, all 5 still good practice; especially AD definition is critical.

Applicability Domain Methods

MethodDefinitionProCon
Ensemble variance (RECOMMENDED for chemprop)Std across N-model ensemble predictionsBuilt-in to chemprop predict; no extra codeAssumes ensemble diversity (random seeds give independent fits)
kNN distanceMean Tanimoto to k nearest in trainingEasy to interpretDoesn't account for label distribution
LeverageHat matrix diagonalStatisticalLinear assumptions
KDE on PCADensity in feature spaceCaptures multivariate structureDensity choice subjective
Mahalanobis distanceCovariance-aware distanceTheoretically motivatedHigh-dim instability
Conformal predictionPer-prediction confidence intervalDistribution-free guaranteeRequires calibration set + sklearn-compatible estimator
Bayesian / MC-dropoutPosterior or dropout varianceDirect uncertaintyComputational cost
Tanimoto coverageAt least 1 NN within thresholdPracticalThreshold subjective

Operational rule: For chemprop, set --ensemble-size 5 at training; at predict time, flag predictions with ensemble std > P95 of the training-set ensemble std distribution as out-of-AD. This is the standard chemprop-native AD measure and avoids the overhead of wrapping with MAPIE/conformal.

chemprop 2.0 Training (CLI)

Goal: Train a 5-fold x 5-model chemprop D-MPNN ensemble with RDKit 2D descriptor features and scaffold-balanced split.

Approach: Invoke chemprop train with --molecule-featurizers rdkit_2d_normalized, --num-folds 5, --ensemble-size 5, and --split scaffold_balanced to produce 25 models for ensemble mean prediction + variance-based applicability domain.

# chemprop 2.x CLI (current): use 'chemprop train' (space; dashes not underscores)
chemprop train \
    --data-path data.csv \
    --task-type classification \
    --save-dir model_dir \
    --molecule-featurizers rdkit_2d_normalized \
    --num-folds 5 \
    --ensemble-size 5 \
    --epochs 50 \
    --batch-size 128 \
    --split scaffold_balanced \
    --split-sizes 0.8 0.1 0.1 \
    --metric roc

# chemprop 1.x legacy CLI (for backwards reference):
# chemprop_train --data_path data.csv --dataset_type classification ...

Key flags (chemprop 2.x):

  • --molecule-featurizers rdkit_2d_normalized: include physchem descriptors (was --features_generator in 1.x)
  • --num-folds 5: 5-fold cross-validation
  • --ensemble-size 5: 5 models per fold for uncertainty
  • --split scaffold_balanced: prevent scaffold leakage (was --split_type in 1.x)
  • --split-sizes 0.8 0.1 0.1: 80/10/10 train/val/test

Total models: 25 (5 folds x 5 ensemble). Use ensemble mean as prediction; ensemble std as uncertainty.

Scaffold-Balanced Split (chemprop 2.0 default)

Goal: Partition a SMILES dataset into train/val/test such that no Bemis-Murcko scaffold appears in more than one split (prevents chemotype leakage).

Approach: Group compounds by scaffold and greedily assign whole scaffolds to train until target fraction, then val, then test. chemprop's --split scaffold_balanced implements this with optional class-stratification.

scaffold_balanced split assigns entire scaffolds to one of train / val / test, preventing chemotype leakage. Modern QSAR uses scaffold split exclusively.

Split typeRandomScaffold balanced
Train AUC0.990.95
Test AUC0.950.75-0.85
GeneralizationOptimisticRealistic

The gap between random and scaffold split is the true generalization gap; report both.

Conformal Prediction for Calibrated Uncertainty

Use conformal prediction when calibrated coverage guarantees matter (regulatory submission, decision thresholds with cost asymmetry). For routine QSAR, chemprop ensemble variance is simpler and sufficient.

# MAPIE expects a scikit-learn-compatible estimator (.fit / .predict / .predict_proba).
# chemprop 2.x is NOT scikit-learn-compatible out of the box -- either wrap chemprop
# in a thin sklearn estimator class or use MAPIE only with the sklearn baseline.
from mapie.regression import MapieRegressor
from sklearn.ensemble import RandomForestRegressor

base = RandomForestRegressor(n_estimators=500, random_state=42)
mapie = MapieRegressor(estimator=base, method='plus', cv=5)
mapie.fit(X_train, y_train)
y_pred, y_intervals = mapie.predict(X_test, alpha=0.1)  # alpha=0.1 -> 90% coverage

Alpha choice: 0.05 (95% coverage) for safety endpoints; 0.10 (90%) for activity ranking. For chemprop, the simpler ensemble-variance route (set --ensemble-size 5 and use prediction std as AD signal) is preferred unless conformal coverage guarantees are specifically required.

SHAP / Atomic Attribution

For mechanistic interpretation:

For a scikit-learn-style model (e.g., Random Forest baseline on ECFP4), SHAP integrates directly:

import shap
from sklearn.ensemble import RandomForestClassifier

# X_train / X_test are Morgan fingerprint arrays (n_samples, n_bits)
model = RandomForestClassifier(n_estimators=500, random_state=42).fit(X_train, y_train)
explainer = shap.TreeExplainer(model)
shap_values = explainer.shap_values(X_test)

# Per-bit contribution; for atomic interpretation, map bits back to
# generating atoms via AllChem.GetMorganFingerprintAsBitVect(mol, ..., bitInfo=bi)
# and aggregate SHAP across all bits triggered by each atom.

For chemprop D-MPNN, SHAP requires a custom wrapper (chemprop is not sklearn-compatible). Use shap.GradientExplainer on the underlying PyTorch module after exposing it via a wrapper class. In practice, chemprop's built-in atom attribution (chemprop predict --uncertainty-method classification with atom-level outputs) is the supported route for atomic interpretation. Use SHAP only when comparing chemprop to a classical baseline.

Bayesian Optimization for Active Learning

from sklearn.gaussian_process import GaussianProcessRegressor

gp = GaussianProcessRegressor(kernel='rbf')
gp.fit(X_train, y_train)
mu, sigma = gp.predict(X_pool, return_std=True)

# Expected Improvement
def expected_improvement(mu, sigma, y_best, xi=0.01):
    z = (mu - y_best - xi) / sigma
    return (mu - y_best - xi) * norm.cdf(z) + sigma * norm.pdf(z)

ei = expected_improvement(mu, sigma, y_train.max())
next_to_test = X_pool[ei.argmax()]

For chemprop + active learning, replace GP with chemprop ensemble + ensemble variance.

Calibration (Platt / Isotonic)

Deep learning native probabilities rarely well-calibrated. Apply Platt scaling on validation set:

from sklearn.calibration import CalibratedClassifierCV
from sklearn.ensemble import RandomForestClassifier

# For sklearn-compatible models, calibration is direct:
base = RandomForestClassifier(n_estimators=500, random_state=42).fit(X_train, y_train)
calibrated = CalibratedClassifierCV(base, method='isotonic', cv='prefit')
calibrated.fit(X_val, y_val)
y_proba_calibrated = calibrated.predict_proba(X_test)

Use Platt (logistic) for binary; isotonic for general. For chemprop, native calibration is applied automatically when training with classification + --metric roc; for further re-calibration, save validation probabilities and fit isotonic regression externally:

from sklearn.isotonic import IsotonicRegression
iso = IsotonicRegression(out_of_bounds='clip').fit(val_chemprop_probs, val_true)
test_calibrated = iso.predict(test_chemprop_probs)

Multi-Task QSAR

Train multiple related endpoints jointly:

df = pd.DataFrame({
    'smiles': [...],
    'CYP1A2_inhibition': [...],
    'CYP2D6_inhibition': [...],
    'CYP3A4_inhibition': [...],
})
df.to_csv('multitask.csv', index=False)
chemprop train --data-path multitask.csv --task-type classification \
               --target-columns CYP1A2_inhibition CYP2D6_inhibition CYP3A4_inhibition \
               --save-dir multitask_model

Joint training improves performance when endpoints are correlated (Wu 2018 MoleculeNet). For independent endpoints, multitask hurts.

Per-Tool Failure Modes

Random split for QSAR

Trigger: Default sklearn train_test_split.

Mechanism: Compounds from same scaffold scatter across train/test; performance optimistic.

Symptom: Cross-validation AUC > 0.95; real-world performance much lower.

Fix: Use --split scaffold_balanced in chemprop 2.x (or --split_type scaffold_balanced in chemprop 1.x legacy); or scaffold_split from chemoinformatics/scaffold-analysis.

Class imbalance not handled

Trigger: 10:1 negative:positive ratio in dataset.

Mechanism: Default loss treats classes equally; model learns majority class.

Symptom: High accuracy but precision/recall on minority class poor.

Fix: Class-weighted loss; SMOTE; or report AUC/F1 not accuracy.

Over-engineered features

Trigger: Including hundreds of descriptors (e.g., rdkit_2d not normalized).

Mechanism: Some descriptors dominate scaling; model overfits.

Symptom: Validation performance differs widely across runs; high feature importance noise.

Fix: Use rdkit_2d_normalized; or restrict to <50 standard descriptors.

Missing AD assessment

Trigger: Predicting on novel chemotypes without AD check.

Mechanism: Model extrapolates; predictions unreliable.

Symptom: Confident predictions but actual values different.

Fix: Always compute kNN Tanimoto distance + ensemble variance; flag low-confidence predictions.

chemprop 1.x vs 2.x confusion

Trigger: Code/tutorial from before late 2024.

Mechanism: Major API change: chemprop_train → chemprop train; Python API redesigned.

Symptom: ImportError or different keyword arguments.

Fix: Use chemprop --version; check 2.x documentation; migrate APIs.

Pretrained Transformer overhead without data benefit

Trigger: Using MolFormer on <500 compounds.

Mechanism: Pretrained representation needs sufficient fine-tuning data to specialize.

Symptom: No improvement over chemprop; slower training.

Fix: For <500 compounds, stick with chemprop or RF. Use pretrained Transformers for >10k.

Validation leakage via standardization

Trigger: Different standardization for train vs test.

Mechanism: Train compounds with R-isomer; test compounds with S-isomer not in train.

Symptom: Test performance better than realistic.

Fix: Standardize unified preprocessing; same canonicalization train + test.

Reconciliation: Classical RF vs chemprop vs Transformer

AspectRF + ECFP4chemprop D-MPNNMolFormer
Data size sweet spot50-1k200-10k>10k
InterpretabilityHigh (Gini importance)Medium (SHAP)Low
UncertaintyBootstrapEnsembleMC-dropout
HardwareCPUGPU recommendedGPU mandatory
OOD performanceMost degradedBetter with ensembleWorst (long-range memorization)
Production deploymentEasy (pickle)OK (ONNX export)Heavy

Common Errors

SymptomCauseFix
chemprop hangs at startGPU OOMReduce batch_size; check CUDA
All predictions same valueConstant targetStandardize labels
AUC mismatched across foldsRandom seed not set--seed 42
Test AUC = train AUCNo held-out dataUse scaffold_balanced split
Ensemble variance always smallAll models identicalSet --seed per fold; check randomness
SHAP fails on D-MPNNCustom architectureUse GradientExplainer not TreeExplainer
MolFormer fine-tune slowAll parameters trainedUse LoRA or freeze early layers
Calibration brokenPre-calibrated alreadySkip calibration step

References

  • Yang et al., J. Chem. Inf. Model. 59:3370 (2019) -- chemprop D-MPNN.
  • Heid et al., J. Chem. Inf. Model. 64:9 (2024) -- chemprop 2.0 redesign.
  • Wu et al., Chem. Sci. 9:513 (2018) -- MoleculeNet benchmark.
  • Ross et al. (2022) -- MoLFormer architecture.
  • Zhou et al., Nat. Commun. 14:3849 (2023) -- Uni-Mol 3D.
  • OECD, "OECD Principles for the Validation of QSAR Models" (2007).
  • OECD, "(Q)SAR Assessment Framework" (2023).
  • Cortés-Ciriano & Bender, J. Chem. Inf. Model. 60:1184 (2020) -- conformal prediction for QSAR applicability domain.
  • Alperstein et al., J. Cheminformatics 15:9 (2023) -- MolBERT chemistry-aware tokenization vs ChemBERTa-2.

Related Skills

  • chemoinformatics/molecular-descriptors - Featurization choices
  • chemoinformatics/molecular-standardization - Mandatory upstream
  • chemoinformatics/scaffold-analysis - Bemis-Murcko split implementation
  • chemoinformatics/admet-prediction - ADMET-specific QSAR
  • chemoinformatics/generative-design - QSAR as scoring component
  • machine-learning/model-validation - General ML validation principles
  • machine-learning/biomarker-discovery - Adjacent ML approaches

What ships with it: 2 files

4.3 KB alongside SKILL.md, 1 of them executable

examples/

Keep looking

Skills are one crate of 325,949. Ordering is by how many stacks a row turns up in, so the top of any crate is what has actually been picked rather than what has the most stars.