agentsclimarketplace

Run2 glm calibration strategy

Skill cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-opus-4-6/temperature-simulation/run2_glm-calibration-strategy

[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.

Install
npx -y skills add cxcscmu/SkillLearnBench --skill run2_glm-calibration-strategy

Assembled from the repository path, not quoted from the project. Check it against their README if it does not work.

What its author says it does

Copied from the file, not written here

Two-phase grid search calibration for GLM parameters with parameter sensitivity analysis and regex-based nml editing.

SKILL.md

2.5 KB, as published. Nobody here has run it

GLM Calibration Strategy

Overview

Calibrate GLM by minimizing RMSE between simulated and observed temperature profiles using a two-phase grid search.

Phase 1: Coarse Search (Most Impactful Parameters)

Test 3-4 values per parameter for Kw, coef_mix_hyp, wind_factor. Fix lw_factor=1.0, ch=0.0013.

for kw in [0.20, 0.25, 0.30, 0.35]:
    for cmh in [0.4, 0.5, 0.6]:
        for wf in [0.85, 0.90, 0.95, 1.0]:
            # ~48 runs, ~2 min each = ~96 min

Phase 2: Fine-Tune Secondary Parameters

Around best Phase 1 point, vary lw_factor and ch:

for lw in [0.93, 0.95, 0.97, 1.0, 1.05]:
    for ch in [0.0010, 0.0012, 0.0013, 0.0015]:

Phase 3 (Optional): Narrow Refinement

Small grid ±0.02 around best values across all 5 params:

# Keep to ~100 combinations max for reasonable runtime

Safe NML Parameter Editing

import re

def set_param(nml_text, param, value):
    """Replace a parameter value in GLM nml text.

    Note: For 'ch', need to be careful not to match 'catchrain'.
    The pattern matches 'param = value' at any position.
    """
    pattern = rf'({param}\s*=\s*)[\d.eE+-]+'
    return re.sub(pattern, rf'\g<1>{value}', nml_text)

Important: Always read the original nml, apply ALL params fresh, then write. Don't accumulate edits.

Scoring Function

Minimize total RMSE across all targets:

score = overall_rmse + annual_deep_rmse + summer_deep_rmse

Lake Mendota Calibrated Values

ParameterDefaultCalibratedEffect
Kw0.300.35Increased extinction → cooler deep water
coef_mix_hyp0.500.45Slightly less hypolimnetic mixing
wind_factor1.000.88Reduced wind → stronger stratification
lw_factor1.000.93Reduced LW radiation → overall cooling
ch0.00130.0012Reduced sensible heat transfer

Results: O=1.34, AD=1.10, SD=1.07 (all well below thresholds of 1.60, 1.55, 1.70)

Key Lessons

  • wind_factor < 1 and Kw slightly above default are the biggest improvements for deep metrics
  • lw_factor < 1 provides overall cooling that helps all metrics
  • The parameter space has multiple local minima; grid search is more robust than gradient descent here
  • Each GLM run takes ~30-60 seconds; budget ~2-3 hours for thorough calibration

Keep looking

Skills are one crate of 328,083. Ordering is by how many stacks a row turns up in, so the top of any crate is what has actually been picked rather than what has the most stars.