agentsclimarketplace

Run3 compute metrics

Skill cxcscmu/SkillLearnBench/skills/b3-teacher-feedback-gemini-3.1-flash-lite-preview/temperature-simulation/run3_compute_metrics

Calculates RMSE between simulated and observed data using exact datetime matching and integer depth binning based on simulation start date 2009-01-01.From its SKILL.md

Install
npx -y skills add cxcscmu/SkillLearnBench --skill run3_compute_metrics

Assembled from the repository path, not quoted from the project. Check it against their README if it does not work.

SKILL.md

1.1 KB, 232 tokens by cl100k_base, as published. Nobody here has run it

When calculating metrics, convert the model time (relative days) to datetime objects starting from 2009-01-01. For each (datetime, rounded_depth) bin, if multiple simulated values exist, calculate their mean before matching with observations.

import pandas as pd
import numpy as np
import xarray as xr

def calculate_metrics(nc_path, obs_csv, lake_depth):
    ds = xr.open_dataset(nc_path)
    obs = pd.read_csv(obs_csv, parse_dates=['datetime'])
    
    # Convert GLM time to datetime
    start_date = pd.Timestamp('2009-01-01')
    ds['datetime'] = start_date + pd.to_timedelta(ds.time.values, unit='D')
    
    # Process layers: depth = lake_depth - z
    # Group by datetime and round(depth)
    # Calculate mean for bins with multiple simulated layers
    # Perform inner merge with obs on (datetime, rounded_depth)
    # Filter for annual_deep (depth >= 13) and summer_deep (month in [6,7,8,9] & depth >= 13)

What ships with it

Read from the repository

Just SKILL.md. No reference files, no scripts.

Keep looking

Skills are one crate of 325,949. Ordering is by how many stacks a row turns up in, so the top of any crate is what has actually been picked rather than what has the most stars.