agentsclimarketplace

Run1 mars cloud data grouping

Skill cxcscmu/SkillLearnBench/skills/b3-teacher-feedback-gemini-3-flash-preview/dbscan-parameter-tuning/run1_mars-cloud-data-grouping

Organizing annotation data by image identifiers to prepare for image-by-image clustering analysis.From its SKILL.md

Install
npx -y skills add cxcscmu/SkillLearnBench --skill run1_mars-cloud-data-grouping

Assembled from the repository path, not quoted from the project. Check it against their README if it does not work.

SKILL.md

0.9 KB, 201 tokens by cl100k_base, as published. Nobody here has run it

The dataset uses a file_rad identifier to link citizen science observations to expert labels. Since evaluations must happen per image:

  1. Load CSVs using pandas.
  2. Pre-group the DataFrames using .groupby('file_rad') and store them in a dictionary for fast lookup.
  3. Ensure that the evaluation loop iterates over all file_rad values present in the expert dataset to account for missed detections (FNs).
import pandas as pd

citsci = pd.read_csv('citsci_train.csv')
expert = pd.read_csv('expert_train.csv')

# Grouping for efficient access
cit_groups = {name: group[['x', 'y']].values for name, group in citsci.groupby('file_rad')}
exp_groups = {name: group[['x', 'y']].values for name, group in expert.groupby('file_rad')}

all_image_ids = expert['file_rad'].unique()

What ships with it

Read from the repository

Just SKILL.md. No reference files, no scripts.

Keep looking

Skills are one crate of 325,949. Ordering is by how many stacks a row turns up in, so the top of any crate is what has actually been picked rather than what has the most stars.