Run1 mars cloud data grouping
Organizing annotation data by image identifiers to prepare for image-by-image clustering analysis.From its SKILL.md
npx -y skills add cxcscmu/SkillLearnBench --skill run1_mars-cloud-data-groupingAssembled from the repository path, not quoted from the project. Check it against their README if it does not work.
SKILL.md
0.9 KB, 201 tokens by cl100k_base, as published. Nobody here has run it
The dataset uses a file_rad identifier to link citizen science observations to expert labels. Since evaluations must happen per image:
- Load CSVs using
pandas. - Pre-group the DataFrames using
.groupby('file_rad')and store them in a dictionary for fast lookup. - Ensure that the evaluation loop iterates over all
file_radvalues present in the expert dataset to account for missed detections (FNs).
import pandas as pd
citsci = pd.read_csv('citsci_train.csv')
expert = pd.read_csv('expert_train.csv')
# Grouping for efficient access
cit_groups = {name: group[['x', 'y']].values for name, group in citsci.groupby('file_rad')}
exp_groups = {name: group[['x', 'y']].values for name, group in expert.groupby('file_rad')}
all_image_ids = expert['file_rad'].unique()
What ships with it
Read from the repository
Just SKILL.md. No reference files, no scripts.