Run1 dbscan custom distance clustering
Implement DBSCAN clustering with a custom weighted Euclidean distance metric controlled by shape_weight parameter. Use this skill to cluster citizen science point annotations on Mars cloud images.From its SKILL.md
npx -y skills add cxcscmu/SkillLearnBench --skill run1_dbscan-custom-distance-clusteringAssembled from the repository path, not quoted from the project. Check it against their README if it does not work.
SKILL.md
1.5 KB, 338 tokens by cl100k_base, as published. Nobody here has run it
Overview
DBSCAN (Density-Based Spatial Clustering of Applications with Noise) is a clustering algorithm that groups points based on density. This implementation uses a custom distance metric that weights x and y dimensions asymmetrically.
Custom Distance Metric
The distance between points a and b is:
d(a, b) = sqrt((w * Δx)² + ((2 - w) * Δy)²)
Where:
w=shape_weightparameter (0.9–1.9)Δx= difference in x-coordinatesΔy= difference in y-coordinates
Interpretation:
- When
w = 1.0: Standard Euclidean distance - When
w > 1.0: y-distances are attenuated (points closer in y-direction are grouped together more easily) - When
w < 1.0: x-distances are attenuated (points closer in x-direction are grouped together more easily)
DBSCAN Parameters
- epsilon (eps): Maximum distance between two points for them to be in the same neighborhood
- min_samples: Minimum number of points in a neighborhood for a point to be considered a core point
Implementation Notes
- Use scikit-learn's DBSCAN with a custom metric function or distance matrix
- Compute cluster centroids as the mean (x, y) of all points in each cluster
- Ignore noise points (label = -1) when computing centroids
- Return only valid clusters (at least 1 cluster found)
What ships with it
Read from the repository
Just SKILL.md. No reference files, no scripts.