agentsclimarketplace

Parallel processing

Skill cxcscmu/SkillLearnBench/skills/b4-skill-creator-claude-opus-4-6/dbscan-parameter-tuning/parallel-processing

Parallel processing with joblib for grid search and batch computations. Use when speeding up computationally intensive tasks across multiple CPU cores.From its SKILL.md

Install
npx -y skills add cxcscmu/SkillLearnBench --skill parallel-processing

Assembled from the repository path, not quoted from the project. Check it against their README if it does not work.

SKILL.md

2.1 KB, 502 tokens by cl100k_base, as published. Nobody here has run it

Parallel Processing with joblib

When to Use

Use this skill when you need to parallelize independent computations — grid search evaluations, batch processing, cross-validation folds, or any embarrassingly parallel workload.

Basic Pattern

from joblib import Parallel, delayed

def evaluate(param1, param2):
    """Single evaluation — must be a pure function."""
    # ... computation ...
    return result

# Generate all parameter combinations
from itertools import product
params = list(product(range_1, range_2))

# Run in parallel
results = Parallel(n_jobs=-1, verbose=1)(
    delayed(evaluate)(p1, p2) for p1, p2 in params
)

Key Parameters

  • n_jobs=-1: Use all available cores
  • n_jobs=-2: Use all cores minus one
  • verbose=1: Show progress bar
  • backend='loky' (default): Process-based, best for CPU-bound work
  • prefer='threads': Thread-based, better for I/O-bound or when sharing large read-only data

Sharing Read-Only Data

For large datasets shared across workers, avoid copying by passing data outside the delayed call:

import numpy as np

# Large shared data — read by all workers
big_array = np.load('data.npy')

def process(idx, data=big_array):
    # data is shared, not copied (with loky backend)
    return data[idx].sum()

results = Parallel(n_jobs=-1)(delayed(process)(i) for i in range(1000))

Grid Search Pattern

from itertools import product

param_grid = {
    'eps': [4, 6, 8, 10],
    'min_samples': [3, 5, 7],
    'weight': [0.9, 1.0, 1.1]
}

combos = list(product(*param_grid.values()))

def evaluate_combo(combo):
    eps, min_samples, weight = combo
    # ... run clustering, compute metrics ...
    return {'eps': eps, 'min_samples': min_samples, 'weight': weight, 'score': score}

results = Parallel(n_jobs=-1, verbose=5)(
    delayed(evaluate_combo)(c) for c in combos
)

import pandas as pd
results_df = pd.DataFrame(results)

What ships with it

Read from the repository

Just SKILL.md. No reference files, no scripts.

Keep looking

Skills are one crate of 325,949. Ordering is by how many stacks a row turns up in, so the top of any crate is what has actually been picked rather than what has the most stars.