Fuzzy string matching
A skill for finding best string matches in datasets using fuzzy matching libraries like fuzzywuzzy or rapidfuzz.From its SKILL.md
npx -y skills add cxcscmu/SkillLearnBench --skill fuzzy-string-matchingAssembled from the repository path, not quoted from the project. Check it against their README if it does not work.
SKILL.md
0.8 KB, 170 tokens by cl100k_base, as published. Nobody here has run it
Fuzzy String Matching
This skill demonstrates how to use fuzzy string matching to find company or fund names in a dataset.
Requirements
fuzzywuzzyorrapidfuzzpython-Levenshtein(optional but recommended for speed)
Installation
pip install rapidfuzz
Usage
import pandas as pd
from rapidfuzz import process, fuzz
df = pd.read_csv('data.tsv', sep='\t')
names = df['NAME_COLUMN'].dropna().unique()
# Find the best match
query = "Renaissance Technologies"
best_match = process.extractOne(query, names, scorer=fuzz.token_sort_ratio)
print(f"Best match: {best_match[0]} with score {best_match[1]}")
What ships with it
Read from the repository
Just SKILL.md. No reference files, no scripts.