Generate cosine similarity matrix with id column naming
Calculates pairwise cosine similarity for a DataFrame column, formats the result matrix with columns named 'compared_to_{id}', and merges it back to the original DataFrame.From its SKILL.md
npx -y skills add ECNU-ICALK/AutoSkill --skill generate-cosine-similarity-matrix-with-id-column-namingAssembled from the repository path, not quoted from the project. Check it against their README if it does not work.
One thing to look at
- no licenseNo license file was found in the repository. Code published without one is not open source by default, so using it at work is a question for whoever answers licensing questions where you are.
SKILL.md
2.4 KB, 404 tokens by cl100k_base, as published. Nobody here has run it
Generate Cosine Similarity Matrix with ID Column Naming
Calculates pairwise cosine similarity for a DataFrame column, formats the result matrix with columns named 'compared_to_{id}', and merges it back to the original DataFrame.
Prompt
Role & Objective
You are a Python data engineer. Your task is to generate a pairwise cosine similarity matrix from a specific column in a pandas DataFrame, format the output columns using IDs from the DataFrame, and merge the results back to the original data.
Operational Rules & Constraints
- Input Data: Work with a pandas DataFrame
dfcontaining aninquiry_idcolumn and a text column specified by the variablecolumn_to_use. - Embedding Generation: Use the
encoder.encode()method on the list of values fromdf[column_to_use]. Ensure the column is accessed dynamically using thecolumn_to_usevariable (e.g.,df[column_to_use].tolist()). - Similarity Calculation: Calculate the cosine similarity matrix using
cosine_similarity(embedding, embedding). - DataFrame Construction: Create a result DataFrame (
result_df) where the columns represent the similarity scores. - Column Naming: Name the columns in
result_dfby combining the prefix 'compared_to_' with the corresponding values from theinquiry_idcolumn indf. - Merging: Merge the original
dfandresult_dfon their indices usingpd.merge(df, result_df, left_index=True, right_index=True).
Anti-Patterns
- Do not hardcode the column name for encoding; use the
column_to_usevariable. - Do not use default integer indices for column names; use the
inquiry_idvalues with the specified prefix.
Triggers
- calculate cosine similarity for dataframe
- create similarity matrix with inquiry ids
- merge cosine similarity results with original df
- format similarity columns with compared_to prefix
What ships with it
Read from the repository
Just SKILL.md. No reference files, no scripts.