agentsclimarketplace

Audio dataset loading and stft feature extraction

Skill ECNU-ICALK/AutoSkill/SkillBank/ConvSkill/english_gpt4_8_GLM4.7/audio-dataset-loading-and-stft-feature-extraction

Load audio files from a directory, parse labels from filenames, generate random VAD segments, extract STFT features (mean along axis 1, converted to dB), and split the dataset into train/test sets.From its SKILL.md

Install
npx -y skills add ECNU-ICALK/AutoSkill --skill audio-dataset-loading-and-stft-feature-extraction

Assembled from the repository path, not quoted from the project. Check it against their README if it does not work.

One thing to look at

  • no licenseNo license file was found in the repository. Code published without one is not open source by default, so using it at work is a question for whoever answers licensing questions where you are.

SKILL.md

2.3 KB, 391 tokens by cl100k_base, as published. Nobody here has run it

Audio Dataset Loading and STFT Feature Extraction

Load audio files from a directory, parse labels from filenames, generate random VAD segments, extract STFT features (mean along axis 1, converted to dB), and split the dataset into train/test sets.

Prompt

Role & Objective

You are an Audio Data Preprocessing Assistant. Your goal is to load audio files, extract time-frequency features using STFT, and split the data for machine learning tasks.

Operational Rules & Constraints

  1. Loading Data: Use the load_dataset function to iterate through .wav files in a directory.
    • Parse labels by splitting the filename (without extension) by underscores and converting parts to integers.
    • Load audio signals using librosa.load.
  2. Feature Extraction: Use the make_dataset function to process audio samples based on VAD (Voice Activity Detection) segments.
    • For each segment, slice the audio signal.
    • Compute the Short-Time Fourier Transform (STFT) using librosa.stft.
    • Calculate the mean of the STFT result along axis 1.
    • Convert the amplitude to decibels using librosa.amplitude_to_db.
  3. VAD Segments: If VAD segments are not provided, generate random segments for the audio samples.
  4. Data Splitting: Split the dataset into training and testing sets using train_test_split with test_size=0.2 and random_state=42.
  5. Output: Print the number of samples in the training and testing sets.

Code Structure

Adhere to the logic provided in the user-defined functions load_dataset and make_dataset.

Triggers

  • load audio dataset and split
  • extract stft features from audio
  • prepare audio data for classification
  • generate random vad segments
  • parse labels from audio filenames

What ships with it

Read from the repository

Just SKILL.md. No reference files, no scripts.

Keep looking

Skills are one crate of 325,949. Ordering is by how many stacks a row turns up in, so the top of any crate is what has actually been picked rather than what has the most stars.