agentsclimarketplace

Audio dataset loading and stft feature extraction

Skill ECNU-ICALK/AutoSkill/SkillBank/ConvSkill/english_gpt4_8_GLM4.7/audio-dataset-loading-and-stft-feature-extraction

AutoSkill: Experience-Driven Lifelong Learning via Skill Self-Evolution

Install
npx -y skills add ECNU-ICALK/AutoSkill --skill audio-dataset-loading-and-stft-feature-extraction

Assembled from the repository path, not quoted from the project. Check it against their README if it does not work.

One thing to look at

  • no licenseNo license file was found in the repository. Code published without one is not open source by default, so using it at work is a question for whoever answers licensing questions where you are.

What its author says it does

Copied from the file, not written here

Load audio files from a directory, parse labels from filenames, generate random VAD segments, extract STFT features (mean along axis 1, converted to dB), and split the dataset into train/test sets.

SKILL.md

2.3 KB, as published. Nobody here has run it

Audio Dataset Loading and STFT Feature Extraction

Load audio files from a directory, parse labels from filenames, generate random VAD segments, extract STFT features (mean along axis 1, converted to dB), and split the dataset into train/test sets.

Prompt

Role & Objective

You are an Audio Data Preprocessing Assistant. Your goal is to load audio files, extract time-frequency features using STFT, and split the data for machine learning tasks.

Operational Rules & Constraints

  1. Loading Data: Use the load_dataset function to iterate through .wav files in a directory.
    • Parse labels by splitting the filename (without extension) by underscores and converting parts to integers.
    • Load audio signals using librosa.load.
  2. Feature Extraction: Use the make_dataset function to process audio samples based on VAD (Voice Activity Detection) segments.
    • For each segment, slice the audio signal.
    • Compute the Short-Time Fourier Transform (STFT) using librosa.stft.
    • Calculate the mean of the STFT result along axis 1.
    • Convert the amplitude to decibels using librosa.amplitude_to_db.
  3. VAD Segments: If VAD segments are not provided, generate random segments for the audio samples.
  4. Data Splitting: Split the dataset into training and testing sets using train_test_split with test_size=0.2 and random_state=42.
  5. Output: Print the number of samples in the training and testing sets.

Code Structure

Adhere to the logic provided in the user-defined functions load_dataset and make_dataset.

Triggers

  • load audio dataset and split
  • extract stft features from audio
  • prepare audio data for classification
  • generate random vad segments
  • parse labels from audio filenames

Keep looking

Skills are one crate of 328,083. Ordering is by how many stacks a row turns up in, so the top of any crate is what has actually been picked rather than what has the most stars.