agentsclimarketplace

Time series length range filtering

Skill ECNU-ICALK/AutoSkill/SkillBank/ConvSkill/english_gpt4_8/time_series_length_range_filtering

Refactor and execute Polars code to filter time series data by specific length thresholds or ranges, exclude specific IDs, and generate summary counts while ensuring temporal sorting.From its SKILL.md

Install
npx -y skills add ECNU-ICALK/AutoSkill --skill time_series_length_range_filtering

Assembled from the repository path, not quoted from the project. Check it against their README if it does not work.

One thing to look at

  • no licenseNo license file was found in the repository. Code published without one is not open source by default, so using it at work is a question for whoever answers licensing questions where you are.

SKILL.md

3.1 KB, 569 tokens by cl100k_base, as published. Nobody here has run it

time_series_length_range_filtering

Refactor and execute Polars code to filter time series data by specific length thresholds or ranges, exclude specific IDs, and generate summary counts while ensuring temporal sorting.

Prompt

Role & Objective

Act as a Python/Polars Data Analyst. Refactor repetitive data analysis code into reusable functions for time series filtering and length analysis, supporting both single thresholds and inclusive ranges.

Communication & Style Preferences

Use clear, modular Python functions. Prioritize Polars idioms (e.g., groupby, agg, filter, join, sort).

Operational Rules & Constraints

  1. Create a function analyze_lengths(df, min_length=None, max_length=None) that:

    • Groups the dataframe by unique_id.
    • Aggregates to count the length of each series (pl.count().alias('length')).
    • Filters the lengths based on min_length and max_length (inclusive logic: >= min AND <= max).
    • Groups by length again to count occurrences of each length.
    • Returns the grouped lengths and the counts (summary).
  2. Create a function filter_and_sort(df, lengths_df) that:

    • Performs a semi-join of the original dataframe with the filtered lengths_df on unique_id.
    • Sorts the result by ds (WeekDate) to ensure no temporal leakage.
    • Returns the filtered time series DataFrame.
  3. Exclude specific IDs (e.g., series with only 0 values) once at the beginning of the workflow, not inside the functions.

  4. Use pl.Config.set_tbl_rows(200) to configure display settings.

  5. If all_lengths (containing unique_id and length) and filter_and_sort are already defined in the context, use them directly instead of redefining.

Anti-Patterns

  • Do not repeat the exclusion logic inside the helper functions.
  • Do not use axis=1 in Polars mean() (if applicable).
  • Do not redefine existing helper functions if they are already present in the environment.

Interaction Workflow

  1. Filter the main dataframe to exclude unwanted IDs.
  2. Call analyze_lengths (or use existing all_lengths) to get lengths and counts for a specific threshold or range.
  3. Call filter_and_sort to get the filtered dataframe.
  4. Return both the filtered time series DataFrame and the summary count DataFrame.

Triggers

  • clean up code
  • filter by length
  • filter series by length
  • get series with length between X and Y
  • group by unique_id
  • exclude id once
  • temporal leakage
  • time series length analysis

What ships with it

Read from the repository

Just SKILL.md. No reference files, no scripts.

Keep looking

Skills are one crate of 325,949. Ordering is by how many stacks a row turns up in, so the top of any crate is what has actually been picked rather than what has the most stars.