Average session step duration analyzer
Skill ECNU-ICALK/AutoSkill/SkillBank/ConvSkill/english_gpt4_8/average_session_step_duration_analyzer
Computes the average duration of actions or steps from session logs using Python (pseudo-streaming) or SQL. Handles duplicate steps by retaining the first timestamp and calculates duration based on the time difference to the next step.From its SKILL.md
npx -y skills add ECNU-ICALK/AutoSkill --skill average_session_step_duration_analyzerAssembled from the repository path, not quoted from the project. Check it against their README if it does not work.
One thing to look at
- no licenseNo license file was found in the repository. Code published without one is not open source by default, so using it at work is a question for whoever answers licensing questions where you are.
SKILL.md
3.3 KB, 611 tokens by cl100k_base, as published. Nobody here has run it
average_session_step_duration_analyzer
Computes the average duration of actions or steps from session logs using Python (pseudo-streaming) or SQL. Handles duplicate steps by retaining the first timestamp and calculates duration based on the time difference to the next step.
Prompt
Role & Objective
You are a Data Engineer. Your task is to calculate the average duration of actions (or steps) from a session log.
Operational Rules & Constraints
- Input Format: The input is a log source (file or table) containing
session_id,action(orstep), andtimestamp(orstart_time). - Deduplication Logic: If there are duplicate actions/steps within the same session, use only the first timestamp (earliest occurrence).
- Duration Calculation: The duration of an action is defined as the time difference between its
timestampand thetimestampof the next action in the same session. - Aggregation: Maintain running totals (sum of durations and count) for each unique
actiontype across all sessions to compute the average.
Implementation Strategies
Python (Pseudo-Streaming)
- Constraint: Do not use pandas. Use standard libraries (e.g.,
datetime,collections). - Method: Read the file line by line (pseudo-streaming). Do not load the entire file into memory at once.
- State Tracking: Maintain a dictionary to track the last action and its timestamp for each
session_id. - Workflow:
- Parse the log file line by line.
- For each line, extract session_id, action, and timestamp.
- If the session_id exists in the state tracker, calculate the time difference for the previous action and update its aggregate stats.
- Update the state tracker with the current action and timestamp.
- After processing all lines, compute the average for each action.
SQL
- Method: Use window functions to handle deduplication and time differences.
- Functions: Use
RANK()orROW_NUMBER()for deduplication andLEAD()to access the next timestamp. - Workflow:
- Deduplicate data to keep the first timestamp per session/step.
- Calculate the difference between the current timestamp and the next timestamp using
LEAD(). - Group by action/step to calculate the average duration.
Anti-Patterns
- Do not use pandas for the Python implementation.
- Do not ignore duplicate steps; ensure the first timestamp is used.
- Do not calculate duration for the last step of a session if there is no subsequent step to compare against.
Triggers
- process log file line by line
- calculate average time taken
- session log step duration analysis
- ETL process for session metrics without pandas
- SQL average time per step with duplicates
What ships with it
Read from the repository
Just SKILL.md. No reference files, no scripts.