Video understanding
Skill worldwonderer/video-recap-skills/skills/video-understanding
把视频分析为结构化理解索引:场景检测、ASR 转写、逐场景 VLM 观察、静音窗口、融合时间线和写作 brief。 用于理解、索引或总结视频,也作为后续创作前的分析阶段。输入视频文件;输出 scenes.json、 asr_result.json、vlm_analysis.json、silence_periods.json、timeline_fusion.json、agent_narration_brief.md。 触发词:视频理解、视频分析、视频索引、video understanding、analyze video、看懂视频。From its SKILL.md
npx -y skills add worldwonderer/video-recap-skills --skill video-understandingAssembled from the repository path, not quoted from the project. Check it against their README if it does not work.
2 things to look at
- reads credentialsReads from 1 credential source: `MIMO_API_KEY`.
- runs commandsInstructs the agent to run 2 commands, including `python3 scripts/understand.py <video> --work-dir <work_dir> [--context "节目名/角色名"] [--scene-threshold 0.1] [--skip-asr] [--mimo-video-overview] [--force]` and 1 more.
SKILL.md
3.8 KB, ~1.2k tokens by cl100k_base, as published. Nobody here has run it
1. 定位
本技能把源视频转成 Agent 与下游阶段可读取的理解索引。它的创作角色是素材观察员 / 场记,不是导演:
- 先观察,再解释;事实与推断分开。
- 除了“发生了什么”,还要让下游看见知识、权力、目标、关系或情绪在哪一刻变化。
- 标出由谁的 POV 承载变化、哪个反应或表演不可替代,以及哪里存在完整台词/动作的自然剪辑边界。
- 证据不足时保留不确定性,不制造戏剧结论。
2. 处理阶段
- 场景检测:写
scenes.json,包含切点、时长和废片段过滤结果。 - 抽帧:为视觉分析提取代表帧。
- ASR:通过
mimo-v2.5-asr写时间戳对白asr_result.json。 - 静音检测:写
silence_periods.json,标注安静窗口与has_speech。 - VLM 观察:写
vlm_analysis.json,包含场景描述、深层分析和frame_facts。 - 时间线融合与创作 brief:写
timeline_fusion.json、asr_writing_chunks.json和agent_narration_brief.md。
各阶段只有在输出产物与 provenance sidecar 同时匹配当前视频及影响结果的设置时才会复用;--force 强制重算。
3. 环境要求
# ffmpeg: brew install ffmpeg | apt install ffmpeg | choco install ffmpeg
export MIMO_API_KEY=***
ASR 使用 mimo-v2.5-asr;VLM 使用 mimo-v2.5。--skip-asr 可跳过对白转写,但完整理解仍需要 MIMO_API_KEY 运行 VLM。--mimo-video-overview 可开启按场景块的视频概览。
若 work_dir/background_research.json 存在,本技能会把剧情梗概和角色名折入 VLM 上下文;--context 可补充一条简短提示。
下面的 scripts/... 均相对于本技能目录。若执行器从仓库根目录启动,请给脚本路径加上本技能的绝对目录。脚本不从其他技能目录读取文件;外部输入仅限命令显式传入的视频、参数与 work_dir 产物。
4. 运行命令
python3 scripts/understand.py <video> --work-dir <work_dir> \
[--context "节目名/角色名"] [--scene-threshold 0.1] [--skip-asr] [--mimo-video-overview] [--force]
5. 输出契约
| 文件 | 内容 |
|---|---|
scenes.json | 场景切点、起止时间与时长 |
asr_result.json | [{start, end, text}] 时间戳对白 |
vlm_analysis.json | 逐场景描述、深层分析与 frame_facts |
silence_periods.json | [{start, end, duration, has_speech}] 安静窗口 |
timeline_fusion.json | VLM、ASR 与静音信息的统一时间线 |
asr_writing_chunks.json | 按句界和场景切分的 ASR 写作块 |
agent_narration_brief.md | Agent 首先阅读的创作简报 |
后续写作阶段根据创作简报与索引制定方案并写 narration.json。
6. 参考资料
- 背景调研:
references/research-guide.md,产出background_research.json。 - JSON 结构:
references/data-schema.md。
7. 能力边界
- 不写解说词,也不做解说评分;只负责生成理解索引与创作简报。
- 不剪辑、不配音、不合成视频。
- 不编造信号无法支持的剧情;当 ASR / VLM 过薄时输出素材警告。
- 不发布、不调度,只向
work_dir写产物并停止。
What ships with it: 25 files
323.4 KB alongside SKILL.md, 22 of them executable
references/
- data-schema.md11.1 KB
- prompt-templates.md1.7 KB
- research-guide.md2.8 KB
scripts/
- agent_brief.pyruns24.0 KB
- agent_text.pyruns14.7 KB
- asr.pyruns10.2 KB
- brief_context.pyruns16.0 KB
- brief_inputs.pyruns10.6 KB
- brief.pyruns370 B
- brief_timeline.pyruns17.4 KB
- consolidate.pyruns24.4 KB
- deslop_qc.pyruns10.8 KB
- detect.pyruns21.5 KB
- extract.pyruns2.5 KB
- lib.pyruns23.9 KB
- narration_lint.pyruns27.1 KB
- speech_ownership.pyruns6.9 KB
- storyboard.pyruns20.6 KB
- timeline_fusion.pyruns7.2 KB
- understanding_brief.pyruns8.1 KB
- understanding_cache.pyruns11.1 KB
- understanding_runner.pyruns14.1 KB
- understanding_storyboard.pyruns6.8 KB
- understand.pyruns192 B
- vlm.pyruns29.5 KB
Gives 1 of the 12 instructions most research analysis skills give in ~1.2k tokens
Counted across 1,213 of the 2,113 authors here whose files we hold, read 2026-09-06
- Cite sources for every important claimin 47 of 1213, across 38 files
- Separate facts from inferences and recommendationshere, and in 21 of 1213, across 12 files
- Write findings to a markdown filein 19 of 1213
- Label every insight with a confidence levelin 18 of 1213, across 8 files
- Read product marketing context before asking questionsin 18 of 1213, across 8 files
- Rank themes by frequency and intensityin 16 of 1213, across 6 files
- Establish research mode before proceedingin 16 of 1213, across 6 files
- Segment survey responses by customer tier or tenurein 16 of 1213, across 6 files
- Categorize support tickets before analyzingin 16 of 1213, across 6 files
- Weight research sources from the last twelve monthsin 16 of 1213, across 6 files
- Use at least five data points per segmentin 15 of 1213, across 5 files
- Extract verbatim quotes for all research findingsin 15 of 1213, across 5 files
Said here and by no other author read
- identify POV changes and key reactions
- mark natural editing boundaries
- preserve uncertainty when evidence is insufficient
- use background research for VLM context
Grouped from the skills themselves: near-identical wordings counted once, and counted by distinct author, so one author publishing three of these counts once. Length counted with cl100k_base; the agent that loads this file may tokenize it differently.