Video understand
Open-source agentic video editing skills — an AI-powered alternative to Opus Clip and CapCut
npx -y skills add WhiteTowerAI/cut-as-code --skill video-understandAssembled from the repository path, not quoted from the project. Check it against their README if it does not work.
One thing to look at
- 8 stars8 stars. Stars are a popularity signal and not a quality one, but at this level it is likely that nobody has read this closely except its author, and you would be relying on your own review.
What its author says it does
Copied from the file, not written here
Use when a video project needs reusable media metadata, word-level transcription, objective speech analysis, or evidence-backed semantic understanding before optional editing skills run.
SKILL.md
3.6 KB, as published. Nobody here has run it
Video Understand
Build the shared evidence layer once. Keep observations in source time and leave editorial decisions to downstream skills.
This skill is a prerequisite for /video-cut, /video-to-shorts,
/video-add-captions, and /video-add-content-cards. Run it first so those skills
consume the same validated evidence and timeline.
Dependencies
Require ffmpeg/ffprobe, Python, and faster-whisper for transcription. Check them before processing media.
Workflow
-
Initialize a project from the original source:
python scripts/init_project.py path/to/source.mp4 path/to/my-video-projectThis creates
input/,review/00-video-understanding/,final/, the minimal machine-facingwork/tree, an identity timeline,project.json, media facts, andSTART-HERE.md. It does not create folders for unselected edit operations. -
Probe again only when the source needs an explicit metadata refresh:
python scripts/probe.py input/original-video.mp4 work/understand/media.json -
Extract 16 kHz mono audio and transcribe it:
ffmpeg -y -i input/original-video.mp4 -ac 1 -ar 16000 work/cache/audio16k.wav python scripts/transcribe.py work/cache/audio16k.wav work/understand/transcript medium ` --lang auto --cache-dir work/cache/faster-whisperUse
--lang autofor unknown or mixed-language speech. Never infer the spoken language from the language of the user's prompt. Pass a fixed language such as--lang zhonly when the audio itself or explicit user metadata establishes it. Keep model downloads in the project-localwork/cache/faster-whisper/cache. Faster-whisper may emit an occasional point-timed word with equal start/end values; the shared timeline mapper preserves it as a 1 ms interval so captions and derivatives do not silently lose text. -
Generate objective metrics and semantic candidates:
python scripts/analyze.py work/understand/transcript.json work/understand/analysis.json -
Read the source, transcript, and analysis. Author
work/understand/understanding.jsonwith factual summaries, source-time ranges, confidence, and transcript evidence. Do not prescribe cuts, cards, or looks. -
Validate before downstream use:
python scripts/validate.py understanding work/understand/understanding.json work/understand/transcript.json -
Create only these useful review artifacts under
review/00-video-understanding/:video-summary.md,transcript.srt, andcontact-sheet.jpg. Verify metadata, timestamps, evidence references, and visible frames. Do not substitute PNG or ad hoc filenames for the protocol names. -
Mark the understanding operation
check.statusaspassonly after the review artifacts and semantic evidence validate. Operation lifecyclestatusand check result are separate; never writeverifiedintocheck.status.
Contracts
- Read project-schema.md when creating or validating
project.json. - Read timeline-schema.md when mapping source and program time.
- Read understanding-schema.md before authoring semantic understanding.
- Use understanding.example.json as a compact example.
All durable machine files live in work/understand/. Treat work/cache/ as disposable.