Song to script
Universal AI Agent OS — audited skills, governance rules, replayable state. One contract, every host agent.
npx -y skills add event4u-app/agent-config --skill song-to-scriptAssembled from the repository path, not quoted from the project. Check it against their README if it does not work.
One thing to look at
- 7 stars7 stars. Stars are a popularity signal and not a quality one, but at this level it is likely that nobody has read this closely except its author, and you would be relying on your own review.
What its author says it does
Copied from the file, not written here
Turn an audio track into a timed `## Scene N` script: song sections → per-scene durations, auto mode adds mood + lip-sync lines. Triggers 'music video', 'from the song', 'cut to the beat'.
SKILL.md
14.2 KB, ~3.5k tokens by cl100k_base, as published. Nobody here has run it
song-to-script
Turn a song into
<project>/script.md— a sequence of## Scene Nblocks whoseduration:values sum to the track length and whose cut points land on real section boundaries. Consumed by/video:from-song, then handed toscene-expanderandvideo-director. Never invents timing — every boundary comes from the audio probe, and the probe'smethodtells this skill how musical (or not) those boundaries actually are.
When to use
- A music-video run needs scenes cut to the song (
/video:from-song). - An existing script must be re-timed to a track after the edit drifted from the beat (the named second consumer — re-time without a full re-author).
Do NOT use when:
- The operator already supplies a
## Scene Nscript withduration:values — feed it straight toscene-expander. - There is no audio — use the operator brief with
scene-expanderdirectly.
Inputs
- Audio analysis — adapter-first, probe as the floor:
- Analysis adapter (when an
audio-analysisprovider is configured — seeaudio-adapter-contract.md):{bpm, beats, downbeats, sections:[{start,end,label,energy?}]}— real musical structure. Beats/downbeats become the candidate cut grid; section labels are musical (verse,chorus, …). - Audio probe — JSON from
scripts/ai-video/lib/probe-audio.sh:{duration, method, warning?, sections:[{start,end,energy,label}]}.method: silence— boundaries are real quiet gaps; trust them as cuts.method: rms— boundaries are energy-delta inflections; usable but coarse.method: interval— the track is structurally flat (brick-walled / sustained); sections are fixed-interval, NOT musical. Whenmethodisinterval(orwarningis set), the emitted script header states that timing is interval-based and the operator should pass--scene-durationsfor musical sync. Never present interval cuts as beat-synced.
- Analysis adapter (when an
- Model capabilities — the chosen video model's renderable envelope
from the multiplexer manifest:
scripts/ai-video/adapters/<provider>.sh capability --model <id>→{min_duration, max_duration, audio_sync, aspect, verified}. Scene durations MUST land inside[min_duration, max_duration]; averified: falsemanifest entry is surfaced in the report, never trusted silently. - Mode —
brief(operator text is the creative source) orauto(infer mood + action from energy). - Brief (brief mode only) — free text: story, settings, look.
- Character lock (optional) —
<project>/character.jsonif a human subject was locked. Absent is normal — abstract / landscape / visualiser videos have no locked subject; see Step 2.
Procedure
Step 1: Map sections → scenes (capability-clamped, beat-aligned)
One ## Scene N per analysis section. duration: = end - start
(rounded to 0.5 s). Then clamp the plan to the chosen model's
renderable envelope — read min_duration / max_duration from the
model-capabilities manifest (<provider>.sh capability --model <id>),
falling back to the provider tuning for single-model adapters:
- Section shorter than
min_duration→ merge into its neighbour. With beat data, merge toward the neighbour that keeps the joined cut on a downbeat (else any beat); without beat data, merge into the shorter neighbour. - Section longer than
max_duration→ split into sub-scenes. With beat data, place every split point on the nearest downbeat (else beat) to the equal-division point — never mid-beat; without beat data, split equally. - No valid plan exists (e.g. the whole song is shorter than
min_duration, or a section cannot be split onto any beat inside the envelope) → halt and surface the conflict with the model id and the violated bound — an unbuildable plan never reaches the renderer.
Every emitted scene satisfies
min_duration ≤ duration ≤ max_duration. When the manifest entry is
verified: false, say so in the report — the envelope is
documented-best-effort, not a smoke-traced fact.
Step 2: Assign mood + action
First decide the subject mode:
- Character mode —
character.jsonexists: every scene'saction:names the locked subject, never a fresh description. - Style mode — no
character.json: scenes describe setting, palette, and motion continuity (the recurring look), not a person. This is the valid abstract / landscape / visualiser path — do not invent a human subject to fill the slot.
Then pick the prompt source per segment — the modality switch:
- Lyric segment (the vocal map places ≥1 transcribed line inside
it) → the scene prompt derives from the lyric line itself: its
imagery, subjects, and verbs seed
mood:+action:(in character mode, acted by the locked subject; in style mode, rendered as setting / weather / palette — never an invented human). The line lands indialogue:per Step 3. - Instrumental segment (no vocal-map line) → the scene prompt
derives from the audio features: section
label+energyvia the intent table below. Never recycle a lyric from another segment into an instrumental one.
Then assign per scene:
-
Brief mode — distribute the brief's beats across scenes in order; the modality switch still applies (lyric segments quote the brief's matching beat through the lyric's lens), and energy modulates pacing. Do not add story the brief did not state.
-
Auto mode — derive mood per section from
energyandlabel(probe labels and musical labels from the analysis adapter both map):label / energy default scene intent intro / low establishing wide, slow camera, calm subject/scene verse / mid narrative motion, medium framing, follow the subject build / rising approach, tightening framing chorus · drop / peak dynamic motion, weather/FX, fast push bridge · breakdown / dip close-up / detail, quiet, single light source outro / fade pull-back, resolve, hold
Energy → cut frequency + motion intensity. Section energy (0..1, relative to the track mean) drives both how often the edit cuts and how hard the camera moves — chorus = faster cuts / more motion:
| energy vs. track mean | cut length target | camera: motion intensity |
|---|---|---|
| ≥ mean + 0.10 (chorus / drop) | short — split the section toward min_duration, one scene per 1–2 downbeat bars | fast push / whip / handheld shake |
| within ±0.10 of mean (verse / build) | medium — one scene per section or per 4-bar phrase | steady dolly, slow tighten |
| ≤ mean − 0.10 (breakdown / outro) | long — merge toward max_duration, hold shots | locked-off or slow drift |
High-energy splitting and low-energy merging both stay inside the Step 1 capability envelope and land on downbeats — the energy table chooses where inside the envelope a scene length falls, never outside it.
Step 3: Vocal map — transcribe, never guess (vocal tracks)
LYRIC TIMING AND SINGER COME FROM THE TRANSCRIBED AUDIO, NEVER FROM A
BRIEF / STORY SKELETON OR A GUESSED STRETCH. NEVER PUT ONE SINGER'S
LINE ON ANOTHER SINGER'S SCENE.
When the track has vocals and the run intends lip-sync, build a
vocal map from the real audio before assigning any dialogue::
- Transcribe the audio to timestamped lines. Adapter-first: a
configured
lyricsprovider (e.g.audio-adapters/whisperx.sh) returns word-level timestamps plus per-line diarization labels (SPEAKER_00, …, or"?"when ambiguous):
No lyrics provider configured → OpenAIecho '{"audio_path":"<vocal-stem-or-song>"}' \ | scripts/ai-video/audio-adapters/whisperx.sh analyze/v1/audio/transcriptions(response_format=verbose_json→segments[].{start,end,text}) or local whisper as before (no speaker labels — every line starts as"?"). Either way the transcript is the only source of lyric timing. - Label the singer per line — map diarization labels (or unlabeled
lines) to cast names via the operator's who-sings reference (a
roster, a brief that names who sings which line, or a character
cast): one diarization label ↦ one cast name, consistently. If a
line's singer is genuinely ambiguous (label
"?", mixed-speaker line, or no roster match), keepsinger: "?"and surface it — never guess a singer to fill the slot. - Emit
<project>/vocal-map.json:[{start, end, text, singer}], timing verbatim from the transcript. - Validate — run the ground-truth enforcer before handing the map
to the sign-off gate:
It rejects re-timed lines, lyrics not in the transcript, and missing singers (exit 7, specific line named). A red validator is a halt — fix the map, never bypass.scripts/ai-video/lib/validate-vocal-map.sh <project>/vocal-map.json \ <project>/transcript.json --roster "<cast names>" - Place lines into the matching scene's
dialogue:block using the transcript timing, tagged with the singer (singer: "<line>"). A scene's lip-sync subject MUST be the line's labelled singer; a"?"line gets NO lip-sync scene until the operator resolves it.
No vocals / no transcript / no lip-sync intent → leave dialogue: empty;
the scene is performance / B-roll. Never fabricate lyrics, never
re-time a line off the brief, and in style mode dialogue: stays empty
(lip-sync needs a character subject). The /video:from-song sign-off
gate (its Step 6) shows this map for approval before any render.
Step 4: Emit + reconcile
Write <project>/script.md (and <project>/vocal-map.json when the
track has vocals). Report the delta, the section→scene map, the probe
method (so the operator sees whether cuts are silence-derived,
energy-derived, or interval-fallback), and whether lyric timing is
transcript-derived (it must be — never brief-derived). If the sum
cannot be reconciled (e.g. provider max-duration forces more time than
the song has), halt and surface the conflict — do not pad silently.
Step 5: Validate before handoff
Concrete checks (all must pass before the script is handed to
scene-expander):
- Assert
Σ(duration) == probe.durationwithin ±1.0 s; report the exact delta. A larger delta → halt, do not pad. - Verify every scene boundary equals a probe section boundary (or a
--scene-durationsvalue) — no invented cut points. - Confirm no scene
duration:exceeds the model'smax_durationor falls belowmin_duration(model-capabilities manifest, or the provider tuning for single-model adapters). - Verify every lyric-segment scene derives its prompt from its own vocal-map line and every instrumental scene from audio features — no cross-segment lyric recycling (modality switch).
- Ensure every
## Scene Ncarries all five keys (duration·mood·action·camera·dialogue), and thatdialogue:is empty in style mode.
Output format
script.mdopens with the derivation header —# <project> — derived from <song-file> (<mode> mode · cuts: <method>)— so the probemethodstays visible downstream.- One
## Scene Nblock per cut carrying exactly the keysduration·mood·action·camera·dialogue—scene-expanderconsumes this verbatim; keep the keys exact. dialogue:stays empty unless operator-supplied lyrics cover the section — detected vocal energy alone never fills it.
# <project> — derived from <song-file> (<mode> mode · cuts: <method>)
## Scene 1
duration: 6.0
mood: establishing, cold, pre-storm
action: <subject from character.json, OR style description in style mode>
camera: slow push-in
dialogue:
## Scene 2
duration: 4.5
mood: build, rising tension
action: close on <subject / detail>, wind picking up
camera: handheld tighten
dialogue:
- "<subject>: \"<lyric line for this section, if any>\""
Gotcha
method: intervalis the brick-walled-master signal, not a bug. A compressed modern master has near-constant RMS and no silence, so the probe degrades to fixed intervals. That is the honest floor — surface it and point the operator at--scene-durations; never dress interval cuts up as beat-synced.- A vocal section without supplied lyrics is B-roll, not lip-sync.
Detected vocal energy alone does not authorise
dialogue:— only operator-supplied lyrics do. - Style mode is the default for a no-character run, not an error path. Landscape / abstract / visualiser videos never get a fabricated human subject.
Do NOT
- Do NOT invent timing. Every cut maps to a probe boundary or a
--scene-durationsvalue — never to taste. - Do NOT present
interval-fallback cuts as beat-synced. Always surface the probemethod. - Do NOT emit a clip outside the provider's min/max duration — split/merge in Step 1 instead.
- Do NOT fabricate lyrics or story beyond the brief / detected vocals.
- Do NOT invent a human subject in style mode; defer identity to
character.jsononly when a lock exists. - Do NOT pad a unreconcilable timing sum — halt and surface it.
See also
/video:from-song— the command that drives this skillscene-expander— consumes the emitted scriptcharacter-consistency— supplies the locked subject referenced inaction:(character mode only)