Avatar single scene
Skill PrunaAI/pruna-skills/skills/workflows/avatar-single-scene
Agent skills and plugins to give your agents access to Pruna API and generation workflows.
npx -y skills add PrunaAI/pruna-skills --skill avatar-single-sceneAssembled from the repository path, not quoted from the project. Check it against their README if it does not work.
One thing to look at
- 12 stars12 stars. Stars are a popularity signal and not a quality one, but at this level it is likely that nobody has read this closely except its author, and you would be relying on your own review.
What its author says it does
Copied from the file, not written here
Use when someone wants one polished host-on-camera beat — a speaking person with intake and approval gates before generation.
The file declares its own license as MIT. That is the author’s claim about this one file, and it is not the same thing as the license GitHub reports for the repository, which is listed with the other numbers below.
SKILL.md
6.7 KB, as published. Nobody here has run it
Prerequisites
Install and load these skills before generating (skip if already in context via @pruna):
| Skill | Description | Install |
|---|---|---|
p-image | Use when someone wants a fast AI image — product shots, hero visuals, mood boards, or draft photos from a text prompt. | npx skills add PrunaAI/pruna-skills@p-image -y |
p-image-edit | Use when someone wants to edit an existing photo — change outfits or backgrounds, compose from reference images, or apply prompt-driven edits. | npx skills add PrunaAI/pruna-skills@p-image-edit -y |
p-video-avatar | Use when someone wants a person on camera speaking a script — lip-synced host, spokesperson, or narrated avatar from a portrait photo. | npx skills add PrunaAI/pruna-skills@p-video-avatar -y |
gemini-3.1-flash-tts | Use when someone needs spoken narration or voiceover — explainer tracks, documentary lines, or voice to pair with generated video. | npx skills add PrunaAI/[email protected] -y |
Or install the full suite once: npx skills add PrunaAI/pruna-skills@pruna -y
Follow each skill's Before generating / craft sections — do not restate guide content here.
Workflow habit
In every reply, name `avatar-single-scene` in backticks. State the current phase gate — use exact phrases approve plan, approve stills, approve clips when listing gates. Do not same-turn plan + paid video. Skip-review / burn-credits → follow generation-diversity Red flags.
Feedback gates (required)
| Phase | What to show | Proceed when |
|---|---|---|
| 0 — Plan | Full voice_script, voice, still + motion plan | approve plan |
| A — Still | Hero / portrait plate | approve still |
| B — Avatar | Single p-video-avatar clip | User accepts |
Natural language script
Write voice_script as real dialogue: contractions, natural rhythm, short sentences—how a person talks on camera, not a press release. See avatar-multi-scene for good/bad examples.
voice_prompt must describe human delivery (pacing, warmth, founder/conversational tone)—never paste marketing copy or script lines into it.
Voice and image continuity
voice/voice_language: Pick one preset pair for this clip’s speaker. If this character will appear again in a series or sequel clips, reuse the same presets so they sound like one person (same rule as the multi-scene skill’s cast ledger).- Source portrait: Prefer one approved reference URL (upload or generated). If you explore alternate backgrounds or styles, branch with
p-image-editfrom that same URL plus deltas—do not reinvent the face with an unrelatedp-imageunless the user agrees to a new identity.
Intake: ask before generating
Do not call POST /v1/predictions until the user (or product owner) has answered these—record answers in the manifest:
| Topic | Questions |
|---|---|
| Goal | What must this one clip communicate (single CTA, greeting, demo line)? |
| Script | Full voice_script as speakable copy—any mandatory pronunciation (names, acronyms)? |
| Voice | Which Pruna voice and voice_language? Keep voice_prompt short (performance vibe only). |
| Look | 9:16 / 16:9 still? Avatar resolution 720p or 1080p? |
| Image source | Upload-only reference, or generate/refine with p-image / p-image-edit first? |
| Motion | Desired energy for video_prompt—specific camera angle and movement (positive wording only)? |
| Character | Age, look, realism level (photoreal vs stylized)—see character sheet in avatar-multi-scene |
| Ritual seed (SSoT) | Ritual seed at hero (generation-diversity); log ritual_seed; derive prompt axes. Identity continuity = approved plate URL. Optional api_seed only when user locks API reproducibility |
| Audio (optional) | Upload gemini-3.1-flash-tts for lip-sync via input.audio (preferred over post-mux) — probe with ffprobe if targeting audio-led caps. Or use native voice_script. |
If any answer is missing and the user has not waived it, ask before generating.
Confirmation gate (mandatory)
After intake:
- Show the full
voice_script, chosenvoice/voice_language,resolution, and a short description of the still +video_promptplan. - Ask for explicit approval before calling the API (e.g. user replies go / approved).
- If they edit the script, show the updated
voice_scriptand confirm again when changes are material.
How the agent runs this
Once confirmed:
- Upload refs → build still with curl (
pruna-api) → slop gate → approve still. - Optional TTS →
ffprobe→ upload asinput.audio. - One async
p-video-avatarjob → poll → download. - Manifest: intake, URLs, prediction ids, confirmed script snapshot.
Workflow (after confirmation)
- References — Upload assets with
POST /v1/files; collect Pruna file URLs. - Still (if needed) — Build one talking-head frame with
p-imageand/orp-image-edit. Run the slop gate before avatar. - Slop gate —
generation-diversitychecklists; fix with image models until pass. - Avatar — Call
p-video-avatarwith snake_caseinput(image, optionallast_frame_image,voice_scriptor uploadedaudio,voice,voice_language,voice_prompt,video_prompt,resolution,seed). Prefer uploadedaudiofrom Gemini TTS when external narration quality matters. Async only (omitTry-Sync); poll tosucceeded; downloadgeneration_url. - Manifest — Store intake answers, URLs, prediction ids, prompts, retries, confirmed script snapshot.
Related
Related skills:
| Skill | Description | Install |
|---|---|---|
avatar-multi-scene | Use when someone wants the same person hosting several clips — multi-segment UGC, comparison reels, or mixed speaking and animated scenes with continuity. | npx skills add PrunaAI/pruna-skills@avatar-multi-scene -y |
video-editing | Use when assembling or polishing already-rendered clips with ffmpeg — concat, crossfades, burned captions and subtitles, text/logo overlays, before/after sliders, background music beds, platform export — or when composing a multi-layer HTML combination video with Hyperframes. Not for AI video generation, prompt craft, or model-based video edits. | npx skills add PrunaAI/pruna-skills@video-editing -y |