agentsclimarketplace

Talking head and piece to camera

Skill social-media-skills/skills/skills/talking-head-and-piece-to-camera

106 social media skills for AI agents - strategy, writing, video, design, platform growth, publishing, and analytics. Works with Claude, Cursor, OpenClaw, Hermes & 40+ agents.

Install
npx -y skills add social-media-skills/skills --skill talking-head-and-piece-to-camera

Assembled from the repository path, not quoted from the project. Check it against their README if it does not work.

2 things to look at

  • 21 days oldThe repository was created 21 days ago. New is not bad, but a brand new repository carrying a familiar-sounding name is the shape a typosquat arrives in, and there has been no time for anyone else to find a problem with it.
  • 4 stars4 stars. Stars are a popularity signal and not a quality one, but at this level it is likely that nobody has read this closely except its author, and you would be relying on your own review.

What its author says it does

Copied from the file, not written here

The on-camera delivery craft — helping a real human film themselves talking to a lens and look like themselves doing it. Use when someone wants a "talking head video" or "piece to camera," says "film myself" or "I look stiff on camera," asks about a teleprompter, framing, lighting, audio, or retakes, or wants to batch-film videos. Uses the TAKES framework. Phone-first: gear is almost never the bottleneck. Reads brand-profile + voice-builder first; takes its script from short-form-video-script (that writes it, this delivers it). The agent coaches setup + delivery, formats prompter/beat-map scripts, and plans batch days; the HUMAN films and picks the take (the agent cannot see footage); WoopSocial publishes the finished file. Camera-shy? Route honestly to heygen/synthesia or faceless formats. Never fabricates "that take looks great." Distinct from scripting-and-storyboarding (the shoot plan), heygen/synthesia (avatars), and captions-and-clipping/capcut/descript (the edit).

SKILL.md

8.6 KB, ~1.9k tokens by cl100k_base, as published. Nobody here has run it

talking-head-and-piece-to-camera

The on-camera delivery craft — tape the setup, anchor the map (not the lines), kick the first 3 seconds, embrace the retake rules, stack the batch. The script comes from short-form-video-script; the human films and picks the take; WoopSocial publishes the finished file.

The POV: presence beats polish, and the phone in your pocket is enough

A talking head works because a real face builds parasocial trust an avatar can't (that's exactly why synthesia routes trust-led founder content here). Three truths most first-timers get backwards. First, gear is not the bottleneck — a phone at eye level, facing a window, with a cheap lav mic outperforms an expensive camera set up wrong; viewers forgive soft video and never forgive bad audio. Second, reading kills it — memorize the map (the beats), not the lines; a word-for-word read shows in the eyes, and a slightly imperfect riff reads as human. Third, the good-enough take ships — take 4 is usually worse than take 2 because energy decays faster than delivery improves; perfectionism is a retention strategy for exactly nobody. Deliver 20% more energy than feels natural, talk to one person, and publish the take where you sound like yourself.

Read these first

  1. brand-profile + voice-builder — who's talking and how they sound off-camera (the on-camera target).
  2. short-form-video-script (or youtube-long-form for long pieces) — the script/beats being delivered; scripting-and-storyboarding if the shoot has multiple scenes.

The framework: TAKES

(Depth: references/the-takes-framework.md.)

  • T — Tape the setup: phone at eye level, arm's-length-plus, lens at the top; face the biggest window (never behind you); mic close (wired lav or phone ≤60cm); quiet room > any mic; clean-but-real background with depth; vertical 9:16, eyes in the top third, caption-safe zones clear.
  • A — Anchor the map, not the lines: memorize 3–5 beats + the first line + the last line verbatim; riff the middle. Teleprompter only if unavoidable — text beside the lens, narrow column, slow scroll, rehearse twice, or the line-at-a-time method. Reading eyes are visible; descript Eye Contact patches a read, not a performance.
  • K — Kick the first 3 seconds: start mid-energy, already talking — no breath, no settle, no "hey guys." Say the hook fresh, first, every session. Smile-then-speak; hands visible; deliver to ONE person behind the lens.
  • E — Embrace the retake rules: retake per beat, not per video; keep rolling and just say the line again (clap between takes to mark them); the three-strike rule — a line that fails 3× is a writing problem, send it back to short-form-video-script; ship the good-enough take.
  • S — Stack the batch: one setup, 4–8 scripts per session, hardest script first, swap tops between scripts so posts don't look same-day; stop at ~60–90 min when energy dies. Plan with batch-content-plan / content-calendar.

The reality (verify-quarterly)

Any recent phone shoots 4K that out-resolves every social feed; audio drives perceived quality more than image (creator consensus — attribute); a below-eye lens reads as looming, backlit windows silhouette you; on-camera energy reads ~20% flatter than it feels (broadcast coaching convention); take quality typically peaks by take 2–3 then decays with energy; batch sessions fade after ~60–90 minutes — directional, attribute, verify-quarterly. Full figures + phone-first setup specifics: references/talking-head-2026-reality.md. Batch-day recipe, setup recipes (desk / walking / car), and camera-shy on-ramps: references/batch-filming-and-recipes.md.

Honest scope (never violate)

  • The agent coaches setup and delivery, formats the script as a beat map or prompter text, writes shot lists and batch plans, and gives a self-review checklist. The human films, performs, and picks the take. The agent cannot see the footage — it never judges a take, never fabricates "that looked natural," and never claims a result it can't observe. WoopSocial publishes the finished file only — it does not film, edit, or analyze footage.
  • Never prescribe buying gear as the fix (phone-first; upgrade only when a named limit is hit), shame a camera-shy human onto camera (route to avatars/faceless honestly), or skip consent for anyone else who appears on camera. AI enhancement of a real human (eye-contact fix, retouch) stays within platform disclosure rules. (Full scope: references/scope-and-connections.md.)

Edge cases (handle honestly)

  • Camera-shy / won't film: legitimate. Route to heygen (creator/social lane) or synthesia (enterprise/ L&D lane) for a disclosed avatar, or to faceless formats (screen-record / B-roll + ai-voiceover). Offer the gentle on-ramp — voice-only first, then hands/desk shots, then face — but never pressure.
  • Perfectionist / 30 takes deep: invoke the good-enough doctrine — cap takes per beat at 3, ship the take where they sound like themselves, and remind them the audience rewards presence, not polish.
  • "Watch my take and tell me it's good": can't — no eyes on footage. Hand over the self-review checklist (hook lands on mute? energy? eyes on lens? audio clean?) and let the human verdict stand.

Distinct from its siblings (route correctly)

talking-head-and-piece-to-camera (this) = the human filming/delivery craft · short-form-video-script = the script this delivers (pair) · scripting-and-storyboarding = the multi-scene shoot plan (this is the shoot-day performance) · heygen / synthesia = synthetic presenters when the human can't/won't film · captions-and-clipping / capcut / descript = the edit after the shoot (descript's Eye Contact patches a read; it doesn't replace delivery) · livestream-and-realtime = live to-camera (no retakes) · ai-voiceover = voice without a face.

Where this connects

Reads first: brand-profile + voice-builder. Takes the script from: short-form-video-script (or youtube-long-form), the plan from scripting-and-storyboarding, batch slots from batch-content-plan + content-calendar. Feeds: captions-and-clipping / capcut / descript (the edit), opus-clip (clipping long pieces), cross-platform-repurposing. Routes away: avatars → heygen / synthesia. Publishes via: edited file → scheduling-and-queue → WoopSocial. Measure with: native + analytics-and-reporting on 3s hold / AVD / completion — never fabricated.

Definition of done

A filmed piece to camera delivered from a beat map (first + last lines verbatim, middle riffed), shot phone-first at eye level facing the light with clean close audio and a caption-safe 9:16 frame, opening mid-energy on the hook with no wind-up, retaken per beat under the three-strike rule and shipped at good-enough rather than sanded lifeless, batched (4–8 scripts, top swaps, ≤90 min) when volume is the goal; camera-shy humans routed honestly to heygen/synthesia or faceless formats; the human filmed and picked the take (the agent never judged footage it can't see, never fabricated praise, never prescribed gear as the fix); consent handled for anyone else in frame; the file edited via captions-and-clipping/capcut/descript and published via scheduling-and-queue → WoopSocial; measured on 3s hold / AVD / completion; and correctly distinguished from short-form-video-script, scripting-and-storyboarding, heygen/synthesia, and the editing skills.

Keep looking

Skills are one crate of 328,083. Ordering is by how many stacks a row turns up in, so the top of any crate is what has actually been picked rather than what has the most stars.