Communitygithub.com

talking-head-and-piece-to-camera

>-

talking-head-and-piece-to-camera 是什么?

talking-head-and-piece-to-camera is a Claude Code agent skill that >-.

兼容平台~Claude Code~Codex CLI~Cursor
npx skills add https://github.com/social-media-skills/skills/tree/main/skills/talking-head-and-piece-to-camera

在你喜欢的 AI 中提问

打开一个已预加载此 Agent Skill 的新对话。

文档

talking-head-and-piece-to-camera 是做什么的?

The on-camera delivery craft — tape the setup, anchor the map (not the lines), kick the first 3 seconds, embrace the retake rules, stack the batch. The script comes from short-form-video-script; the human films and picks the take; WoopSocial publishes the finished file.

The POV: presence beats polish, and the phone in your pocket is enough

A talking head works because a real face builds parasocial trust an avatar can't (that's exactly why synthesia routes trust-led founder content here). Three truths most first-timers get backwards. First, gear is not the bottleneck — a phone at eye level, facing a window, with a cheap lav mic outperforms an expensive camera set up wrong; viewers forgive soft video and never forgive bad audio. Second, reading kills it — memorize the map (the beats), not the lines; a word-for-word read shows in the eyes, and a slightly imperfect riff reads as human. Third, the good-enough take ships — take 4 is usually worse than take 2 because energy decays faster than delivery improves; perfectionism is a retention strategy for exactly nobody. Deliver 20% more energy than feels natural, talk to one person, and publish the take where you sound like yourself.

Read these first

  1. brand-profile + voice-builder — who's talking and how they sound off-camera (the on-camera target).
  2. short-form-video-script (or youtube-long-form for long pieces) — the script/beats being delivered; scripting-and-storyboarding if the shoot has multiple scenes.

The framework: TAKES

(Depth: references/the-takes-framework.md.)

  • T — Tape the setup: phone at eye level, arm's-length-plus, lens at the top; face the biggest window (never behind you); mic close (wired lav or phone ≤60cm); quiet room > any mic; clean-but-real background with depth; vertical 9:16, eyes in the top third, caption-safe zones clear.
  • A — Anchor the map, not the lines: memorize 3–5 beats + the first line + the last line verbatim; riff the middle. Teleprompter only if unavoidable — text beside the lens, narrow column, slow scroll, rehearse twice, or the line-at-a-time method. Reading eyes are visible; descript Eye Contact patches a read, not a performance.
  • K — Kick the first 3 seconds: start mid-energy, already talking — no breath, no settle, no "hey guys." Say the hook fresh, first, every session. Smile-then-speak; hands visible; deliver to ONE person behind the lens.
  • E — Embrace the retake rules: retake per beat, not per video; keep rolling and just say the line again (clap between takes to mark them); the three-strike rule — a line that fails 3× is a writing problem, send it back to short-form-video-script; ship the good-enough take.
  • S — Stack the batch: one setup, 4–8 scripts per session, hardest script first, swap tops between scripts so posts don't look same-day; stop at ~60–90 min when energy dies. Plan with batch-content-plan / content-calendar.

The reality (verify-quarterly)

Any recent phone shoots 4K that out-resolves every social feed; audio drives perceived quality more than image (creator consensus — attribute); a below-eye lens reads as looming, backlit windows silhouette you; on-camera energy reads ~20% flatter than it feels (broadcast coaching convention); take quality typically peaks by take 2–3 then decays with energy; batch sessions fade after ~60–90 minutes — directional, attribute, verify-quarterly. Full figures + phone-first setup specifics: references/talking-head-2026-reality.md. Batch-day recipe, setup recipes (desk / walking / car), and camera-shy on-ramps: references/batch-filming-and-recipes.md.

Honest scope (never violate)

  • The agent coaches setup and delivery, formats the script as a beat map or prompter text, writes shot lists and batch plans, and gives a self-review checklist. The human films, performs, and picks the take. The agent cannot see the footage — it never judges a take, never fabricates "that looked natural," and never claims a result it can't observe. WoopSocial publishes the finished file only — it does not film, edit, or analyze footage.
  • Never prescribe buying gear as the fix (phone-first; upgrade only when a named limit is hit), shame a camera-shy human onto camera (route to avatars/faceless honestly), or skip consent for anyone else who appears on camera. AI enhancement of a real human (eye-contact fix, retouch) stays within platform disclosure rules. (Full scope: references/scope-and-connections.md.)

Edge cases (handle honestly)

  • Camera-shy / won't film: legitimate. Route to heygen (creator/social lane) or synthesia (enterprise/ L&D lane) for a disclosed avatar, or to faceless formats (screen-record / B-roll + ai-voiceover). Offer the gentle on-ramp — voice-only first, then hands/desk shots, then face — but never pressure.
  • Perfectionist / 30 takes deep: invoke the good-enough doctrine — cap takes per beat at 3, ship the take where they sound like themselves, and remind them the audience rewards presence, not polish.
  • "Watch my take and tell me it's good": can't — no eyes on footage. Hand over the self-review checklist (hook lands on mute? energy? eyes on lens? audio clean?) and let the human verdict stand.

Distinct from its siblings (route correctly)

talking-head-and-piece-to-camera (this) = the human filming/delivery craft · short-form-video-script = the script this delivers (pair) · scripting-and-storyboarding = the multi-scene shoot plan (this is the shoot-day performance) · heygen / synthesia = synthetic presenters when the human can't/won't film · captions-and-clipping / capcut / descript = the edit after the shoot (descript's Eye Contact patches a read; it doesn't replace delivery) · livestream-and-realtime = live to-camera (no retakes) · ai-voiceover = voice without a face.

Where this connects

Reads first: brand-profile + voice-builder. Takes the script from: short-form-video-script (or youtube-long-form), the plan from scripting-and-storyboarding, batch slots from batch-content-plan + content-calendar. Feeds: captions-and-clipping / capcut / descript (the edit), opus-clip (clipping long pieces), cross-platform-repurposing. Routes away: avatars → heygen / synthesia. Publishes via: edited file → scheduling-and-queue → WoopSocial. Measure with: native + analytics-and-reporting on 3s hold / AVD / completion — never fabricated.

Definition of done

A filmed piece to camera delivered from a beat map (first + last lines verbatim, middle riffed), shot phone-first at eye level facing the light with clean close audio and a caption-safe 9:16 frame, opening mid-energy on the hook with no wind-up, retaken per beat under the three-strike rule and shipped at good-enough rather than sanded lifeless, batched (4–8 scripts, top swaps, ≤90 min) when volume is the goal; camera-shy humans routed honestly to heygen/synthesia or faceless formats; the human filmed and picked the take (the agent never judged footage it can't see, never fabricated praise, never prescribed gear as the fix); consent handled for anyone else in frame; the file edited via captions-and-clipping/capcut/descript and published via scheduling-and-queue → WoopSocial; measured on 3s hold / AVD / completion; and correctly distinguished from short-form-video-script, scripting-and-storyboarding, heygen/synthesia, and the editing skills.

Individual skills in this repo

This repo contains 20 individual skills — each has its own dedicated page.

相关技能