Communitygithub.com

beachyphotoandfilm-arch/video-skills

Cut vertical Instagram reels out of a long-form talking-head video, hook first. Picks the strongest standalone moments from the transcript, opens each reel on its punchiest line, builds vertical timelines in DaVinci Resolve framed on the speaker's face with the LUT applied, and finishes them: a hook (a held text card, or a word-by-word "build" hook with punch-in) and burned-look captions on V3; this-or-that reels get product graphics and a comment-keyword CTA card. You trim and render, nothing to import. Trigger on "make reels", "clip this for Instagram", "cut some verticals", or after a YouTube video is cut.

Was ist video-skills?

video-skills is a Claude Code agent skill that cut vertical Instagram reels out of a long-form talking-head video, hook first. Picks the strongest standalone moments from the transcript, opens each reel on its punchiest line, builds vertical timelines in DaVinci Resolve framed on the speaker's face with the LUT applied, and finishes them: a hook (a held text card, or a word-by-word "build" hook with punch-in) and burned-look captions on V3; this-or-that reels get product graphics and a comment-keyword CTA card. You trim and render, nothing to import. Trigger on "make reels", "clip this for Instagram", "cut some verticals", or after a YouTube video is cut.

Funktioniert mit~Claude Code~Codex CLI~Cursor
npx skills add https://github.com/beachyphotoandfilm-arch/video-skills/tree/HEAD/skills/video-reels

In Ihrer bevorzugten KI fragen

Öffnet einen neuen Chat, in dem dieser Agent-Skill bereits geladen ist.

Dokumentation


name: video-reels description: Cut vertical Instagram reels out of a long-form talking-head video, hook first. Picks the strongest standalone moments from the transcript, opens each reel on its punchiest line, builds vertical timelines in DaVinci Resolve framed on the speaker's face with the LUT applied, and finishes them: a hook (a held text card, or a word-by-word "build" hook with punch-in) and burned-look captions on V3; this-or-that reels get product graphics and a comment-keyword CTA card. You trim and render, nothing to import. Trigger on "make reels", "clip this for Instagram", "cut some verticals", or after a YouTube video is cut.

Reels

Vertical clips from a long-form video (or standalone reels filmed as reels).

Config: set your brand here

Edit these once before the first run. Everything below refers to them by name.

settingwheredefault
accent color--brand at the top of scripts/motion.html, scripts/thisorthat.html, scripts/hook.html#3b82f6
display font (hook cards, titles, keywords)--display-font in the same filesAnton (free, Google Fonts)
caption fontscripts/caption.htmlMontserrat ExtraBold 800 (free)
camera LUT"lut" in reel-params.json, and LUT env var for preview_reel.py<PATH_TO_YOUR_LUT>.cube
new-footage drop folderstep "Where new footage lands"<YOUR_FOOTAGE_DRIVE>/<INBOX_FOLDER>/
b-roll libraryanimated style, step 6<YOUR_FOOTAGE_DRIVE>/<shoot>/<cam>/full/
reel trackerStep 0drafts/social/reel-tracker.md in your project
whisper modelWHISPER_MODEL env var~/.cache/whisper-models/ggml-small.en.bin
comment-keyword CTActa plan itemsyour own keyword (e.g. GUIDE)
background music referencemusic stepa reference track you like (e.g. a Spotify track ID for SearchRecordings)

Fonts: install Montserrat and your display font into ~/Library/Fonts (or keep the Google Fonts @import in each HTML file and render online). If your display font is licensed, keep it out of the repo.

Tools: whisper-cli (whisper.cpp), ffmpeg, Node + puppeteer-core, Google Chrome, DaVinci Resolve Studio with scripting enabled (and a Resolve MCP with execute_resolve_code), macOS for cutout.swift.

Step 0: read the scoreboard first

Before picking or cutting anything, read the reel tracker.

  1. Fill in stats for any row that has been live 7+ days and is still blank. Pull them from your analytics tool (e.g. Metricool getAnalyticsDataByMetrics for Instagram reels). Views, average watch time, 3-second hold if available, shares, saves, comments, follows.
  2. Compare by what we control: format, hook style (card / build / title), length, CTA type, punch-ins, topic. Only call something a pattern with at least 3 reels on each side, and say how big the gap is.
  3. Write the takeaway under "What the numbers say so far" in the tracker, dated.
  4. Apply it to this batch and say so in the reel list you show the user ("build hooks are holding 1.4x longer at 3s, so I used build on 6 of 8"). If a winner contradicts a locked style, ask before changing it.
  5. After scheduling, add a row per reel with how it was edited.

The rule that matters: hook first

The first 3 seconds decide distribution, up to half of viewers leave inside that window, and a hook has to load instantly rather than build. A reel that opens with "tip number two is..." is dead.

So do not cut a chronological excerpt. Build each reel as:

  1. HOOK (2-4s): the punchiest standalone line, taken from anywhere in the video, including the end.
  2. BUILD: the context and the explanation, in whatever order reads best.
  3. PAYOFF: a short closing line if one exists.

Chronological chunks produce weak hooks. Reordering the same footage, without re-shooting, fixes it. Reordering is the whole job.

Good hooks, as a pattern: "we don't skip leg day anymore", "a plan you skip is not a plan", "10 minutes a day is 60 hours a year". Short, declarative, no setup.

Where new footage lands

New reel batches show up in <YOUR_FOOTAGE_DRIVE>/<INBOX_FOLDER>/, often as untitled folder. Look there first. Once you know what's in it, rename the folder date first, then a name: 9-28-2026 gear reels. Transcribe before renaming, then fix the paths in clips.ndjson.

Standalone reels: animated style

For reels filmed as reels (no long video, no paper edit), the animated style is more engaging: product in frame, photo cards, an arrow or ring on a button, typewriter notes up top, small image pop-ins above the head, a big bold hook with a small subline. Build one first and get approval before doing the batch.

  1. No paper edit. Write segments-v2.json as [] and give every part "raw": true, "exact": true.
  2. Don't trust whole-clip whisper on these. Takes with long resets between lines drift 2-5s (whisper can place a phrase inside a 5.4s silence). Run GAP=0.18 SUF=.fine blobs.py <work> <clip> and cut parts on blob edges. For the last word, check a 50ms RMS map: a dip mid-word can be a fricative ("f" in "free"), so re-read the caption tail. End the last part after the word's tail drops below -50 dB. Whisper also invents words past the cut on the last card, so check the SRT's last cue against the audio.
  3. Punch-in rhythm: alternate parts between wide (zoom 1.25, tilt 0) and punched (zoom 1.6, tilt -90) on native-vertical footage. Pan any part where the speaker leans (e.g. hook: pan 180).
  4. Animation layer on V2: motion.json (element list timed to the words in reels-built.json) -> render_motion.mjs (frame by frame, ~50s for 30s). Element types are documented in motion.html: word builds, slam, typewriter notes, wobble, product cutouts (pop / fly-in / shrink), pulsing rings on buttons, a strike-through, and a b-roll card that plays real footage (graded with the LUT via ffmpeg lut3d). One visual per line, all in the band above the head (y 200-550; at y 100 the top words get cut off). To make that room, tilt every part down 100 more than the base framing (wide 1.25 / tilt -100, punch 1.6 / tilt -190). Key words use the blue class (it uses --brand), which draws a thick white border behind the fill so the accent doesn't vanish on a neutral wall. Place with resolve_captions.py, "track": 2.
  5. Product images: Wikimedia Commons API first, then brand OG images. Cut out with scripts/cutout, crop off table reflections. Keep them in <work>/gear/.
  6. Your footage for b-roll: <YOUR_FOOTAGE_DRIVE>/<shoot>/<cam>/full/. If the footage is log, use the full-res files plus the LUT; baked proxies often look crunchy.
  7. Sound effects on A2: build_sfx.py motion.json out.wav. It uses video-reels/sfx/<name>.wav when present (pop, whoosh, slam, tick, click, type, swipe, wobble, shrink), else a synthesized stand-in, so it works out of the box with no files. ding is always scripts/ding.wav (ships as a synthesized ding; replace it with your own pick sound). The synthesized pop is a camera shutter (a pitch-sweep pop tends to sound "alien"), and MASTER = -3 dB keeps every effect a touch under the voice. For better sounds, create sfx/ and drop in your own licensed effects, then tune per-file gains in LIBGAIN. With the Epidemic Sound connector: search with SearchSoundEffects, fetch with DownloadSoundEffect (WAV), then curl the assetUrl. 7b. Soft background music on A3: pick one vibe and stick with it (e.g. dreamy jazz / sax, no vocals). With Epidemic Sound: SearchRecordings with query.externalID = {type: SPOTIFY_TRACK, id: <REFERENCE_SPOTIFY_TRACK_ID>} and vocals: false, skip anything tagged weird/scary/drone/orchestral, and give each reel a different track. Then EditRecording with targetDurationMs = reel length + ~0.3s and forceDuration, so the track ends instead of getting chopped. Bake -21 dB and a 1.2s fade-out into the wav with ffmpeg, because the Resolve API can't set clip volume (SetProperty("Volume") returns False). That sits the bed about 15 dB under the voice.
  8. Preview before Resolve review: preview_reel.py (same crops, LUT, overlays, sfx, music) gives an mp4 you can watch in seconds. Trim each piece to Resolve's exact frame count (fps, then trim=end_frame, then atrim=duration); plain -ss/-t inputs run ~0.45s long and chop the last line off the preview. Proof it with a 1.3s contact sheet, and whisper the last 1.5s to confirm the ending is all there.
  9. Keep a sign-off line if the speaker adds one after the CTA: cut on the silence after it, not after the CTA word. build_reels.py pads 1s of silence after the audio, because whisper drops the last words from the captions otherwise. Re-renders go to new filenames (-v3) so Resolve can't serve a cached old clip.

Batches of standalone reels

Get ONE reel's style approved first, then run the rest in parallel:

  1. One folder per reel (~/Movies/reels-<MMDD>/<NN>-<slug>/) with symlinked <clip>.wav/.json, clips.ndjson, segments-v2.json = []. Separate folders stop reels overwriting each other's reels-built.json.
  2. Write a short brief pointing at the approved reel's files (motion.json, framing, sfx) and launch one background agent per reel, 4 at a time (more chokes a laptop on 4K decodes). Give each agent the clip's transcript, the product's exact model name, and any lines to cut (wrong facts, restarts, chatter).
  3. Agents stop at a no-music preview. Do music centrally (one edit per reel, a different track each), level to about mean -38 dB with a fade, render the final preview, then place in Resolve one reel at a time.
  4. Proof every agent's contact sheet yourself before building. Collect the final previews into one previews/ folder for review.
  5. Framing differs per setup. Example: wide zoom 1.2 / tilt -190, punch 1.5 / tilt -380 on one setup; 1.25/-100, 1.6/-190 on a looser one. Check stills per clip.
  6. Collect every agent's "to check" items (mishears, facts it cut, claims on screen) into your summary. Whisper commonly mishears fast speech (e.g. "free" as "for you").

Get every reel out of the video

Mine the whole video, never repeat content.

  1. Inventory first. Walk the whole transcript and list every idea unit that could stand alone: each tip, story, stat, script, confession, contrarian line, and customer win. Expect about one reel per 1.2-1.5 min of final cut.
  2. Look for these hook types:
    • a number: "300 loaves", "40% of my grocery bill", "5K to 50K followers"
    • a confession: "I used to just sit back and hope they'd reach out"
    • a contrarian line: "don't just give it away... that's a poor giveaway"
    • an insider line: "the secret to this that nobody does"
    • a customer win: a named result (with permission)
    • a script said out loud: the DM, the referral ask
    • a "never": "I've never had anybody tell me no to my face"
  3. The hook has to name the topic. "The secret nobody does" fails on its own because the viewer can't tell what it's about. Pick the line that says what the reel is about in the first second.
  4. No line in two reels. Two reels can share a theme if the hook and angle differ.
  5. Present it in tiers: the A list to build, the B list as spares, and the leftovers with a reason each ("second half of Reel 4", "no hook", "rambly on camera"). The user picks.
  6. Stop when it would repeat. Past that point every reel dilutes the others.

The shape every reel follows

"The hook needs to hook people in, and then the rest of the reel explains the hook and explains the value, the meat of it."

  1. Hook: stops the scroll and names the topic.
  2. Explain the hook: the next line answers the question the hook raised. If the hook is "I lost 20 pounds without a gym", what follows says how.
  3. The meat: the actual value, the how-to, the script, the proof. This is most of the reel.
  4. Stop on the payoff line.

A great hook followed by lines that don't pay it off is worse than a plain hook, because the viewer feels baited. Judge each reel on both halves: would the hook stop me, and does everything after it make sense and deliver what it promised?

Flow check: read it as a stranger before building

Write each reel out as plain sentences in order, then read it as someone who never saw the long video. Only then look up timings. These are the rules hand trims keep landing on:

  1. End on the payoff, once. Stop at the strongest line. Don't add a slogan after it or a second pass at the point.
  2. Nothing can point to something the reel didn't set up. "I just simply asked them and they almost always share it" with no "asked them what" in the reel has to go. Check every "she", "they", "this", "that" and "so" against what came before in the reel.
  3. Stories run in story order. A win as the spoken hook, with the person named later, confuses the listener. Name the person first, then the win. The text card does the hooking.
  4. One idea per reel. Two ideas plus a restatement should become one idea plus one proof.
  5. Cut asides about the speaker and garbles inside a part. Split the part around them.
  6. The spoken hook names the topic in words. "It's also an incredible way to reach a new community" leaves "it" undefined. "As you keep meal prepping on Sundays..." says what the reel is about.
  7. Target 20-35s. Anything past 40 needs a reason.

Ship it tight: no ums, no dead air, no restarts. Before building:

  • Lapel/live-room audio: run segments.py (from paper-edit) with NOISE_DB=-30. ffmpeg's silencedetect is peak-based, and room noise on a lapel peaks above the -36 default, so real 0.5-0.8s pauses go undetected at -36.
  • whisper drops "um" entirely, even with a filler prompt. Check each part's first voiced blob: a short (0.2-0.7s) isolated blob before the first real word is an um or a false start. Cut to the next pause.
  • Restarts inside a part show up in the caption read-through as doubled words ("from that from that", "it's not it's not"). Split the part around the first attempt.

Trailing verbal tics ("okay?", "right?"). Two traps: a part end on the sentence's end time keeps it, and a part start snapped 0.6s back pulls in the previous sentence's tic. Whisper's word times for these can be off by up to 0.8s, so find the real gap with silencedetect=noise=-40dB:d=0.06 over a 2s window, set the edge in the gap, and mark it "exact":"end" or "exact":"start". Then re-read the captions: the tic should be gone.

Lines cut from the long video: add "raw": true to a part to take it straight from the raw clip instead of intersecting it with the locked cut. Raw parts get their own pause trim (silencedetect -38dB, pauses over 0.5s cut to 0.22s).

Reel CTA clip: when the video has one (reel_cta in your state file), add it as the last part of every reel with "exact": true, and symlink its <clip>.json/.wav into the work dir so build_reels.py can read it. Import the MP4 with ImportMedia if the builder reports one clip fewer than the plan has segments.

When a transcribed word flips the meaning ("nobody" vs "everybody"), re-transcribe that 3-5s alone with a second of lead-in before trusting it.

Run it

Requires paper-edit to have run first (needs segments-v2.json and the whisper transcripts in the same work dir).

Separate lapel audio (live talks, events): run lapel-sync first. It fills the same work dir, and each line of clips.ndjson gets an "audio" field pointing at that clip's lapel slice. When that field is there, reel timelines take picture from the MP4 (mediaType: 1) and sound from the slice (mediaType: 2), same in and out points, because slice time equals clip time. Captions and pause-snapping already read the lapel through <work>/<clip>.wav and <clip>.json.

  1. Read the transcript and pick the moments. Write <work>/reels.json:
[{"name":"Reel 2 - skip the meal plan app","parts":[
  {"clip":"C0001","start":518.04,"end":520.60,"why":"HOOK: 'we don't skip leg day anymore'"},
  {"clip":"C0001","start":308.08,"end":316.40,"why":"my goal: book a call or tell me no"},
  {"clip":"C0001","start":176.68,"end":183.40,"why":"most people follow up once or twice"}]}]

Times are seconds in the raw clip. Each part becomes a yellow marker on the timeline so the structure is visible.

  1. Build segments and captions:
python3 <skill>/scripts/build_reels.py "<work>" "<out dir for SRTs>"

Pause trims are inherited from segments-v2.json (or SEGMENTS=segments-locked.json after a hand trim of v2), so reels match the long cut.

  • Part edges snap to the nearest pause, not to whisper's word times. Those run late on the long file (~0.4s), so starts search 0.6s back and ends 0.35s back. Add "exact": true to a part to keep an edge set by hand.
  • Each reel's audio gets a 1s lead-in before transcription, because whisper drops the first word when the audio starts abruptly.
  • Captions come from re-transcribing each assembled reel, so they are exact and no edge words get dropped. Known mishears go in <work>/caption-fixes.json ({"JON":"JOHN","CUSTOM":"COSTUME"}). Take people's names from your own notes, never from whisper.
  • Read every reel's full caption text before building in Resolve. It is the fastest way to hear the cut: a stray word at the start or end ("CLIENTS. SO ASKING...") means an edge bled into the next sentence, and a missing word means it clipped. Nudge that part's start/end by 0.1-0.3s and rebuild.
  1. Build the timelines in Resolve:
PARAMS = "<work>/reel-params.json"
exec(open("<skill>/scripts/resolve_reels.py").read())

reel-params.json: {"work":"<work>","files":[...],"lut":"<YOUR_LUT_FOLDER>/<YOUR_LUT>.cube","zoom":3.2,"pan":150,"tilt":-30,"replace":true}. Optional "grade_from" names the timeline whose hand grade gets copied onto every reel clip.

  1. Render the hook cards (see below), then place them:
OUT_DIR="$HOME/Movies/<video>-hooks" HOOKS="<work>/hooks.json" node <skill>/scripts/render_hooks.mjs
PARAMS = "<work>/hook-params.json"
exec(open("<skill>/scripts/resolve_hooks.py").read())
  1. Render and place the captions. They go on as alpha clips, not as a subtitle track:
OUT_DIR="$HOME/Movies/<video>-captions" SRTS="<srt1>,<srt2>" node <skill>/scripts/render_captions.mjs
PARAMS = "<work>/caption-params.json"
exec(open("<skill>/scripts/resolve_captions.py").read())

Now the reel is finished on the timeline: footage on V1, hook card on V2, captions on V3. Trim, then render straight out. No subtitle import, no burn-in setting.

Captions as overlay clips, not a subtitle track

Resolve cannot import an SRT through the API, and burn-in has two separate switches that both fail silently. So captions are rendered the same way as the overlays: one full-length 1080x1920 alpha .mov per reel, dropped on V3 at frame 0 and trimmed to the timeline length.

It is fast because captions are static: render_captions.mjs screenshots one PNG per cue, then assembles them with ffmpeg's concat demuxer at each cue's duration. A 30s reel costs ~27 screenshots instead of ~720 frames, about 15 seconds of wall clock.

  • Default style: Montserrat ExtraBold (800), white, 64px, uppercase, no shadow and no outline. Change to taste in caption.html. Montserrat is free from Google Fonts; install it in ~/Library/Fonts.
  • Position 0.74 of frame height, which is the lower third and clear of the Instagram UI.
  • Override with POS and SIZE env vars.
  • Trim to the timeline, never past it. The mov is slightly longer than the reel because the concat tail needs a final entry.
  • To reword: edit the SRT and re-render that one reel. Do not retype into Resolve, or the next run overwrites it.

Fallback, to type captions by hand: import the SRT by hand (File > Import > Subtitle) and style in the Inspector at ~72-78% vertical. Read the timecode trap below first, because it bites every time.

The SRT trap: timeline start timecode

Reel timelines must start at 00:00:00:00. Resolve defaults every new timeline to 01:00:00:00, and SRT cues start at 00:00:00,000. Import an SRT onto a 01:00:00:00 timeline and Resolve places the subtitles one hour before the first frame: the track is created, it looks empty, and nothing appears in the viewer. It looks exactly like a failed import.

resolve_reels.py calls SetStartTimecode("00:00:00:00") on every reel timeline. To repair timelines built before that:

for i in range(project.GetTimelineCount()):
    t = project.GetTimelineByIndex(i+1)
    if t.GetName().startswith("Reel"):
        t.SetStartTimecode("00:00:00:00")

Clips keep their positions; only the ruler changes. Re-import the SRT afterwards.

Burning in a hand-imported subtitle track, in Deliver:

  1. Make sure the subtitle track's eye icon is on in the Edit timeline. A hidden track renders blank.
  2. Deliver > Subtitle Settings > tick Export Subtitle, then set Format to Burn into Video. If you leave it on "As a separate file" you get a sidecar .srt and clean video, which is what Instagram cannot read.
  3. Render 1080x1920 H.264. Check the first few frames of the output before scheduling.

Hook style: card or build (ask which, or split the batch)

Two ways to open a reel. Keep both until your numbers pick a winner.

stylewhat it iswhen
cardthe white hook card below: all caps, top of frame, held 5s (talking heads: size 52, held 7s)the default
buildthe opening question builds on screen word by word as it is said, the key words land big in the accent color with the pick sound, and the shot punches in ~12% on the question, then cuts back to normal framing on the answer. Clears at 7s. The same look as the CTA card at the end.talking heads that open on a question or a quote ("do you really need a sales call?")

How to build it (it rides the this-or-that overlay, render_thisorthat.mjs, on V2; no V4 card):

  1. Split the opening part at the pause after the question, and give the question part "zoom": <framing + 0.12>, "tilt": -60 in reels.json. build_reels.py passes per-part zoom/pan/tilt through to Resolve.
  2. Add this as the first item of the overlay plan, with q = that part's why:
{"q":"HOOK question","cta":{"pre":"do you really need a","word":"sales call?","post":"",
  "size":150,"preSize":64,"top":880,"until":7.0,
  "say":["do","you","really","need","a","sales"]}}

pre builds in white, word lands huge in the accent color with the pick sound, post builds in italics under it. say lists the spoken word that reveals each token, in order (use the words actually said). Long key words need a smaller size (two long words fit at 112). top 880 keeps it on the chest on close-ups. 3. Build the timeline without the V4 hook card.

Tracking the test: log which reels shipped with which style in the reel tracker (date, reel, style), so an analytics pull can compare 3-second hold and views later.

Text hook cards

For the first 5 seconds, a held white card over the top of the frame, so a muted scroller reads the promise before a word is said. Same renderer family as video-overlays, sized 1080x1920 with alpha.

<work>/hooks.json:

[{"id":"reel1","text":"No Sale Is Real Until They Pay","dur":5.0}]

Default style: white rounded card with dark text in the display font, all caps, 64px, card top at 165px, which sits just above the head at the standard reel framing. Set the font with --display-font in hook.html.

Optional per hook: size, top, lh (line height).

  • Six words or fewer, written in Title Case in hooks.json (the template renders it in caps). It is a promise, not a sentence. Match the reel's spoken hook without quoting it word for word.
  • Default position clears the head at the standard reel framing (zoom 3.2 / tilt -30, pan set per shoot). Verify with export_current_frame at ~1.5s if the framing changed.
  • 5 seconds, and it goes on V2, so the footage runs underneath untouched. Set "dur":5.0 in every hooks.json entry (3s reads as too short).
  • If the card length changes after reels are built, do it the right way: swap the V2 clip in each Resolve timeline (keep the trims, never touch V3/V4) and re-render from Resolve. Never overlay a new card onto an already-rendered file.
  • Keep it separate from the captions. The card is the promise; the captions are the words.

Framing numbers

These are for landscape 3840x2160 footage. Native-vertical footage (a camera shooting vertical, which Resolve reads as 2160x3840) is different: it already fills 1080x1920 at zoom 1.0, so there is no 3.16 crop. Zoom is a punch-in on the 2160x3840 frame, up to about 2x with no resolution loss. At a live talk the speaker walks around, so set pan per part, not once per shoot, and check each with export_current_frame.

  • Zoom 3.2, not 3.16. 3.16 makes the picture 1919.7px tall, so any tilt uncovers a black strip at the top (8px at tilt -26). 3.2 is the widest zoom that survives the tilt. Check the top rows of an exported still for black before handing reels over.
  • Zoom 3.16 is the widest possible that still fills 1080x1920 from a 3840x2160 source. Less than that gives black bars.
  • Pan changes with each shoot day, so check it on every video (typical range 95-150 for a speaker slightly off centre). Lower pan moves the subject left in the frame.
  • Tilt -30 puts the eyes near the upper third and leaves the lower third clear for captions.
  • Verify with export_current_frame and draw a centre line on it before declaring it centred. Don't eyeball it from memory.
  • No resolution is lost: a 1215x2160 slice of 4K downsampling into 1080x1920 is still oversampled.

Also

  • Apply the LUT (your camera LUT on node 1). Raw log footage looks grey and flat.
  • Then copy the hand grade from v2 onto every reel clip. v2_item.CopyGrades(all_reel_V1_items) works across timelines and leaves zoom/pan/tilt alone. resolve_reels.py does this when grade_from names the graded timeline. Verify each item's node 1 shows the grade tools, not just the LUT. Include the CTA clip: same shoot, same grade.
  • Never carry the YouTube overlays into a reel. Build reels from the raw clips, which is what these scripts do.

Sources

This-or-that reels (two voices, products on the desk)

One person asks off camera ("Canon or Sony?"), the speaker picks. If you keep scripts and on-screen wording in a doc (e.g. Canva <YOUR_CANVA_SCRIPTS_DOC_ID>), read it for each reel's title, audience line and edition name.

1. Speech map, not whole-clip whisper. An off-mic voice on the same lapel track sits ~8 dB under the main voice. Whole-clip whisper puts those questions up to 5s off or drops them. Run scripts/blobs.py <work> <clips>: energy finds every voiced stretch, each is transcribed alone. Parts in reels.json take blob edges with "exact": true; questions carry "say": "canon or sony" (written caption, since whisper can't hear the off-mic voice reliably).

2. Captions: NORM=1 CASE=lower WORDS=5 build_reels.py (part-by-part transcription, lowercase, phrase-aware cards), then POS=0.56 SIZE=50 render_captions.mjs.

3. Overlay on V2: product cutouts from product shots (brand CDNs, Wikimedia Commons originals; retailer pages usually block bots) cut with scripts/cutout (swiftc -O cutout.swift -o cutout, macOS subject lift). Plan -> build_thisorthat.py <work> plan.json spec.json -> render_thisorthat.mjs. Place with resolve_captions.py and "track": 2. The builder also writes dings-<reel>.wav for A2.

Layout defaults (change to taste):

  • Title white (an accent-colored title can vanish on a busy background), audience line above it (e.g. "for beginners"), edition in italic under it. Title leaves after 10s (title_for).
  • Instagram safe zone: nothing below y 1500 (caption, handle, audio bar cover the bottom ~420px) and nothing right of x 950 past y 1000 (like/comment/share buttons). Top ~200px is clear.
  • Products sit on the desk: zoom the footage (native vertical, 2160x3840) to 1.2 with Tilt 160 so the desk edge rises to ~1290 and labels (y 1225) and products (bottom 1490) land on it. Positive Tilt moves the picture up; past (zoom-1)*960 it shows black.
  • The pick label turns the --brand color and bounces (8-frame scale pop), with a pick sound on each (scripts/ding.wav; ding-synth.wav is the synthesized original to fall back to).
  • Comment-keyword CTA card: when the speaker gives a comment CTA, a card as big as the opening title: "COMMENT" (white Montserrat), the keyword huge in the accent color and display font, a short italic line under it. Every word appears the moment it is said and pops; the keyword gets the pick sound. Plan item: {"q":"CTA: GUIDE","cta":{"pre":"comment","word":"guide","post":"for my free guide","say":["comment","guide","for","my","free","guide"]}}. Needs the words list build_reels.py saves in reels-built.json.
  • The user may trim while you work. Before re-placing anything on a reel that has been touched, export the live V1 layout to tlmap-<reel>.json (start/dur/left per item + frames) and set "tlmap" in the plan: build_thisorthat.py remaps every state and pick sound onto the trimmed cut. Never re-place the V3 captions over hand trims; they ripple with the edit already.
  • Talking heads in the same batch: the hook goes in the classic white hook card at the top (render_hooks.mjs, all caps, size 52, 7s), placed on V4 with resolve_captions.py "track": 4. No big title on these. The overlay on V2 carries only the CTA card (layout {"ctaTop":880,...}, on the chest since close-ups fill the top half); captions at POS=0.72. Follow-only CTAs: "cta":{"pre":"","word":"follow","post":"for something free"}. This-or-that follow CTAs can use a Following / Not Following pair with a screenshot of your own profile card in gear/profile-card.png.
  • Resolve API: timeline.DeleteClips only works on the CURRENT timeline (project.SetCurrentTimeline first); otherwise it returns True and does nothing. Name keys as "Reel 1 -" not "Reel 1", or Reel 1 matches Reel 10-12.

Individual skills in this repo

This repo contains 7 individual skills — each has its own dedicated page.

beachyphotoandfilm-arch/video-skills

Sync camera footage to a separate lapel or field-recorder track (e.g. Tascam DR-10L and similar) and lay it up in DaVinci Resolve. Stitches the recorder's split files, finds each clip's exact place in the audio by matching words then waveforms, cuts one lapel slice per clip so every later skill works unchanged, checks for clipping, and builds a synced timeline in a new Resolve project. Trigger on "sync the lapel", "sync my audio", "line up the recorder", "I used a lav", or any shoot with a separate audio folder (live talks, events, workshops). Run it before paper-edit or video-reels.

beachyphotoandfilm-arch/video-skills

Turn raw talking-head footage into an assembled rough cut inside DaVinci Resolve. Pulls the clips off your footage drive, transcribes them locally, removes failed takes, restarts, asides and dead air, then creates a new Resolve project with two paper-edit timelines ready for your fine trim. Trigger on "paper edit", "cut my raw footage", "make a rough cut", "clean up this video", or when the user points at a folder of raw camera files for a YouTube video.

beachyphotoandfilm-arch/video-skills

Design a unique Instagram cover for each reel in your house style. Reads what the reel is about, writes a short hook headline, picks one of your own photos from a tagged photo library, lays it out like your existing covers, and proofs it in grid view at phone size. Renders locally, never with AI image generation. Can be called by a scheduling skill before reels are scheduled. Trigger on "make covers", "reel covers", "cover photo for this reel", "design the IG covers".

beachyphotoandfilm-arch/video-skills

Render finished reels out of DaVinci Resolve, write captions in the creator's voice, attach a custom Instagram cover from reel-covers, and schedule them to Instagram, TikTok and YouTube Shorts through Metricool. Handles the Drive upload hop, collision checks against the existing calendar, and cleanup. Trigger on "schedule the reels", "post these", "put these on the calendar", or after reels are approved.

beachyphotoandfilm-arch/video-skills

Build animated brand overlays (motion graphics) for a talking-head video, timed to the speaker's exact words, plus a DaVinci Resolve import file. House style is in-scene type (text beside and behind the speaker's head) plus diagrams. Transcribes the rough cut locally, picks the moments, renders transparent .mov clips, and writes an FCPXML that drops every clip onto the timeline already in position. Trigger on "make animations for this video", "add overlays/motion graphics", "animate this", "b-roll graphics for my YouTube video", or when the user shares a rough cut and asks for graphics.

beachyphotoandfilm-arch/video-skills

Design YouTube thumbnails and titles for a talking-head video. Pulls graded full-res frames out of the raw footage, renders brand-styled variants in two proven layouts, and proofs them at phone size where the click actually gets decided. Trigger on "make thumbnails", "thumbnail for this video", "title and thumbnail", or after a video is cut.

beachyphotoandfilm-arch/video-skills

Run the whole YouTube video pipeline end to end, one stage at a time, stopping for the creator's approval at every gate. Paper edit, punch-ins, reels, reel covers, scheduling the reels, overlays, thumbnail and title, then scheduling the long-form video, all through Metricool. Keeps state per video so you can stop anywhere and pick up later. Trigger on "/youtube", "let's do the whole video", "run the video pipeline", "continue the <name> video", or when handed raw footage for a YouTube video.

Verwandte Skills