Communitygithub.com

beachyphotoandfilm-arch/video-skills

Build animated brand overlays (motion graphics) for a talking-head video, timed to the speaker's exact words, plus a DaVinci Resolve import file. House style is in-scene type (text beside and behind the speaker's head) plus diagrams. Transcribes the rough cut locally, picks the moments, renders transparent .mov clips, and writes an FCPXML that drops every clip onto the timeline already in position. Trigger on "make animations for this video", "add overlays/motion graphics", "animate this", "b-roll graphics for my YouTube video", or when the user shares a rough cut and asks for graphics.

Was ist video-skills?

video-skills is a Claude Code agent skill that build animated brand overlays (motion graphics) for a talking-head video, timed to the speaker's exact words, plus a DaVinci Resolve import file. House style is in-scene type (text beside and behind the speaker's head) plus diagrams. Transcribes the rough cut locally, picks the moments, renders transparent .mov clips, and writes an FCPXML that drops every clip onto the timeline already in position. Trigger on "make animations for this video", "add overlays/motion graphics", "animate this", "b-roll graphics for my YouTube video", or when the user shares a rough cut and asks for graphics.

Funktioniert mit✓Claude Code~Codex CLI~Cursor
npx skills add https://github.com/beachyphotoandfilm-arch/video-skills/tree/HEAD/skills/video-overlays

In Ihrer bevorzugten KI fragen

Öffnet einen neuen Chat, in dem dieser Agent-Skill bereits geladen ist.

Dokumentation

Video overlays

Turns a rough cut into a set of animated overlays that land on the exact frame the speaker says each thing, plus a Resolve import file.

Set your brand here

Edit these once, in scripts/scene.html (the :root block at the top). Everything else reads from them.

settingCSS variabledefault
accent color--brand-accent (+ --brand-accent-rgb as r,g,b)#4a90e2
light text color--brand-light (+ --brand-light-rgb)#f2f2f2
dark ink color--brand-ink#0d131a
display font (headlines, numbers)--font-display'Anton','Bebas Neue',Impact,sans-serif
body font--font-body'Inter',-apple-system,sans-serif

The display font must be installed locally (e.g. in ~/Library/Fonts), since clips render offline in headless Chrome. Anton and Bebas Neue are free on Google Fonts. The in-code name BLUE just means "the accent color".

House style: in-scene type. No cards. Text lives in the room next to the speaker's head, big titles flank the head, questions type on as they are said. Always your brand colors and display font.

The one rule that matters most: keep graphics behind the speaker's head or to the side. Nothing in front of the face, and no background color. Diagrams live on the wall beside the speaker too. Everything is matted, so hands pass in front of the graphics.

One run produces: N transparent .mov clips, PLACEMENT.txt, two .fcpxml files, and a short PREVIEW.mp4.

No sound by default. Don't add music or sound effects unless asked.

Style defaults (change to taste)

  1. Nothing appears before it's said. Anchor every clip to the word-level timestamp of the phrase, plus ~0.12s so it lands just after the words. Section-level timing reads as early.
  2. Short. 2.5 to 6.5 seconds. Most around 3. It should cut back and forth like social, not sit there.
  3. Behind the head or to the side, never over the speaker. No cards, no backgrounds, no dimming. That retires the corner chips, every full-* card and define for new videos.
  4. Still show the mechanism. Reach for a diagram before a sentence. Use the in-scene diagram types (dots, flow, bars, compare, math), which sit on the wall beside the speaker.
  5. Split anything with two halves. Old way and new way are two clips. 10% and 20% are two clips. Three options are three clips, each landing as it's named.
  6. Every line lands on its word. Lists, side lines and typed words carry per-line offsets (at, t) from the word timestamps, not an even stagger. scripts/plan.example.py shows how: quote the phrase, name the cue words, and it does the lookup.
  7. A good mix (about 48 clips over 11 minutes): roughly a quarter in-scene diagrams (mostly flow), then side, pair, type, list2, num. Save pair for the punchlines.
  8. Big titles go around the head, not behind it. pair splits each line at a word boundary into two halves that flank the head, with a gap of 1.9x the face-box width (hair and ears included), scaled so both halves stay on screen. So every pair line needs 2+ words ("ONE HOUR", not "HOUR"): a single word can't split and ends up hidden. "around": false in data brings back the centred look. The outlined line is 150px with a 5px stroke and a light accent fill (35%) so it stays readable.

Before you start: is the cut locked?

Overlays are anchored to exact words, so any re-trim afterwards shifts every clip. If the user hands over raw footage, or the cut isn't locked, run the paper-edit skill first and place overlays on the timeline it produces.

Pipeline

1. Get the cut and probe it

Ask for an export from Resolve (e.g. to ~/Movies/), or raw footage on <YOUR_FOOTAGE_DRIVE> for a paper edit first. A video file is fine; no need to export audio separately. When Resolve is open, read the timeline through the MCP instead of asking for an export.

ffprobe -v error -show_entries format=duration -show_entries stream=codec_type,r_frame_rate -of default=nw=1 "<file>"

Note the frame rate (the renderer assumes 23.976 = 24000/1001; change FPSN/FPSD in render.mjs for other rates) and ask whether the timeline starts at 00:00:00:00 or 01:00:00:00 (Resolve's default is 01).

2. Transcribe with word-level timestamps

ffmpeg -y -i "<file>" -ac 1 -ar 16000 -c:a pcm_s16le cut.wav
whisper-cli -m ~/.cache/whisper-models/ggml-small.en.bin -f cut.wav -oj -ml 1 -sow -of cut

-ml 1 -sow gives one token per segment, which is what makes exact anchoring possible. Then find the start time of each phrase by matching token runs in cut.json (transcription[].offsets.from is ms).

3. Pick the moments

Read the whole transcript first. Aim for roughly one clip per 15 seconds of video. For each: the anchor phrase, a type, and the shortest duration that lets the animation land and breathe.

Ground the content in what was actually said, not in written notes. If notes say "30 minutes" and on camera it's "half an hour or so", the recording wins, and flag the difference to the user.

4. Write clips.json

Copy scripts/clips.example.json as the model. One object per clip:

{"id":"06-math-10","start":52.96,"dur":3.6,"type":"dots","cue":"\"if 10 out of 100 meals are home cooked\"",
 "data":{"lit":10,"pct":"10%","count":"10 meals"}}

Keep it outside the skill folder (a scratchpad path is fine) and pass it with CLIPS=. start is the raw phrase time in seconds; the renderer adds the 0.12s lead itself. id order sets file order, so number them.

5. Render a clean cut (needed by the in-scene types)

The in-scene types read the head position and cut the speaker out from a copy of the cut with no overlays on it. Render the locked timeline from Resolve with every overlay clip disabled, H.264 1080p, into the scratchpad. In the API, SetTrackEnable on V2 did not take in Resolve 19; item.SetClipEnabled(False) on each clip in the track did. The raw footage has to be online, so the footage drive must be mounted; if it isn't, ask the user to plug it in and watch for the mount rather than rendering a finished master (old overlays would be baked into it).

6. Render

cd <skill>/scripts && npm i                   # first time only, installs puppeteer-core
swiftc -O person.swift -o person              # first time only, the matte + head finder
TITLE="<video name>" CLIPS=/path/to/clips.json FOOTAGE=/path/to/clean.mp4 OUT_DIR="$HOME/Movies/<video-name>-overlays" node render.mjs
# one or a few clips only:
ONLY=06-math-10,07-math-30 OUT_DIR=... node render.mjs

Roughly 10 to 15 seconds of wall clock per second of overlay (pair is slower, it mattes every frame). Backgrounding it and checking in is the right move. It writes PLACEMENT.txt and both .fcpxml files at the end.

Always proof over the real footage before a full render. node proof.mjs <clean.mp4> <clips.json> out.jpg [email protected] [email protected] ... renders one frame per clip (prefix@seconds-into-clip) over the actual shot, matte included, tiled into one image. Look at it. Run-off-frame text, a title crossing the eyes, text over a busy background: all invisible in code and obvious in the image. For diagrams, compositing over grey is still fine.

Proof frames: extract to PNG first. Overlaying a qtrle .mov straight onto a lavfi color source with -ss gives a blank frame. Pull the frame out (ffmpeg -ss 2.8 -i clip.mov -frames:v 1 f.png), then composite the PNG.

Placing in Resolve: duplicate the locked timeline as YouTube Final, import the movs into an Overlays bin, and AppendToTimeline each with trackIndex: 2 and recordFrame from the IN timecode in PLACEMENT.txt (((h*3600+m*60+s)*24+ff), so 01:00:00:00 = 86400). Parse PLACEMENT.txt one line at a time (line.split()[0] is the id, [1] the timecode). A whole-file regex can read cue text like "60-second" as a clip id and steal the next line's timecode, so one clip silently goes unplaced. Check that the placed count equals the clip count.

7. Preview, then hand off

Build a 10-15 second preview over the real footage so the user can judge before it goes in:

ffmpeg -y -ss <t> -t 11 -i "<cut>" -itsoffset <a> -i "<clip1>.mov" -itsoffset <b> -i "<clip2>.mov" \
  -filter_complex "[0:v][1:v]overlay=0:0:eof_action=pass[v1];[v1][2:v]overlay=0:0:eof_action=pass[vo];[vo]scale=1280:-2[v]" \
  -map "[v]" -map 0:a -c:v libx264 -crf 20 -preset veryfast -c:a aac PREVIEW.mp4

The original audio comes along with the footage, which is all the preview needs. Then give the import steps (File > Import > Timeline > Import AAF, EDL, XML, pick the starts-01 file, copy the clips up onto a track above the footage at 01:00:00:00).

Working inside Resolve directly (preferred when Resolve is open)

If a resolve MCP server is installed (for example the third-party barckley75/resolve-claude-mcp), it talks to Resolve Studio's scripting API. If the paper-edit skill ships an ensure_resolve.sh, run it first so Resolve is launched if it's closed. With Preferences > System > General > External scripting using: Local, prefer it over the export-and-import dance.

What it changes:

  • No exported cut needed. Read the open timeline for its real timecodes, frame rate, start timecode and clip layout, then anchor overlays to that. This fixes "timecodes shift every time I re-trim."
  • No manual import. Place the rendered .mov clips onto a video track directly.
  • Better moment picking. Seeing where the cuts already fall means overlays can land with the edit instead of fighting it.
  • It can also kick off the render when the overlays are in.

Rules for using it:

  1. Read before you write. Confirm project name, timeline name, start timecode and frame rate first, and say them back to the user.
  2. Never write to the master timeline. Duplicate it first and work on the copy, so a mistake costs nothing.
  3. Avoid execute_resolve_code unless a normal tool can't do the job. It runs arbitrary Python against the user's projects. If it is the only way, show the code first.
  4. Local sessions only. Resolve has to be running on this Mac, so this never works in a cloud or scheduled run. Fall back to the FCPXML there.
  5. The FCPXML path still works and is still the fallback. Keep generating both files; they cost seconds.

Audio transcription: some Resolve MCP servers ship their own mlx-whisper transcription, but the whisper.cpp step above is proven with this pipeline.

Clip types

Everything below is retired for new videos (kept so old videos can be re-cut). Use the in-scene types in the next section. Full-frame types paint a solid ink background with a soft accent glow, so they hide the speaker completely. Chips sit bottom-left.

typeuse it fordata
full-dotsretired, use dotslit, pct, count
full-flowretired, use flowlabel, steps[], good
full-compareretired, use comparetitle, left{label,value,sub}, right{...,hi}, vs
full-barsretired, use barstitle, bars[{label,value,value_label,hi}]
full-mathretired, use mathparts[] (operators x = + → render in the accent color), sub
full-packretired, use bars or sidename, note, size 0-1, hi
full-statretired, use pairlabel, value, sub, hi
full-titleretired, use pair or typetitle (<br> allowed), size, sub
full-listretired, use list2title, items[], step seconds
chipretired, use side or numnum, title, sub, hi
chip-stackretired, use list2items[]

The retired types still render, for re-cutting old videos.

In-scene types (the default)

No cards, text lives in the room next to the speaker's head, and big titles go around the head.

typeuse it fordata
paira punchline title: outlined line above the head, solid line at forehead heighttop, bottom, outline (top/bottom)
side2-3 short lines beside the head, each landing on its wordlines[], accent (index, big + accent color), at[]
typea quote or question typed on word by word as it's saidwords[{w,t,hi,br}]
defineretired: dims the shot. Use side or pairword, sub
numoutlined number draws on, words beside it ("01 plan the week")num, lines[]
list2items building on the left and/or right of the headleft[], right[], at[] (in order, left first)

In-scene diagrams (same wall placement and matte):

typeuse it fordata
dotsa rate out of 100, grid fills beside the speakerlit, pct, count
flowa process, steps chained top to bottom, last in the accent colorlabel, steps[], at[] optional
barstwo or three numbers by sizetitle, bars[{label,value,value_label,hi}]
comparetwo options, one each side of the headleft{label,value,sub}, right{...,hi}
mathan equation read down the wall, operators lead each lineparts[], sub

at and t are seconds from clip start; build them from the word timestamps so every line lands on its word.

These need FOOTAGE=, the cut with no overlays (render the timeline from Resolve with overlay clips disabled). The renderer reads the face box and a background brightness grid for each clip, so text finds the empty wall, flips side or shrinks if it would run off frame, and switches to ink on a light background and light text on a dark one. It mattes the speaker out of every frame of every in-scene clip (Apple Vision person segmentation, person.swift), so titles sit behind the head and hands pass in front of the graphics. Compile once: swiftc -O person.swift -o person. Without FOOTAGE they still render, just with a guessed head position and no matte.

Adding a type: add a function to the R registry in scene.html, keyed by name. It gets T (time), D (duration) and DATA. Use life(in_) for opacity, pr(start,dur) for progress, backdrop() for full-frame.

Output format

Output is qtrle (QuickTime Animation) with a real alpha channel: lossless, editor-friendly, and far smaller than ProRes 4444 for flat graphics. Brand colors and fonts are set in the :root block of scene.html (see "Set your brand here").

Gotchas

  • --force-device-scale-factor=1 and omitBackground: true on every screenshot, or the alpha comes back wrong.
  • qtrle needs -pix_fmt argb. Without it the alpha is silently dropped.
  • Timecode at 23.976 is counted at 24 frames per second of timecode: frames = round(seconds / 1.001 * 24).
  • Chrome path: render.mjs and proof.mjs launch /Applications/Google Chrome.app/.... Change the path on other systems.
  • Don't move the output folder after Resolve imports it, or every clip goes offline.
  • white-space:nowrap on chip text. Wrapping looks broken at this size.
  • Paths with spaces: build filesystem paths with fileURLToPath(import.meta.url), not new URL(...).pathname (that one is URL-encoded and spawn can't find the file).
  • The face box starts at the eyebrows, not the hairline. pair places its solid line at head.y - 0.1*head.h so it runs at forehead height; anchoring at the box center put the words across the eyes.
  • Light vs dark backgrounds. White text vanishes on a light wall, so in-scene text samples the brightness behind itself: ink on light areas, light text on dark areas.
  • The FCPXML places everything on lane="1" inside one gap. Resolve imports it as a standalone timeline; copy the track up onto the edit.

When the cut changes

Timecodes shift with every trim. For diagrams, only the start values in clips.json need updating; the clips don't need re-rendering. In-scene clips do (side, pair, type, num, list2): their placement and matte come from the footage under them, so re-render the clean cut and re-render those clips with ONLY=. Re-run with ONLY= set to nothing and it will rebuild the FCPXML from the existing files in seconds.

Related

  • paper-edit for the rough cut that comes first
  • Your own voice/style notes for how on-screen copy should sound

Individual skills in this repo

This repo contains 7 individual skills — each has its own dedicated page.

beachyphotoandfilm-arch/video-skills

Sync camera footage to a separate lapel or field-recorder track (e.g. Tascam DR-10L and similar) and lay it up in DaVinci Resolve. Stitches the recorder's split files, finds each clip's exact place in the audio by matching words then waveforms, cuts one lapel slice per clip so every later skill works unchanged, checks for clipping, and builds a synced timeline in a new Resolve project. Trigger on "sync the lapel", "sync my audio", "line up the recorder", "I used a lav", or any shoot with a separate audio folder (live talks, events, workshops). Run it before paper-edit or video-reels.

beachyphotoandfilm-arch/video-skills

Turn raw talking-head footage into an assembled rough cut inside DaVinci Resolve. Pulls the clips off your footage drive, transcribes them locally, removes failed takes, restarts, asides and dead air, then creates a new Resolve project with two paper-edit timelines ready for your fine trim. Trigger on "paper edit", "cut my raw footage", "make a rough cut", "clean up this video", or when the user points at a folder of raw camera files for a YouTube video.

beachyphotoandfilm-arch/video-skills

Design a unique Instagram cover for each reel in your house style. Reads what the reel is about, writes a short hook headline, picks one of your own photos from a tagged photo library, lays it out like your existing covers, and proofs it in grid view at phone size. Renders locally, never with AI image generation. Can be called by a scheduling skill before reels are scheduled. Trigger on "make covers", "reel covers", "cover photo for this reel", "design the IG covers".

beachyphotoandfilm-arch/video-skills

Render finished reels out of DaVinci Resolve, write captions in the creator's voice, attach a custom Instagram cover from reel-covers, and schedule them to Instagram, TikTok and YouTube Shorts through Metricool. Handles the Drive upload hop, collision checks against the existing calendar, and cleanup. Trigger on "schedule the reels", "post these", "put these on the calendar", or after reels are approved.

beachyphotoandfilm-arch/video-skills

Cut vertical Instagram reels out of a long-form talking-head video, hook first. Picks the strongest standalone moments from the transcript, opens each reel on its punchiest line, builds vertical timelines in DaVinci Resolve framed on the speaker's face with the LUT applied, and finishes them: a hook (a held text card, or a word-by-word "build" hook with punch-in) and burned-look captions on V3; this-or-that reels get product graphics and a comment-keyword CTA card. You trim and render, nothing to import. Trigger on "make reels", "clip this for Instagram", "cut some verticals", or after a YouTube video is cut.

beachyphotoandfilm-arch/video-skills

Design YouTube thumbnails and titles for a talking-head video. Pulls graded full-res frames out of the raw footage, renders brand-styled variants in two proven layouts, and proofs them at phone size where the click actually gets decided. Trigger on "make thumbnails", "thumbnail for this video", "title and thumbnail", or after a video is cut.

beachyphotoandfilm-arch/video-skills

Run the whole YouTube video pipeline end to end, one stage at a time, stopping for the creator's approval at every gate. Paper edit, punch-ins, reels, reel covers, scheduling the reels, overlays, thumbnail and title, then scheduling the long-form video, all through Metricool. Keeps state per video so you can stop anywhere and pick up later. Trigger on "/youtube", "let's do the whole video", "run the video pipeline", "continue the <name> video", or when handed raw footage for a YouTube video.

Verwandte Skills