Video overlays
Turns a rough cut into a set of animated overlays that land on the exact frame the speaker says each thing, plus a Resolve import file.
Set your brand here
Edit these once, in scripts/scene.html (the :root block at the top). Everything else reads from them.
| setting | CSS variable | default |
|---|---|---|
| accent color | --brand-accent (+ --brand-accent-rgb as r,g,b) | #4a90e2 |
| light text color | --brand-light (+ --brand-light-rgb) | #f2f2f2 |
| dark ink color | --brand-ink | #0d131a |
| display font (headlines, numbers) | --font-display | 'Anton','Bebas Neue',Impact,sans-serif |
| body font | --font-body | 'Inter',-apple-system,sans-serif |
The display font must be installed locally (e.g. in ~/Library/Fonts), since clips render offline in headless Chrome. Anton and Bebas Neue are free on Google Fonts. The in-code name BLUE just means "the accent color".
House style: in-scene type. No cards. Text lives in the room next to the speaker's head, big titles flank the head, questions type on as they are said. Always your brand colors and display font.
The one rule that matters most: keep graphics behind the speaker's head or to the side. Nothing in front of the face, and no background color. Diagrams live on the wall beside the speaker too. Everything is matted, so hands pass in front of the graphics.
One run produces: N transparent .mov clips, PLACEMENT.txt, two .fcpxml files, and a short PREVIEW.mp4.
No sound by default. Don't add music or sound effects unless asked.
Style defaults (change to taste)
- Nothing appears before it's said. Anchor every clip to the word-level timestamp of the phrase, plus ~0.12s so it lands just after the words. Section-level timing reads as early.
- Short. 2.5 to 6.5 seconds. Most around 3. It should cut back and forth like social, not sit there.
- Behind the head or to the side, never over the speaker. No cards, no backgrounds, no dimming. That retires the corner chips, every
full-*card anddefinefor new videos. - Still show the mechanism. Reach for a diagram before a sentence. Use the in-scene diagram types (
dots,flow,bars,compare,math), which sit on the wall beside the speaker. - Split anything with two halves. Old way and new way are two clips. 10% and 20% are two clips. Three options are three clips, each landing as it's named.
- Every line lands on its word. Lists, side lines and typed words carry per-line offsets (
at,t) from the word timestamps, not an even stagger.scripts/plan.example.pyshows how: quote the phrase, name the cue words, and it does the lookup. - A good mix (about 48 clips over 11 minutes): roughly a quarter in-scene diagrams (mostly
flow), thenside,pair,type,list2,num. Savepairfor the punchlines. - Big titles go around the head, not behind it.
pairsplits each line at a word boundary into two halves that flank the head, with a gap of 1.9x the face-box width (hair and ears included), scaled so both halves stay on screen. So everypairline needs 2+ words ("ONE HOUR", not "HOUR"): a single word can't split and ends up hidden."around": falseindatabrings back the centred look. The outlined line is 150px with a 5px stroke and a light accent fill (35%) so it stays readable.
Before you start: is the cut locked?
Overlays are anchored to exact words, so any re-trim afterwards shifts every clip. If the user hands over raw footage, or the cut isn't locked, run the paper-edit skill first and place overlays on the timeline it produces.
Pipeline
1. Get the cut and probe it
Ask for an export from Resolve (e.g. to ~/Movies/), or raw footage on <YOUR_FOOTAGE_DRIVE> for a paper edit first. A video file is fine; no need to export audio separately. When Resolve is open, read the timeline through the MCP instead of asking for an export.
ffprobe -v error -show_entries format=duration -show_entries stream=codec_type,r_frame_rate -of default=nw=1 "<file>"
Note the frame rate (the renderer assumes 23.976 = 24000/1001; change FPSN/FPSD in render.mjs for other rates) and ask whether the timeline starts at 00:00:00:00 or 01:00:00:00 (Resolve's default is 01).
2. Transcribe with word-level timestamps
ffmpeg -y -i "<file>" -ac 1 -ar 16000 -c:a pcm_s16le cut.wav
whisper-cli -m ~/.cache/whisper-models/ggml-small.en.bin -f cut.wav -oj -ml 1 -sow -of cut
-ml 1 -sow gives one token per segment, which is what makes exact anchoring possible. Then find the start time of each phrase by matching token runs in cut.json (transcription[].offsets.from is ms).
3. Pick the moments
Read the whole transcript first. Aim for roughly one clip per 15 seconds of video. For each: the anchor phrase, a type, and the shortest duration that lets the animation land and breathe.
Ground the content in what was actually said, not in written notes. If notes say "30 minutes" and on camera it's "half an hour or so", the recording wins, and flag the difference to the user.
4. Write clips.json
Copy scripts/clips.example.json as the model. One object per clip:
{"id":"06-math-10","start":52.96,"dur":3.6,"type":"dots","cue":"\"if 10 out of 100 meals are home cooked\"",
"data":{"lit":10,"pct":"10%","count":"10 meals"}}
Keep it outside the skill folder (a scratchpad path is fine) and pass it with CLIPS=. start is the raw phrase time in seconds; the renderer adds the 0.12s lead itself. id order sets file order, so number them.
5. Render a clean cut (needed by the in-scene types)
The in-scene types read the head position and cut the speaker out from a copy of the cut with no overlays on it. Render the locked timeline from Resolve with every overlay clip disabled, H.264 1080p, into the scratchpad. In the API, SetTrackEnable on V2 did not take in Resolve 19; item.SetClipEnabled(False) on each clip in the track did. The raw footage has to be online, so the footage drive must be mounted; if it isn't, ask the user to plug it in and watch for the mount rather than rendering a finished master (old overlays would be baked into it).
6. Render
cd <skill>/scripts && npm i # first time only, installs puppeteer-core
swiftc -O person.swift -o person # first time only, the matte + head finder
TITLE="<video name>" CLIPS=/path/to/clips.json FOOTAGE=/path/to/clean.mp4 OUT_DIR="$HOME/Movies/<video-name>-overlays" node render.mjs
# one or a few clips only:
ONLY=06-math-10,07-math-30 OUT_DIR=... node render.mjs
Roughly 10 to 15 seconds of wall clock per second of overlay (pair is slower, it mattes every frame). Backgrounding it and checking in is the right move. It writes PLACEMENT.txt and both .fcpxml files at the end.
Always proof over the real footage before a full render. node proof.mjs <clean.mp4> <clips.json> out.jpg [email protected] [email protected] ... renders one frame per clip (prefix@seconds-into-clip) over the actual shot, matte included, tiled into one image. Look at it. Run-off-frame text, a title crossing the eyes, text over a busy background: all invisible in code and obvious in the image. For diagrams, compositing over grey is still fine.
Proof frames: extract to PNG first. Overlaying a qtrle .mov straight onto a lavfi color source with -ss gives a blank frame. Pull the frame out (ffmpeg -ss 2.8 -i clip.mov -frames:v 1 f.png), then composite the PNG.
Placing in Resolve: duplicate the locked timeline as YouTube Final, import the movs into an Overlays bin, and AppendToTimeline each with trackIndex: 2 and recordFrame from the IN timecode in PLACEMENT.txt (((h*3600+m*60+s)*24+ff), so 01:00:00:00 = 86400). Parse PLACEMENT.txt one line at a time (line.split()[0] is the id, [1] the timecode). A whole-file regex can read cue text like "60-second" as a clip id and steal the next line's timecode, so one clip silently goes unplaced. Check that the placed count equals the clip count.
7. Preview, then hand off
Build a 10-15 second preview over the real footage so the user can judge before it goes in:
ffmpeg -y -ss <t> -t 11 -i "<cut>" -itsoffset <a> -i "<clip1>.mov" -itsoffset <b> -i "<clip2>.mov" \
-filter_complex "[0:v][1:v]overlay=0:0:eof_action=pass[v1];[v1][2:v]overlay=0:0:eof_action=pass[vo];[vo]scale=1280:-2[v]" \
-map "[v]" -map 0:a -c:v libx264 -crf 20 -preset veryfast -c:a aac PREVIEW.mp4
The original audio comes along with the footage, which is all the preview needs.
Then give the import steps (File > Import > Timeline > Import AAF, EDL, XML, pick the starts-01 file, copy the clips up onto a track above the footage at 01:00:00:00).
Working inside Resolve directly (preferred when Resolve is open)
If a resolve MCP server is installed (for example the third-party barckley75/resolve-claude-mcp), it talks to Resolve Studio's scripting API. If the paper-edit skill ships an ensure_resolve.sh, run it first so Resolve is launched if it's closed. With Preferences > System > General > External scripting using: Local, prefer it over the export-and-import dance.
What it changes:
- No exported cut needed. Read the open timeline for its real timecodes, frame rate, start timecode and clip layout, then anchor overlays to that. This fixes "timecodes shift every time I re-trim."
- No manual import. Place the rendered
.movclips onto a video track directly. - Better moment picking. Seeing where the cuts already fall means overlays can land with the edit instead of fighting it.
- It can also kick off the render when the overlays are in.
Rules for using it:
- Read before you write. Confirm project name, timeline name, start timecode and frame rate first, and say them back to the user.
- Never write to the master timeline. Duplicate it first and work on the copy, so a mistake costs nothing.
- Avoid
execute_resolve_codeunless a normal tool can't do the job. It runs arbitrary Python against the user's projects. If it is the only way, show the code first. - Local sessions only. Resolve has to be running on this Mac, so this never works in a cloud or scheduled run. Fall back to the FCPXML there.
- The FCPXML path still works and is still the fallback. Keep generating both files; they cost seconds.
Audio transcription: some Resolve MCP servers ship their own mlx-whisper transcription, but the whisper.cpp step above is proven with this pipeline.
Clip types
Everything below is retired for new videos (kept so old videos can be re-cut). Use the in-scene types in the next section. Full-frame types paint a solid ink background with a soft accent glow, so they hide the speaker completely. Chips sit bottom-left.
| type | use it for | data |
|---|---|---|
full-dots | retired, use dots | lit, pct, count |
full-flow | retired, use flow | label, steps[], good |
full-compare | retired, use compare | title, left{label,value,sub}, right{...,hi}, vs |
full-bars | retired, use bars | title, bars[{label,value,value_label,hi}] |
full-math | retired, use math | parts[] (operators x = + → render in the accent color), sub |
full-pack | retired, use bars or side | name, note, size 0-1, hi |
full-stat | retired, use pair | label, value, sub, hi |
full-title | retired, use pair or type | title (<br> allowed), size, sub |
full-list | retired, use list2 | title, items[], step seconds |
chip | retired, use side or num | num, title, sub, hi |
chip-stack | retired, use list2 | items[] |
The retired types still render, for re-cutting old videos.
In-scene types (the default)
No cards, text lives in the room next to the speaker's head, and big titles go around the head.
| type | use it for | data |
|---|---|---|
pair | a punchline title: outlined line above the head, solid line at forehead height | top, bottom, outline (top/bottom) |
side | 2-3 short lines beside the head, each landing on its word | lines[], accent (index, big + accent color), at[] |
type | a quote or question typed on word by word as it's said | words[{w,t,hi,br}] |
define | retired: dims the shot. Use side or pair | word, sub |
num | outlined number draws on, words beside it ("01 plan the week") | num, lines[] |
list2 | items building on the left and/or right of the head | left[], right[], at[] (in order, left first) |
In-scene diagrams (same wall placement and matte):
| type | use it for | data |
|---|---|---|
dots | a rate out of 100, grid fills beside the speaker | lit, pct, count |
flow | a process, steps chained top to bottom, last in the accent color | label, steps[], at[] optional |
bars | two or three numbers by size | title, bars[{label,value,value_label,hi}] |
compare | two options, one each side of the head | left{label,value,sub}, right{...,hi} |
math | an equation read down the wall, operators lead each line | parts[], sub |
at and t are seconds from clip start; build them from the word timestamps so every line lands on its word.
These need FOOTAGE=, the cut with no overlays (render the timeline from Resolve with overlay clips disabled). The renderer reads the face box and a background brightness grid for each clip, so text finds the empty wall, flips side or shrinks if it would run off frame, and switches to ink on a light background and light text on a dark one. It mattes the speaker out of every frame of every in-scene clip (Apple Vision person segmentation, person.swift), so titles sit behind the head and hands pass in front of the graphics. Compile once: swiftc -O person.swift -o person. Without FOOTAGE they still render, just with a guessed head position and no matte.
Adding a type: add a function to the R registry in scene.html, keyed by name. It gets T (time), D (duration) and DATA. Use life(in_) for opacity, pr(start,dur) for progress, backdrop() for full-frame.
Output format
Output is qtrle (QuickTime Animation) with a real alpha channel: lossless, editor-friendly, and far smaller than ProRes 4444 for flat graphics. Brand colors and fonts are set in the :root block of scene.html (see "Set your brand here").
Gotchas
--force-device-scale-factor=1andomitBackground: trueon every screenshot, or the alpha comes back wrong.- qtrle needs
-pix_fmt argb. Without it the alpha is silently dropped. - Timecode at 23.976 is counted at 24 frames per second of timecode:
frames = round(seconds / 1.001 * 24). - Chrome path:
render.mjsandproof.mjslaunch/Applications/Google Chrome.app/.... Change the path on other systems. - Don't move the output folder after Resolve imports it, or every clip goes offline.
white-space:nowrapon chip text. Wrapping looks broken at this size.- Paths with spaces: build filesystem paths with
fileURLToPath(import.meta.url), notnew URL(...).pathname(that one is URL-encoded andspawncan't find the file). - The face box starts at the eyebrows, not the hairline.
pairplaces its solid line athead.y - 0.1*head.hso it runs at forehead height; anchoring at the box center put the words across the eyes. - Light vs dark backgrounds. White text vanishes on a light wall, so in-scene text samples the brightness behind itself: ink on light areas, light text on dark areas.
- The FCPXML places everything on
lane="1"inside one gap. Resolve imports it as a standalone timeline; copy the track up onto the edit.
When the cut changes
Timecodes shift with every trim. For diagrams, only the start values in clips.json need updating; the clips don't need re-rendering. In-scene clips do (side, pair, type, num, list2): their placement and matte come from the footage under them, so re-render the clean cut and re-render those clips with ONLY=. Re-run with ONLY= set to nothing and it will rebuild the FCPXML from the existing files in seconds.
Related
paper-editfor the rough cut that comes first- Your own voice/style notes for how on-screen copy should sound