caption-burn
Captions timed to what is actually said in the finished video, not to the script's estimate.
Run
python transcribe.py --media reel.mp4 --out reel.words.json # paid, cents
python captions.py --video reel.mp4 --beats cutlist.aligned.json \
--words reel.words.json --out final.mp4 [--style plate|outline] [--highlight Brand]
--beats supplies the lines (vo), their timing and, for split layouts, seam, size
and each beat's state. Any file with beats: [{start, end, vo}] works.
Placement and style
-
--anchor seam(the default when the beats have a seam): onsplitbeats the plate is pinned to the seam, positioned by the plate, not the text: 25% of the plate above the line, 75% below. Full-frame beats use--full-y. -
--anchor fixed --y 0.62: every caption's plate centred at that fraction of the height. -
plate(default): white bold on a dark grey rounded plate, 1–2 words, cap ~0.019 H. -
outline: white bold with a dark outline, no plate, 1–3 words, cap ~0.034 H. -
--highlight WORDcolours that word yellow (the CTA keyword). Repeatable. -
serif-word: ONE word at a time, heavy serif (Georgia Bold), white with a black outline, on a fixed baseline at 0.77 H (the screen-insert look). The highlight word is quoted. -
--card "LINE ONE|LINE TWO" --card-until 4.7: a white rounded hook card with two lines of heavy red capitals near the top, for the opening seconds. ~14 characters a line.
Captions for a format with no voice (plates.py)
python plates.py --video walk.mp4 --beats cutlist.json --out captioned.mp4 [--logo logo.png]
Each beat's caption (a string or list of lines) shows for the whole beat on ONE black
block (square rectangles unioned, then rounded as a single silhouette: rounding each line
leaves seams), lines left-aligned, the block centred on its widest line. It goes in the
emptiest band of that beat's frame unless the beat pins cap_y. logo: true on a beat
hangs the logo tile under the block. Write lines a person would type: the same short
"fragment. fragment." shape three times reads as AI-written. No emoji twice.
Rules
- Transcribe the FINISHED audio. Joining takes and aligning lines shifts timing; only the final video's audio gives correct cues.
- The last caption holds to the last frame. It is the call to action.
- A word Whisper writes differently ("200" for "two hundred") is interpolated between its neighbours rather than dropped or stretched over the whole line.
- Without
--wordstiming is estimated from syllables. Use that to judge placement, never to ship. - Fonts: a bold sans is found on macOS, Linux or Windows; if none is present, Roboto
Bold is fetched once into
~/.cache/gooseworks/fonts.--fontorGW_CAPTION_FONToverrides.