Lapel sync
Camera clips plus a separate recorder track in, one lapel slice per clip and a synced Resolve timeline out. First tested on a live talk: 5 camera clips against a 59-minute recorder track. Every clip landed within 20 ms, well under one frame.
Why it matters: at a live event the camera mic is across a room. Its audio is unusable, and the lapel is the only clean sound. After this skill, <work>/<clip>.wav and <clip>.json are the lapel, so paper-edit and video-reels read clean audio and clean transcripts without knowing anything changed.
Setup: fill these in
| Placeholder | What it is | Example |
|---|---|---|
<FOOTAGE_DRIVE> | Where raw footage lives, one folder per video, recorder files in an audio/ subfolder | /Volumes/MyWorkDrive |
<SYNC_OUT> | Where full-quality lapel slices go | ~/Movies/<video>-sync |
<LUT_PATH> | Optional LUT, relative to Resolve's LUT folder | MyLuts/SLog3-to-709.cube |
WHISPER_MODEL (env) | whisper.cpp model file | default ~/.cache/whisper-models/ggml-small.en.bin |
Requirements: ffmpeg, ffprobe, whisper-cli (whisper.cpp), Python 3 (numpy and scipy are installed automatically into ~/.cache/lapel-sync-venv), DaVinci Resolve with a Resolve MCP server exposing execute_resolve_code, and the paper-edit skill (for ensure_resolve.sh).
Run it
Work dir: the session scratchpad, e.g. <scratchpad>/<video>. Skill dir below is .claude/skills/lapel-sync. Full-quality slices go to <SYNC_OUT>.
zsh gotcha: in inline shell, $ss:end_sample parses :e as a modifier and silently eats the text. Always write ${ss}.
1. Find the footage and the audio
One folder per video on <FOOTAGE_DRIVE>. The recorder files sit in an audio/ subfolder:
find "<FOOTAGE_DRIVE>" -maxdepth 4 -iname "*<keyword>*" -not -path "*/.*"
ls "<folder>" "<folder>/audio"
A Tascam DR-10L, for example, writes 000_YYMMDD.wav, 001_..., split every 15 minutes, plus _D twins. The _D files are the dual-record safety track, about 20 dB quieter. Use the main files; keep _D for patching clipped spots.
2. Stitch the recorder files
bash <skill>/scripts/stitch.sh "<folder>/audio" "<SYNC_OUT>/lav-master.wav"
It reads each file's BWF time_reference (samples since midnight) and checks that every file starts exactly where the last one ended. When they are sample-contiguous it is a lossless stream copy. If there is a gap it inserts that much silence, so the master stays on wall-clock time, and warns with the size.
3. Transcribe both sides
bash <skill>/scripts/transcribe_sources.sh "<folder>" "<SYNC_OUT>/lav-master.wav" "<work>"
Everything lands in <work>/sync/, kept out of <work> itself so paper-edit's sentences.py never reads it. Camera audio is boosted hard first (highpass=f=120,volume=40dB,alimiter, 16k mono), because it is often around -65 dB. Skips anything done, so re-run freely.
The lapel takes about 8 minutes per hour of audio. Run it in the background. Whisper streams word timings into <work>/sync/lav.log, and align.py reads that log when lav.json isn't written yet, so alignment can start before the lapel finishes. Clips past the transcribed point just report "only N word matches"; re-run once it's done.
4. Align
python3 <skill>/scripts/align.py "<work>"
Writes <work>/sync/sync.json. First run creates ~/.cache/lapel-sync-venv with numpy and scipy (system python3 often has neither); after that it just re-runs itself inside it.
How it works, so you can trust or question the output:
- Words first. Unique 4-word runs that appear in both transcripts. Offset = lapel time minus camera time. Hundreds of matches per clip; densest cluster, keep within 0.5s, median. Good to about 0.1s.
- Waveform refine. Each 20s chunk of camera audio (6s chunks on clips under 40s), band-passed 300-3000 Hz at 8 kHz, is cross-correlated against the lapel within +-0.6s of the word offset. Peaks over 6x the median count. Chunks within 20 ms of each other agree; on clips over 5 minutes a line fit gives offset plus drift.
Read every line it prints: offset, chunks agreeing, spread, drift. Any clip marked CHECK BY EAR (fewer than 3 agreeing chunks) may be up to 0.1s off; scrub it in Resolve before trusting it. Drift over a frame across a clip gets flagged too.
5. Slice the lapel per clip
python3 <skill>/scripts/slice.py "<work>" "<SYNC_OUT>/lav-master.wav" "<SYNC_OUT>"
For each clip it writes:
<SYNC_OUT>/<clip>_lav.wav: 24-bit, source rate, sample-exact to the clip's length, so lapel time equals clip time<work>/<clip>.wav(16k mono) and<work>/<clip>.json(whisper of the slice): what paper-edit and video-reels read<work>/clips.ndjsonwith an"audio"field pointing at the slice
If the recorder started late or stopped early for a clip, that part is filled with silence and it warns. Whisper on the slices takes a few minutes; don't also run paper-edit's transcribe.sh afterwards (it would skip anyway, since the jsons exist).
6. Check for clipping
slice.py counts samples at or above 0.99 full scale per slice and prints the longest run. Hundreds of short runs (tens of samples) are usually the recorder's limiter catching peaks, not distortion. Only when it lists spots to listen to (runs of 1 ms or more), or the user says it sounds crunchy, patch those moments from the _D twin (same timing, bring it up about 20 dB).
7. Build it in Resolve
bash .claude/skills/paper-edit/scripts/ensure_resolve.sh
Write <work>/sync/params.json:
{"project":"<Video> - YYYY-MM-DD","work":"<work>","name":"Synced - Full Talk",
"lut":"<LUT_PATH>"}
lut is optional; drop it if you don't use one. Then one execute_resolve_code call:
PARAMS = "<work>/sync/params.json"
exec(open("<skill>/scripts/resolve_sync_timeline.py").read())
It creates the project (or opens it if the name exists), imports the camera clips and the lapel slices, and builds Synced - Full Talk: each clip on V1 as picture only (mediaType: 1), its slice on A1 as sound only (mediaType: 2), same in and out frames. recordFrame = timeline start + round((offset - earliest offset) * 24000/1001), so the real gaps between clips stay in. LUT on node 1. Camera scratch audio is left off. Timeline is 1080x1920 when the clips are vertical, 1920x1080 otherwise (override with width/height). If the timeline name is taken it adds v2 rather than touching the existing one.
Have the user scrub a few spots with clear consonants (a "p" or "t") on each clip. That is the real check.
8. Hand off
Tell the user the timeline exists, the per-clip numbers from step 4 in one line each, anything flagged, and the clipping verdict. Then move on: paper-edit (skip its transcribe step; <work> is already set up) or video-reels.
Facts worth keeping
- Waveform correlation alone failed. Whole-clip and 20s-chunk correlation over a wide window gave inconsistent garbage when the camera mic was about -65 dB mean across a reverberant room. Words first, then a narrow waveform search, is what works.
- The camera clock is often wrong. In testing, the camera was about 149s behind the recorder. Some cameras (e.g. Sony) write
creation_timein UTC (ends inZ), while a recorder's BWFtime_referenceis local. Treat creation_time as a rough hint at best. The word match doesn't need it. - Drift is negligible on the tested gear: about 0.5 ms/min, 11 ms over a 21-minute clip. No time-stretch needed. Short clips report 0 because a fit over a few seconds is just noise.
- Re-running the scripts reproduced hand-built slices to within 6 ms.
- Vertical footage: e.g. a Sony shooting vertical records 4K 23.976 with a rotation=90 display matrix. Resolve reads it as 2160x3840 with no rotation fix needed.
- Example recorder, Tascam DR-10L: 48k 24-bit mono BWF, new file every 15 minutes,
_Dsafety twin about 20 dB down. - paper-edit's
resolve_build.pystill takes sound from the camera file. On a lapel shoot its timelines carry the camera's scratch audio until it learns to read the"audio"field. Until then, tell the user, or relink A1 to the_lav.wavslices; the cut points are the same because slice time equals clip time. - Whisper runs with
-oj -ml 1 -sowfor word timings.
Related
paper-edit: the rough cut, run after thisvideo-reels: readsclips.ndjson; the"audio"field tells it to take sound from the slice (if installed)