Paper edit
Raw camera files in, assembled rough cut in Resolve out. On its first real test video: raw 14:17, out 10:16, against the creator's own hand edit of 10:32.
The division of labor: Claude does the broad cut, the creator does the fine trim. Don't try to deliver a finished edit, and don't be so conservative that they have to redo the obvious parts.
Setup: fill these in
| Placeholder | What it is | Example |
|---|---|---|
<FOOTAGE_DRIVE> | Where raw footage lives, one folder per video | /Volumes/MyWorkDrive |
<LUT_PATH> | Optional LUT, relative to Resolve's LUT folder | MyLuts/SLog3-to-709.cube |
WHISPER_MODEL (env) | whisper.cpp model file | default ~/.cache/whisper-models/ggml-small.en.bin |
RESOLVE_APP (env, optional) | Resolve app path if non-standard | default /Applications/DaVinci Resolve/DaVinci Resolve.app |
Requirements: ffmpeg, ffprobe, whisper-cli (whisper.cpp), Python 3, DaVinci Resolve with a Resolve MCP server that exposes execute_resolve_code.
Run it
Work dir: use the session scratchpad, e.g. <scratchpad>/paperedit. Skill dir below is .claude/skills/paper-edit.
1. Find the footage
Footage lives on <FOOTAGE_DRIVE>, one folder per video, often named loosely.
find "<FOOTAGE_DRIVE>" -maxdepth 4 -iname "*<keyword>*" -not -path "*/.*"
Confirm the folder with the user if more than one matches. Don't copy the files; Resolve links to them on the drive.
2. Transcribe every clip
bash <skill>/scripts/transcribe.sh "<FOOTAGE_DRIVE>/<folder>" "<work>"
Writes <clip>.json (word-level), <clip>.wav, and clips.ndjson with durations. About 90s for 15 minutes of footage. Skips clips already done, so it is safe to re-run.
3. Read the transcript yourself
python3 <skill>/scripts/sentences.py "<work>"
Read all of it. This is the part that cannot be scripted. A repeated-phrase detector was tried and failed badly: it wanted to cut 428 of 737 seconds, because teachers repeat phrases on purpose ("here's what works, here's what doesn't").
4. Write the cut list
<work>/cuts.json, one object per removal, times in seconds from the start of that clip:
[{"clip":"C0002","start":32.58,"end":39.76,"why":"stumble: 'what I realized was even sometimes...'"}]
Cut these:
- Asides and camera talk: "that was pretty good, not gonna lie", "this camera's about to die", "this'll probably be a seven minute video"
- False starts, keeping the last attempt. Speakers often restate a sentence twice, better the second time.
- Failed takes of the CTA. There are often 3-4 at the end; keep the final one.
- Long silences between sections, bracketed noise (
[BLANK_AUDIO],(keyboard clicking),(shuffling)) - Sentences abandoned mid-thought and rebuilt ("the reason I used to not do this was because I was scared to, and...")
- Short restarts inside a sentence, marked by a comma and a restated clause: "all I asked for in return is a simple review and this is, and adding the review part..." Cut from the abandoned clause to the restart. These are 1-2s and easy to read past.
- Whisper word timings drift about 0.5s. On a false start that restates the same words ("I love to come in offering, I love to start off"), end the cut at the start of the second "I love to", not the end of the first attempt, or a stray word survives.
- Start a cut at the kept word's
tooffset, never itsfrom. In one test a cut started at 40.30 while the kept word ran 40.19-40.92, so the hook's last word was chopped and had to be dragged back by hand. Print the offsets withfrom|tofor the word before every cut. - After a long
[BLANK_AUDIO]run, timings are unreliable. Whisper squeezed the words of the first hook take together after 30s of silence, and the cut left a word fragment as the very first second of the video. When a take starts after dead air, check the audio level (ffmpegastats) to find where speech really starts instead of trusting the word offsets.
Keep:
- Deliberate repetition used for teaching rhythm
- Tangents that still teach something, unless the user says otherwise
- Anything you are unsure about. Flag it in your summary instead of cutting it.
5. Build the segments
python3 <skill>/scripts/segments.py "<work>"
Produces segments-v1.json (judgment cuts only), segments-v2.json (also trims pauses over 0.50s down to 0.22s), and cutlist.txt. Tunables via env: MIN_GAP, KEEP_PAUSE, NOISE_DB.
It also drops no-speech blips: v2 segments under 1s with no audio at speaking level. They are mouse clicks and glances at a laptop, left as islands between trimmed pauses. Tunables: BLIP_MAX, SPEECH_DB.
Read every join: line in cutlist.txt before showing the user. It prints how the sentence reads across each cut. If the same phrase appears on both sides of the |, the cut is too short.
Both passes matter. With only the judgment cuts, the feedback was "there are big pauses throughout, especially at the start." The pause trim is what took 11:32 down to 10:16. Don't drop KEEP_PAUSE below ~0.20s; it starts to sound rushed.
6. Show the cut list before building
Paste the summary: raw length, v1, v2, cut count, and the notable cuts with reasons. The user vetoes what they want.
7. Build it in Resolve
Requires Preferences > System > General > External scripting using: Local (set once, it sticks). Don't ask the user to open Resolve. Run this first; it launches Resolve if it's closed and waits until the API answers:
bash <skill>/scripts/ensure_resolve.sh
If it times out with Resolve running, scripting is off; that is the one thing to ask the user to fix.
Write <work>/params.json:
{"project":"<Video Name> - YYYY-MM-DD","work":"<work>",
"files":["<FOOTAGE_DRIVE>/<folder>/C0001.MP4","..."],
"fps":23.976,"width":1920,"height":1080}
Then one execute_resolve_code call:
PARAMS = "<work>/params.json"
exec(open("<skill>/scripts/resolve_build.py").read())
Several videos: build them one at a time, never as parallel tool calls. Resolve has one current project, so two builds at once collide. Once two builds ran together and one v2 came out empty, even though the script reported every item added. After each build, check each timeline's item count yourself.
Name the project <Video> - YYYY-MM-DD so same-named projects never get confused. It creates a new project, imports the clips, and builds Paper Edit v1 and Paper Edit v2 (tight). Nothing existing is touched. This is a sanctioned use of execute_resolve_code; append_to_timeline cannot take in/out points, so there is no alternative tool. Show the user the code the first time in a session.
Fallback if Resolve can't be reached (cloud/scheduled run, or ensure_resolve.sh failed): write an FCPXML of asset-clip elements, one per segment, with start and duration in frames*1001/24000s, and let the user import it. Same result, one extra click.
8. Hand off
Tell the user both timelines exist, which one to scrub first (v2), and that v1 is the safety net if v2 feels clipped. Then ask whether the cut is locked, because overlays come after (see video-overlays, if you have it).
9. Punch-ins (optional third timeline)
Static framing variation makes a talking head watchable. Resolve's API has no keyframe access, so slow pushes are impossible to script; static per-clip framing ("every now and then one of the clips is a little more zoomed in") is what you can do.
Tie every change to the transcript, never to a timer. A timer-based version felt random. Punch in on the heavy-hitting lines or where a new topic starts, so it makes sense.
Write <work>/framing.json:
[{"clip":"C0002","at":361.40,"framing":"med","why":"TIP 3"},
{"clip":"C0002","at":371.80,"framing":"close","why":"'this is the part that matters'"}]
at is the sentence start from sentences.py. Then python3 <skill>/scripts/punchins.py "<work>" and set "punchins": true in params.json. It splits a segment wherever a mark falls inside it, at the quietest point near the sentence start, because talking heads often run long without a pause and otherwise the change slides late.
If the user hand-trimmed v2 in Resolve, their timeline is the cut. Read its items (GetLeftOffset/GetDuration, frames x 1.001/24 = seconds) into segments-locked.json, run with SEGMENTS=segments-locked.json, and add only v3 to the existing project with "timelines": ["v3"]. Aim for one change every ~20-40s: wide at each new topic, medium for teaching, close on the heavy line. The builder drops a blue marker at each change naming the reason, so the logic is visible.
Framings: wide 1.00, med 1.11 (tilt -12), close 1.21 (tilt -24).
If the user graded v2 while trimming it, v3 is built from scratch with only the LUT, so copy the v2 grade onto v3 right after building it:
src = v2.GetItemListInTrack("video", 1)[5] # any graded v2 item
src.CopyGrades(v3.GetItemListInTrack("video", 1)) # works across timelines; zoom/tilt survive
Verify with GetNodeGraph().GetToolsInNode(1) on every v3 item. Don't use the gallery route (GrabStill + ExportStills to .drx): the export returns False through the API. If GrabStill was used, delete the still so the gallery stays clean.
Reading a locked v2: GetLeftOffset() and GetDuration() give source frames. The first segment tells you if the hook was trimmed. Diff the locked segments against segments-v2.json and note what changed, to calibrate the next edit.
10. Colour
Set "lut": "<LUT_PATH>" in params.json to apply it to every clip on node 1. If the camera shoots a log profile (e.g. S-Log3), ungraded footage looks flat and grey, so a LUT is worth setting. The API cannot add a second node; for a multi-node grade, grade one clip by hand and push it out with CopyGrades.
11. Photos over the talk (optional)
When the speaker talks about something they have stills of, put them on V2 of v3, on a white backdrop.
- Cards, not a slideshow. Render 1920x1080 PNGs on pure white: three portraits side by side (scale to h840, 50px gaps), two portraits (h880, 60px gap), or one landscape (h900). Group photos by what is being said.
- Time each card to a sentence start from the locked transcript, and switch cards on the next sentence. About 2-5s per card. Stay on the photos only while they are being described.
- Import them as exact-length ProRes .mov, not stills. Stills come into Resolve at a fixed 2s and overlap each other. Numbered names (
card1.png,card2.png) import as a single image sequence, so use word names. Render withffmpeg -loop 1 -framerate 24000/1001 -i x.png -frames:v N -c:v prores_ks -profile:v 3. - For a .mov,
AppendToTimelineneedsendFrame= the frame count (not count-1), or every card lands a frame short and the speaker flashes on screen between cards. Pass"mediaType": 1so no audio comes along. - Name the track " photos", put the cards in their own media pool folder, and never put the LUT on them (they're already graded).
- Show the user a contact sheet of the cards in chat before placing them.
Facts worth keeping
- Calibration from real edits: after a v2, the creator typically removes 10-20s by hand from a ~10-minute cut. Most of it is frame-level tightening at segment edges, plus restated lines, where a point is said and then said again better within a few seconds. When a sentence restates the one before it, cut the first version. This is different from deliberate teaching repetition, which repeats a phrase across sections, not back to back. When a point is made twice far apart, flag the weaker take as a possible cut.
- The usual worry is over-cutting so things have to be added back, so tighten only on evidence, never on a hunch.
- Example source format: 4K 23.976 MP4 with 48kHz audio; delivery timeline 1920x1080 23.976, starting at 01:00:00:00. Adjust
fps/width/heightinparams.jsonfor your footage. - Timecode at 23.976 counts 24 frames per timecode second:
frames = round(seconds / 1.001 * 24). AppendToTimelinetreatsendFrameas exclusive: pass the frame after the last one you want. An "inclusive, subtract 1" rule drops a frame from every segment (about 4s over a 10-minute timeline, and any 1-frame segment disappears).- Some speakers use almost no filler words; check before building around filler removal.
- Roughly 20% of raw talking-head footage is removable, and about three quarters of that is judgment (takes and restarts), one quarter is dead air.
- Whisper model:
ggml-small.en.binis accurate enough; it does mishear a few words, which does not matter for cut points but does matter if you quote the speaker.
Related
video-overlays: animated graphics, run after the cut is locked (if installed)video-reels: vertical cuts from the same footage (if installed)video-thumbnails: thumbnails and titles (if installed)lapel-sync: run first when audio was recorded on a separate lapel/recorder