Communitygithub.com

beachyphotoandfilm-arch/video-skills

Turn raw talking-head footage into an assembled rough cut inside DaVinci Resolve. Pulls the clips off your footage drive, transcribes them locally, removes failed takes, restarts, asides and dead air, then creates a new Resolve project with two paper-edit timelines ready for your fine trim. Trigger on "paper edit", "cut my raw footage", "make a rough cut", "clean up this video", or when the user points at a folder of raw camera files for a YouTube video.

video-skills란 무엇인가요?

video-skills is a Claude Code agent skill that turn raw talking-head footage into an assembled rough cut inside DaVinci Resolve. Pulls the clips off your footage drive, transcribes them locally, removes failed takes, restarts, asides and dead air, then creates a new Resolve project with two paper-edit timelines ready for your fine trim. Trigger on "paper edit", "cut my raw footage", "make a rough cut", "clean up this video", or when the user points at a folder of raw camera files for a YouTube video.

지원 대상✓Claude Code~Codex CLI~Cursor
npx skills add https://github.com/beachyphotoandfilm-arch/video-skills/tree/HEAD/skills/paper-edit

즐겨 사용하는 AI에게 물어보기

이 에이전트 스킬이 미리 로드된 새 채팅을 엽니다.

문서

Paper edit

Raw camera files in, assembled rough cut in Resolve out. On its first real test video: raw 14:17, out 10:16, against the creator's own hand edit of 10:32.

The division of labor: Claude does the broad cut, the creator does the fine trim. Don't try to deliver a finished edit, and don't be so conservative that they have to redo the obvious parts.

Setup: fill these in

PlaceholderWhat it isExample
<FOOTAGE_DRIVE>Where raw footage lives, one folder per video/Volumes/MyWorkDrive
<LUT_PATH>Optional LUT, relative to Resolve's LUT folderMyLuts/SLog3-to-709.cube
WHISPER_MODEL (env)whisper.cpp model filedefault ~/.cache/whisper-models/ggml-small.en.bin
RESOLVE_APP (env, optional)Resolve app path if non-standarddefault /Applications/DaVinci Resolve/DaVinci Resolve.app

Requirements: ffmpeg, ffprobe, whisper-cli (whisper.cpp), Python 3, DaVinci Resolve with a Resolve MCP server that exposes execute_resolve_code.

Run it

Work dir: use the session scratchpad, e.g. <scratchpad>/paperedit. Skill dir below is .claude/skills/paper-edit.

1. Find the footage

Footage lives on <FOOTAGE_DRIVE>, one folder per video, often named loosely.

find "<FOOTAGE_DRIVE>" -maxdepth 4 -iname "*<keyword>*" -not -path "*/.*"

Confirm the folder with the user if more than one matches. Don't copy the files; Resolve links to them on the drive.

2. Transcribe every clip

bash <skill>/scripts/transcribe.sh "<FOOTAGE_DRIVE>/<folder>" "<work>"

Writes <clip>.json (word-level), <clip>.wav, and clips.ndjson with durations. About 90s for 15 minutes of footage. Skips clips already done, so it is safe to re-run.

3. Read the transcript yourself

python3 <skill>/scripts/sentences.py "<work>"

Read all of it. This is the part that cannot be scripted. A repeated-phrase detector was tried and failed badly: it wanted to cut 428 of 737 seconds, because teachers repeat phrases on purpose ("here's what works, here's what doesn't").

4. Write the cut list

<work>/cuts.json, one object per removal, times in seconds from the start of that clip:

[{"clip":"C0002","start":32.58,"end":39.76,"why":"stumble: 'what I realized was even sometimes...'"}]

Cut these:

  • Asides and camera talk: "that was pretty good, not gonna lie", "this camera's about to die", "this'll probably be a seven minute video"
  • False starts, keeping the last attempt. Speakers often restate a sentence twice, better the second time.
  • Failed takes of the CTA. There are often 3-4 at the end; keep the final one.
  • Long silences between sections, bracketed noise ([BLANK_AUDIO], (keyboard clicking), (shuffling))
  • Sentences abandoned mid-thought and rebuilt ("the reason I used to not do this was because I was scared to, and...")
  • Short restarts inside a sentence, marked by a comma and a restated clause: "all I asked for in return is a simple review and this is, and adding the review part..." Cut from the abandoned clause to the restart. These are 1-2s and easy to read past.
  • Whisper word timings drift about 0.5s. On a false start that restates the same words ("I love to come in offering, I love to start off"), end the cut at the start of the second "I love to", not the end of the first attempt, or a stray word survives.
  • Start a cut at the kept word's to offset, never its from. In one test a cut started at 40.30 while the kept word ran 40.19-40.92, so the hook's last word was chopped and had to be dragged back by hand. Print the offsets with from|to for the word before every cut.
  • After a long [BLANK_AUDIO] run, timings are unreliable. Whisper squeezed the words of the first hook take together after 30s of silence, and the cut left a word fragment as the very first second of the video. When a take starts after dead air, check the audio level (ffmpeg astats) to find where speech really starts instead of trusting the word offsets.

Keep:

  • Deliberate repetition used for teaching rhythm
  • Tangents that still teach something, unless the user says otherwise
  • Anything you are unsure about. Flag it in your summary instead of cutting it.

5. Build the segments

python3 <skill>/scripts/segments.py "<work>"

Produces segments-v1.json (judgment cuts only), segments-v2.json (also trims pauses over 0.50s down to 0.22s), and cutlist.txt. Tunables via env: MIN_GAP, KEEP_PAUSE, NOISE_DB.

It also drops no-speech blips: v2 segments under 1s with no audio at speaking level. They are mouse clicks and glances at a laptop, left as islands between trimmed pauses. Tunables: BLIP_MAX, SPEECH_DB.

Read every join: line in cutlist.txt before showing the user. It prints how the sentence reads across each cut. If the same phrase appears on both sides of the |, the cut is too short.

Both passes matter. With only the judgment cuts, the feedback was "there are big pauses throughout, especially at the start." The pause trim is what took 11:32 down to 10:16. Don't drop KEEP_PAUSE below ~0.20s; it starts to sound rushed.

6. Show the cut list before building

Paste the summary: raw length, v1, v2, cut count, and the notable cuts with reasons. The user vetoes what they want.

7. Build it in Resolve

Requires Preferences > System > General > External scripting using: Local (set once, it sticks). Don't ask the user to open Resolve. Run this first; it launches Resolve if it's closed and waits until the API answers:

bash <skill>/scripts/ensure_resolve.sh

If it times out with Resolve running, scripting is off; that is the one thing to ask the user to fix.

Write <work>/params.json:

{"project":"<Video Name> - YYYY-MM-DD","work":"<work>",
 "files":["<FOOTAGE_DRIVE>/<folder>/C0001.MP4","..."],
 "fps":23.976,"width":1920,"height":1080}

Then one execute_resolve_code call:

PARAMS = "<work>/params.json"
exec(open("<skill>/scripts/resolve_build.py").read())

Several videos: build them one at a time, never as parallel tool calls. Resolve has one current project, so two builds at once collide. Once two builds ran together and one v2 came out empty, even though the script reported every item added. After each build, check each timeline's item count yourself.

Name the project <Video> - YYYY-MM-DD so same-named projects never get confused. It creates a new project, imports the clips, and builds Paper Edit v1 and Paper Edit v2 (tight). Nothing existing is touched. This is a sanctioned use of execute_resolve_code; append_to_timeline cannot take in/out points, so there is no alternative tool. Show the user the code the first time in a session.

Fallback if Resolve can't be reached (cloud/scheduled run, or ensure_resolve.sh failed): write an FCPXML of asset-clip elements, one per segment, with start and duration in frames*1001/24000s, and let the user import it. Same result, one extra click.

8. Hand off

Tell the user both timelines exist, which one to scrub first (v2), and that v1 is the safety net if v2 feels clipped. Then ask whether the cut is locked, because overlays come after (see video-overlays, if you have it).

9. Punch-ins (optional third timeline)

Static framing variation makes a talking head watchable. Resolve's API has no keyframe access, so slow pushes are impossible to script; static per-clip framing ("every now and then one of the clips is a little more zoomed in") is what you can do.

Tie every change to the transcript, never to a timer. A timer-based version felt random. Punch in on the heavy-hitting lines or where a new topic starts, so it makes sense.

Write <work>/framing.json:

[{"clip":"C0002","at":361.40,"framing":"med","why":"TIP 3"},
 {"clip":"C0002","at":371.80,"framing":"close","why":"'this is the part that matters'"}]

at is the sentence start from sentences.py. Then python3 <skill>/scripts/punchins.py "<work>" and set "punchins": true in params.json. It splits a segment wherever a mark falls inside it, at the quietest point near the sentence start, because talking heads often run long without a pause and otherwise the change slides late.

If the user hand-trimmed v2 in Resolve, their timeline is the cut. Read its items (GetLeftOffset/GetDuration, frames x 1.001/24 = seconds) into segments-locked.json, run with SEGMENTS=segments-locked.json, and add only v3 to the existing project with "timelines": ["v3"]. Aim for one change every ~20-40s: wide at each new topic, medium for teaching, close on the heavy line. The builder drops a blue marker at each change naming the reason, so the logic is visible.

Framings: wide 1.00, med 1.11 (tilt -12), close 1.21 (tilt -24).

If the user graded v2 while trimming it, v3 is built from scratch with only the LUT, so copy the v2 grade onto v3 right after building it:

src = v2.GetItemListInTrack("video", 1)[5]          # any graded v2 item
src.CopyGrades(v3.GetItemListInTrack("video", 1))    # works across timelines; zoom/tilt survive

Verify with GetNodeGraph().GetToolsInNode(1) on every v3 item. Don't use the gallery route (GrabStill + ExportStills to .drx): the export returns False through the API. If GrabStill was used, delete the still so the gallery stays clean.

Reading a locked v2: GetLeftOffset() and GetDuration() give source frames. The first segment tells you if the hook was trimmed. Diff the locked segments against segments-v2.json and note what changed, to calibrate the next edit.

10. Colour

Set "lut": "<LUT_PATH>" in params.json to apply it to every clip on node 1. If the camera shoots a log profile (e.g. S-Log3), ungraded footage looks flat and grey, so a LUT is worth setting. The API cannot add a second node; for a multi-node grade, grade one clip by hand and push it out with CopyGrades.

11. Photos over the talk (optional)

When the speaker talks about something they have stills of, put them on V2 of v3, on a white backdrop.

  • Cards, not a slideshow. Render 1920x1080 PNGs on pure white: three portraits side by side (scale to h840, 50px gaps), two portraits (h880, 60px gap), or one landscape (h900). Group photos by what is being said.
  • Time each card to a sentence start from the locked transcript, and switch cards on the next sentence. About 2-5s per card. Stay on the photos only while they are being described.
  • Import them as exact-length ProRes .mov, not stills. Stills come into Resolve at a fixed 2s and overlap each other. Numbered names (card1.png, card2.png) import as a single image sequence, so use word names. Render with ffmpeg -loop 1 -framerate 24000/1001 -i x.png -frames:v N -c:v prores_ks -profile:v 3.
  • For a .mov, AppendToTimeline needs endFrame = the frame count (not count-1), or every card lands a frame short and the speaker flashes on screen between cards. Pass "mediaType": 1 so no audio comes along.
  • Name the track " photos", put the cards in their own media pool folder, and never put the LUT on them (they're already graded).
  • Show the user a contact sheet of the cards in chat before placing them.

Facts worth keeping

  • Calibration from real edits: after a v2, the creator typically removes 10-20s by hand from a ~10-minute cut. Most of it is frame-level tightening at segment edges, plus restated lines, where a point is said and then said again better within a few seconds. When a sentence restates the one before it, cut the first version. This is different from deliberate teaching repetition, which repeats a phrase across sections, not back to back. When a point is made twice far apart, flag the weaker take as a possible cut.
  • The usual worry is over-cutting so things have to be added back, so tighten only on evidence, never on a hunch.
  • Example source format: 4K 23.976 MP4 with 48kHz audio; delivery timeline 1920x1080 23.976, starting at 01:00:00:00. Adjust fps/width/height in params.json for your footage.
  • Timecode at 23.976 counts 24 frames per timecode second: frames = round(seconds / 1.001 * 24).
  • AppendToTimeline treats endFrame as exclusive: pass the frame after the last one you want. An "inclusive, subtract 1" rule drops a frame from every segment (about 4s over a 10-minute timeline, and any 1-frame segment disappears).
  • Some speakers use almost no filler words; check before building around filler removal.
  • Roughly 20% of raw talking-head footage is removable, and about three quarters of that is judgment (takes and restarts), one quarter is dead air.
  • Whisper model: ggml-small.en.bin is accurate enough; it does mishear a few words, which does not matter for cut points but does matter if you quote the speaker.

Related

  • video-overlays: animated graphics, run after the cut is locked (if installed)
  • video-reels: vertical cuts from the same footage (if installed)
  • video-thumbnails: thumbnails and titles (if installed)
  • lapel-sync: run first when audio was recorded on a separate lapel/recorder

Individual skills in this repo

This repo contains 7 individual skills — each has its own dedicated page.

beachyphotoandfilm-arch/video-skills

Sync camera footage to a separate lapel or field-recorder track (e.g. Tascam DR-10L and similar) and lay it up in DaVinci Resolve. Stitches the recorder's split files, finds each clip's exact place in the audio by matching words then waveforms, cuts one lapel slice per clip so every later skill works unchanged, checks for clipping, and builds a synced timeline in a new Resolve project. Trigger on "sync the lapel", "sync my audio", "line up the recorder", "I used a lav", or any shoot with a separate audio folder (live talks, events, workshops). Run it before paper-edit or video-reels.

beachyphotoandfilm-arch/video-skills

Design a unique Instagram cover for each reel in your house style. Reads what the reel is about, writes a short hook headline, picks one of your own photos from a tagged photo library, lays it out like your existing covers, and proofs it in grid view at phone size. Renders locally, never with AI image generation. Can be called by a scheduling skill before reels are scheduled. Trigger on "make covers", "reel covers", "cover photo for this reel", "design the IG covers".

beachyphotoandfilm-arch/video-skills

Render finished reels out of DaVinci Resolve, write captions in the creator's voice, attach a custom Instagram cover from reel-covers, and schedule them to Instagram, TikTok and YouTube Shorts through Metricool. Handles the Drive upload hop, collision checks against the existing calendar, and cleanup. Trigger on "schedule the reels", "post these", "put these on the calendar", or after reels are approved.

beachyphotoandfilm-arch/video-skills

Build animated brand overlays (motion graphics) for a talking-head video, timed to the speaker's exact words, plus a DaVinci Resolve import file. House style is in-scene type (text beside and behind the speaker's head) plus diagrams. Transcribes the rough cut locally, picks the moments, renders transparent .mov clips, and writes an FCPXML that drops every clip onto the timeline already in position. Trigger on "make animations for this video", "add overlays/motion graphics", "animate this", "b-roll graphics for my YouTube video", or when the user shares a rough cut and asks for graphics.

beachyphotoandfilm-arch/video-skills

Cut vertical Instagram reels out of a long-form talking-head video, hook first. Picks the strongest standalone moments from the transcript, opens each reel on its punchiest line, builds vertical timelines in DaVinci Resolve framed on the speaker's face with the LUT applied, and finishes them: a hook (a held text card, or a word-by-word "build" hook with punch-in) and burned-look captions on V3; this-or-that reels get product graphics and a comment-keyword CTA card. You trim and render, nothing to import. Trigger on "make reels", "clip this for Instagram", "cut some verticals", or after a YouTube video is cut.

beachyphotoandfilm-arch/video-skills

Design YouTube thumbnails and titles for a talking-head video. Pulls graded full-res frames out of the raw footage, renders brand-styled variants in two proven layouts, and proofs them at phone size where the click actually gets decided. Trigger on "make thumbnails", "thumbnail for this video", "title and thumbnail", or after a video is cut.

beachyphotoandfilm-arch/video-skills

Run the whole YouTube video pipeline end to end, one stage at a time, stopping for the creator's approval at every gate. Paper edit, punch-ins, reels, reel covers, scheduling the reels, overlays, thumbnail and title, then scheduling the long-form video, all through Metricool. Keeps state per video so you can stop anywhere and pick up later. Trigger on "/youtube", "let's do the whole video", "run the video pipeline", "continue the <name> video", or when handed raw footage for a YouTube video.

관련 스킬