Communitygithub.com

beachyphotoandfilm-arch/video-skills

Design YouTube thumbnails and titles for a talking-head video. Pulls graded full-res frames out of the raw footage, renders brand-styled variants in two proven layouts, and proofs them at phone size where the click actually gets decided. Trigger on "make thumbnails", "thumbnail for this video", "title and thumbnail", or after a video is cut.

video-skills 是什么?

video-skills is a Claude Code agent skill that design YouTube thumbnails and titles for a talking-head video. Pulls graded full-res frames out of the raw footage, renders brand-styled variants in two proven layouts, and proofs them at phone size where the click actually gets decided. Trigger on "make thumbnails", "thumbnail for this video", "title and thumbnail", or after a video is cut.

兼容平台~Claude Code~Codex CLI~Cursor
npx skills add https://github.com/beachyphotoandfilm-arch/video-skills/tree/HEAD/skills/video-thumbnails

在你喜欢的 AI 中提问

打开一个已预加载此 Agent Skill 的新对话。

文档

Thumbnails and titles

Set your brand here

Edit the :root block at the top of both scripts/layout-fullbleed.html and scripts/layout-panel.html:

settingCSS variabledefault
accent color--brand-accent (+ --brand-accent-rgb as r,g,b)#4a90e2
light text color--brand-light#f2f2f2
dark ink color (stroke, panel background)--brand-ink#0d131a
display font (big text)--font-display'Anton','Bebas Neue',Impact,sans-serif
body font (small text)--font-body'Inter',sans-serif

The display font must be installed locally, since rendering runs in headless Chrome from file://. Anton and Bebas Neue are free on Google Fonts.

Also set your grading LUT in scripts/frames.sh (LUT_DEFAULT), or pass LUT=/path/to.cube per run.

Style defaults (change to taste)

  1. Never put anything over the face. No scrim gradient across the subject. Either a hard-edged panel with the face fully clear, or full-bleed with text in empty space.
  2. Pull frames from the RAW 4K, not an export. 1080p export frames look soft.
  3. Grade the frame first. Log footage (S-Log3 etc.) is flat and grey ungraded. Bake the LUT in with lut3d (see frames.sh). Skip the LUT if the camera shoots a finished profile.

Rules that came from research

  • Three to five words, and fewer is better. The thumbnail is not where the video gets explained.
  • A face wins. Most breakout videos use one. Strong expression beats a neutral one.
  • One dominant subject, two or three colours, high contrast.
  • Judge it at ~210px. Most viewing decisions happen on a phone. If it doesn't work there, nothing else matters.
  • Title and thumbnail should not repeat each other. Two halves of one idea. Keep titles short and specific (roughly 50-60 characters so they don't truncate).
  • A/B testing in YouTube Studio beats arguing about it, so give the user 2-3 real options.

Which layout

Full-bleed wins at phone size and is the default. Side by side at 210px, the face reads about twice as large as in the panel version, and panel sublines disappear entirely.

  • layout-fullbleed.html: graded frame edge to edge, 1-3 words in the display font with a heavy ink stroke, placed in empty space (wall, door) and never on the face. Presets stab, stat, bottom, band (a solid accent band across the bottom). Legacy example variants e, f, g, h.
  • layout-panel.html: subject in a clean right-hand square panel, text on a dark left half. Use when the idea needs a diagram (the dot grid) or two lines of text. Example variants a-d.

Run it

# 1. graded frames from the raw, at expressive moments (use the transcript to find them)
bash <skill>/scripts/frames.sh "<YOUR_FOOTAGE_DRIVE>/<folder>/C0001.MP4" "<work>" 178 310 459 503

# 2. look at them before choosing - a contact sheet beats guessing
ffmpeg -i hi-*.jpg -filter_complex hstack=... contact.jpg

# 3. render variants + the phone-size proof sheet
WORK=<work> OUT=<out> node <skill>/scripts/render.mjs variants.json

variants.json:

[{"id":"E-no-time","layout":"fullbleed","v":"e","img":"full-C0001-503.jpg"},
 {"id":"S-custom","layout":"fullbleed","img":"full-C0001-310.jpg",
  "data":{"preset":"stab","head":"TOO BUSY?","kicker":"TRY THIS"}},
 {"id":"B-3x","layout":"panel","v":"b","img":"cut-C0001-97.jpg"}]

Full-bleed text can be passed per variant as data (preset, head, kicker, side, size), or added as a new legacy v block in the layout HTML. Panel text lives in the v blocks of layout-panel.html; add a new variant by adding a block, keeping it to 3 words.

Needs puppeteer-core: reuse ../video-overlays/scripts/node_modules or npm i in the scripts dir. render.mjs launches /Applications/Google Chrome.app/...; change the path on other systems.

Picking the frame

Look for raised eyebrows, an open hand, mid-word mouth. Use the transcript to find a moment with energy, then pull a frame a second in. Dead-eyed frames kill it. The crops in frames.sh assume a 4K (3840x2160) source with the subject at about 46% of frame width; adjust the crop= offsets if your framing differs.

Output

Render at 1920x1080 (YouTube downsamples cleanly), keep under 2MB.

Related

  • paper-edit, video-overlays: the rest of the video pipeline

Individual skills in this repo

This repo contains 7 individual skills — each has its own dedicated page.

beachyphotoandfilm-arch/video-skills

Sync camera footage to a separate lapel or field-recorder track (e.g. Tascam DR-10L and similar) and lay it up in DaVinci Resolve. Stitches the recorder's split files, finds each clip's exact place in the audio by matching words then waveforms, cuts one lapel slice per clip so every later skill works unchanged, checks for clipping, and builds a synced timeline in a new Resolve project. Trigger on "sync the lapel", "sync my audio", "line up the recorder", "I used a lav", or any shoot with a separate audio folder (live talks, events, workshops). Run it before paper-edit or video-reels.

beachyphotoandfilm-arch/video-skills

Turn raw talking-head footage into an assembled rough cut inside DaVinci Resolve. Pulls the clips off your footage drive, transcribes them locally, removes failed takes, restarts, asides and dead air, then creates a new Resolve project with two paper-edit timelines ready for your fine trim. Trigger on "paper edit", "cut my raw footage", "make a rough cut", "clean up this video", or when the user points at a folder of raw camera files for a YouTube video.

beachyphotoandfilm-arch/video-skills

Design a unique Instagram cover for each reel in your house style. Reads what the reel is about, writes a short hook headline, picks one of your own photos from a tagged photo library, lays it out like your existing covers, and proofs it in grid view at phone size. Renders locally, never with AI image generation. Can be called by a scheduling skill before reels are scheduled. Trigger on "make covers", "reel covers", "cover photo for this reel", "design the IG covers".

beachyphotoandfilm-arch/video-skills

Render finished reels out of DaVinci Resolve, write captions in the creator's voice, attach a custom Instagram cover from reel-covers, and schedule them to Instagram, TikTok and YouTube Shorts through Metricool. Handles the Drive upload hop, collision checks against the existing calendar, and cleanup. Trigger on "schedule the reels", "post these", "put these on the calendar", or after reels are approved.

beachyphotoandfilm-arch/video-skills

Build animated brand overlays (motion graphics) for a talking-head video, timed to the speaker's exact words, plus a DaVinci Resolve import file. House style is in-scene type (text beside and behind the speaker's head) plus diagrams. Transcribes the rough cut locally, picks the moments, renders transparent .mov clips, and writes an FCPXML that drops every clip onto the timeline already in position. Trigger on "make animations for this video", "add overlays/motion graphics", "animate this", "b-roll graphics for my YouTube video", or when the user shares a rough cut and asks for graphics.

beachyphotoandfilm-arch/video-skills

Cut vertical Instagram reels out of a long-form talking-head video, hook first. Picks the strongest standalone moments from the transcript, opens each reel on its punchiest line, builds vertical timelines in DaVinci Resolve framed on the speaker's face with the LUT applied, and finishes them: a hook (a held text card, or a word-by-word "build" hook with punch-in) and burned-look captions on V3; this-or-that reels get product graphics and a comment-keyword CTA card. You trim and render, nothing to import. Trigger on "make reels", "clip this for Instagram", "cut some verticals", or after a YouTube video is cut.

beachyphotoandfilm-arch/video-skills

Run the whole YouTube video pipeline end to end, one stage at a time, stopping for the creator's approval at every gate. Paper edit, punch-ins, reels, reel covers, scheduling the reels, overlays, thumbnail and title, then scheduling the long-form video, all through Metricool. Keeps state per video so you can stop anywhere and pick up later. Trigger on "/youtube", "let's do the whole video", "run the video pipeline", "continue the <name> video", or when handed raw footage for a YouTube video.

相关技能