Communitygithub.com

gooseworks-ai/goose-skills

Burned-in captions for a finished vertical video, three kinds. transcribe.py gets word timings from the video's own audio through the GooseWorks proxy (fal Whisper, bills the Ads agent, cents); captions.py burns one to three words at a time with Pillow + ffmpeg (no libass needed), either pinned to a split-screen seam (plate 25% above / 75% below) or at a fixed height, in a plate, outline or one-word serif style, with an optional red hook card; plates.py burns per-beat caption blocks (black, one union silhouette, placed in the emptiest band) for formats with no voice. The last caption (the CTA) holds to the final frame. Use as the last step of any video ad.

goose-skills とは?

goose-skills is a Claude Code agent skill that burned-in captions for a finished vertical video, three kinds. transcribe.py gets word timings from the video's own audio through the GooseWorks proxy (fal Whisper, bills the Ads agent, cents); captions.py burns one to three words at a time with Pillow + ffmpeg (no libass needed), either pinned to a split-screen seam (plate 25% above / 75% below) or at a fixed height, in a plate, outline or one-word serif style, with an optional red hook card; plates.py burns per-beat caption blocks (black, one union silhouette, placed in the emptiest band) for formats with no voice. The last caption (the CTA) holds to the final frame. Use as the last step of any video ad.

対応~Claude Code~Codex CLI~Cursor
npx skills add https://github.com/gooseworks-ai/goose-skills/tree/HEAD/skills/ads/capabilities/caption-burn

お気に入りのAIに質問する

このエージェントスキルを事前に読み込んだ状態で新しいチャットを開きます。

ドキュメント

caption-burn

Captions timed to what is actually said in the finished video, not to the script's estimate.

Run

python transcribe.py --media reel.mp4 --out reel.words.json          # paid, cents
python captions.py --video reel.mp4 --beats cutlist.aligned.json \
    --words reel.words.json --out final.mp4 [--style plate|outline] [--highlight Brand]

--beats supplies the lines (vo), their timing and, for split layouts, seam, size and each beat's state. Any file with beats: [{start, end, vo}] works.

Placement and style

  • --anchor seam (the default when the beats have a seam): on split beats the plate is pinned to the seam, positioned by the plate, not the text: 25% of the plate above the line, 75% below. Full-frame beats use --full-y.

  • --anchor fixed --y 0.62: every caption's plate centred at that fraction of the height.

  • plate (default): white bold on a dark grey rounded plate, 1–2 words, cap ~0.019 H.

  • outline: white bold with a dark outline, no plate, 1–3 words, cap ~0.034 H.

  • --highlight WORD colours that word yellow (the CTA keyword). Repeatable.

  • serif-word: ONE word at a time, heavy serif (Georgia Bold), white with a black outline, on a fixed baseline at 0.77 H (the screen-insert look). The highlight word is quoted.

  • --card "LINE ONE|LINE TWO" --card-until 4.7: a white rounded hook card with two lines of heavy red capitals near the top, for the opening seconds. ~14 characters a line.

Captions for a format with no voice (plates.py)

python plates.py --video walk.mp4 --beats cutlist.json --out captioned.mp4 [--logo logo.png]

Each beat's caption (a string or list of lines) shows for the whole beat on ONE black block (square rectangles unioned, then rounded as a single silhouette: rounding each line leaves seams), lines left-aligned, the block centred on its widest line. It goes in the emptiest band of that beat's frame unless the beat pins cap_y. logo: true on a beat hangs the logo tile under the block. Write lines a person would type: the same short "fragment. fragment." shape three times reads as AI-written. No emoji twice.

Rules

  1. Transcribe the FINISHED audio. Joining takes and aligning lines shifts timing; only the final video's audio gives correct cues.
  2. The last caption holds to the last frame. It is the call to action.
  3. A word Whisper writes differently ("200" for "two hundred") is interpolated between its neighbours rather than dropped or stretched over the whole line.
  4. Without --words timing is estimated from syllables. Use that to judge placement, never to ship.
  5. Fonts: a bold sans is found on macOS, Linux or Windows; if none is present, Roboto Bold is fetched once into ~/.cache/gooseworks/fonts. --font or GW_CAPTION_FONT overrides.

Individual skills in this repo

This repo contains 7 individual skills — each has its own dedicated page.

gooseworks-ai/goose-skills

Assemble a cartoon / animated / hand-crafted music-video ad from a config — a sung song carries the whole narrative while N per-bar i2v clips (one recurring animated character, one look pack) are each cut to their BAR window from librosa beat-tracking and hard-concatenated on the bar, VEED-whisper white bold-sans captions in the BOTTOM third (Alignment 2, above the logo bug, no pill) burned from the song's word timings re-spelled against the locked lyrics, a persistent brand logo bug held over the body (suppressed on the end card), and closed on a solid-brand-color PIL end card with the song still playing under it — never AI-rendered text. This is the FREE deterministic assembly stage (cut-to-bar + hard concat + logo bug + captions + end card + song mux); the song, character, keyframes, and clips come from create-music-elevenlabs / create-image-fal / create-video-fal. Use for the cartoon-music-video format.

gooseworks-ai/goose-skills

Assemble a cinematic live-action-style music-video ad from a config — an original sung anthem carries the whole narrative while N 35mm-film-look i2v clips are each cut to their lyric window and hard-concatenated on the beat as a 3-act arc, the anthem muxed at loudnorm I=-14, cinematic lower-third serif captions built from the song's OWN word timings (never Whisper) with the hook line landing on the chorus drop, and closed on a brand end card composited from the real asset — never AI-rendered text. This is the FREE deterministic assembly stage (cut-to-window + hard concat + anthem mux + captions + end card); the anthem, keyframes, and clips come from create-music-elevenlabs / create-image-gpt-image-fal / create-video-fal. Use for the cinematic-music-video format.

gooseworks-ai/goose-skills

Assemble an editorial-motion podcast-clip ad from a config — a real clipped podcast MP3 carries the narrative while N flat limited-palette editorial-illustration keyframes (one look pack) are animated NOT by generative i2v but by DETERMINISTIC ffmpeg ken-burns (zoompan) + hard cuts (no crossfades, which expose geometric drift), each beat snapped to its spoken line, the real audio muxed, Whisper-driven captions burned only mid-sentence, and closed on a PIL brand end card — never AI-rendered text. This is the FREE deterministic assembly stage (ffmpeg ken-burns + hard concat + audio mux + captions + end card); the real audio is clipped from source and the keyframes come from create-image-fal. Use for the editorial-motion-podcast format.

gooseworks-ai/goose-skills

Generate a single 4-15s vertical video clip with ByteDance Seedance 2.0 reference-to-video via fal.ai. Multi-image reference (avatar + product + setting), native lip-synced VO + ambient audio (generate-audio on by default), internal multi-cut handling within one render. Routes through the GooseWorks FAL proxy (bills the Ads agent). The default clip atom for AI-creator UGC ads built on the NB2 + Seedance architecture. Validated on beauty-by-earth/video-01.

gooseworks-ai/goose-skills

Mandatory pre-publish review gate for a UGC video render. Transcribes the finished render's AUDIO with Whisper and word-diffs it against the approved spoken script, then gates set_final_render — blocking a render whose generated audio mis-voices a word (e.g. the approved "human-vetted" spoken as "human witted"), drops an approved phrase, or comes back silent. Runnable, gating counterpart to content-goose's review-transcript-integrity atom. Every ugc-video-formats recipe runs this after render and BEFORE set_final_render.

gooseworks-ai/goose-skills

Render pixel-accurate Apple Notes (iPhone, light mode) screenshot mockups from a JSON note spec. Outputs HTML + PNG at the iPhone 16/15 Pro native 1180×2556. Supports paragraphs, images, checklists, dividers, autocorrect underline, smart quotes, and an optional iOS keyboard chrome overlay used by the parent video-ad molecule.

gooseworks-ai/goose-skills

Find speakers, hosts, and guest profiles at conferences and events on Luma. Two modes - free direct scrape for hosts, or Apify-powered search for full guest profiles with LinkedIn/Twitter/bio.

関連スキル