Communitygithub.com

gooseworks-ai/goose-skills

Assemble an editorial-motion podcast-clip ad from a config — a real clipped podcast MP3 carries the narrative while N flat limited-palette editorial-illustration keyframes (one look pack) are animated NOT by generative i2v but by DETERMINISTIC ffmpeg ken-burns (zoompan) + hard cuts (no crossfades, which expose geometric drift), each beat snapped to its spoken line, the real audio muxed, Whisper-driven captions burned only mid-sentence, and closed on a PIL brand end card — never AI-rendered text. This is the FREE deterministic assembly stage (ffmpeg ken-burns + hard concat + audio mux + captions + end card); the real audio is clipped from source and the keyframes come from create-image-fal. Use for the editorial-motion-podcast format.

Qu'est-ce que goose-skills ?

goose-skills is a Claude Code agent skill that assemble an editorial-motion podcast-clip ad from a config — a real clipped podcast MP3 carries the narrative while N flat limited-palette editorial-illustration keyframes (one look pack) are animated NOT by generative i2v but by DETERMINISTIC ffmpeg ken-burns (zoompan) + hard cuts (no crossfades, which expose geometric drift), each beat snapped to its spoken line, the real audio muxed, Whisper-driven captions burned only mid-sentence, and closed on a PIL brand end card — never AI-rendered text. This is the FREE deterministic assembly stage (ffmpeg ken-burns + hard concat + audio mux + captions + end card); the real audio is clipped from source and the keyframes come from create-image-fal. Use for the editorial-motion-podcast format.

Compatible avec~Claude Code~Codex CLI~Cursor
npx skills add https://github.com/gooseworks-ai/goose-skills/tree/HEAD/skills/ads/capabilities/render-editorial-motion-podcast

Demander à votre IA préférée

Ouvre une nouvelle conversation avec cette compétence d'agent déjà préchargée.

Documentation

render-editorial-motion-podcast

Assemble an editorial-motion podcast-clip ad from a config: a real clipped podcast audio line carries the whole narrative and every visual beat is timed to the sentence it describes, in ONE flat, strictly limited-palette editorial-illustration look pack ("a magazine spot-illustration that moves" — the style and palette are the caller's choice). The motion is not generative video but deterministic ffmpeg ken-burns on static keyframes, so it reads as a printed page that moves. This capability is that FREE, deterministic assembly — the ffmpeg motion, hard-concat, audio mux, caption burn, and PIL end card.

scripts/config.example.json is the worked example (Klarify "Rat Park", ~40.8s 1080×1920 9:16, 6 beats — its 2-tone Niemann look, cream/charcoal/sage palette and Rat Park metaphor are that demo's picks, never defaults); scripts/PIPELINE.md maps every config block to its source step and scripts/README.md documents the free assembly.

Choices

The creative calls are the caller's (the video-format recipe asks the user); this assembly never picks them. The demo's value is an example only:

  • Illustration style — demo: bold flat 2-tone silhouettes, halftone (Niemann / Steinberg lineage).
  • Palette — demo: cream #F4EBD9 + charcoal #1A1A1A + one sage #86987A accent.
  • Central metaphor — demo: the Rat Park study.
  • Tone / voice — demo: warm and reflective, the real podcast host (no generated voice).

Run

This is the FREE, deterministic assembly stage — it spends nothing on the motion layer. The paid inputs are separate: the real podcast MP3 is clipped from source (free ffmpeg) with its Whisper word timings, and one editorial-illustration keyframe per beat (chained ref images so cage/character geometry holds) comes from create-image-fal (Nano Banana). Given the clipped audio + words.json + the per-beat keyframes + the real brand wordmark PNG, render-editorial-motion-podcast renders each keyframe as a ken-burns segment, hard-concats on the beat, muxes the real audio, burns the mid-sentence captions, and composites the PIL end card → the master. Re-cuts reuse the existing audio / keyframes and cost $0.

Contract (the free assembly)

  • A spoken narration carries the whole spot — no generated SONG. Mux the provided narration MP3 (-map 0:v:0 -map 1:a:0) — a real clipped podcast line (preferred) OR an approved generated VO (create-vo-elevenlabs). Never a sung/generated track. (Clip-vs-generate is the recipe's STEP-0 intake decision — if no source episode is supplied, ASK the user.)
  • NO generative i2v — deterministic ffmpeg ken-burns only. Animate each static keyframe with zoompan (push-in / pull-back, 1.0→~1.06×, 24fps); Seedance/Kling are photoreal-trained and invent naturalistic middle states that collapse the flat limited-palette look look. Never -loop 1 with zoompan d=N (it balloons the duration); feed a single image and clamp with -t + trim.
  • Hard cuts on the beat — no crossfades. Crossfades ghost two drifting cages through each other; hard-concat each beat's segments and split long beats into micro-cuts (target 8–10 distinct visual moments). Each beat's visual STARTS within ~0.5s of its spoken line.
  • Captions from Whisper word-timestamps, ON only mid-sentence. Burn frosted-subtle captions while the speaker talks; leave silent/reflective beats and the end card uncaptioned. THREE mandatory rules (each bit us in prod — bake them in):
    1. NON-OVERLAP — clamp every line to END before the next STARTS (end = min(last_word_end + ~0.15, next_start - 0.03)). Two boxes must never stack at the same spot; an end-tail bleeding into the next window is the #1 caption bug.
    2. SAFE AREA — captions sit in the lower third, so the keyframe's subject must stay in the upper ~75% (see the recipe's look_pack.caption_safe_area). If a finished keyframe's subject intrudes into the caption band, deterministically shift the subject UP into the empty top space (PIL: paste up ~0.24H onto a canvas pre-filled with the exact paper color from a clean corner) — never let the box sit on the subject.
    3. BURN ENGINE — prefer libass (ass/subtitles filter), but check ffmpeg -filters first: many builds (Homebrew) lack libass/drawtext. If absent, use the deterministic overlay fallback — render each line as a transparent PNG (frosted rounded box + white text, PIL) and composite via the ffmpeg overlay filter with timed enable='between(t,st,en)' windows. Same look, no libass.
  • End card via PIL from the real wordmark PNG — never AI-render brand text. The lockup is composited deterministically (stretched-gradient bg + feathered mascot crop + wordmark + tagline with a system font); a diffusion model garbles a wordmark ("therapits"). The video runs a ~1.5s silent hold past the audio on the end card (fade first/last 0.3s).
  • FFmpeg composite, deterministic, FREE. Ken-burns each keyframe, hard-concat, mux the real audio, burn the captions, hold on the end card → a 1080×1920 h264+aac master. No paid calls.

Individual skills in this repo

This repo contains 7 individual skills — each has its own dedicated page.

gooseworks-ai/goose-skills

Burned-in captions for a finished vertical video, three kinds. transcribe.py gets word timings from the video's own audio through the GooseWorks proxy (fal Whisper, bills the Ads agent, cents); captions.py burns one to three words at a time with Pillow + ffmpeg (no libass needed), either pinned to a split-screen seam (plate 25% above / 75% below) or at a fixed height, in a plate, outline or one-word serif style, with an optional red hook card; plates.py burns per-beat caption blocks (black, one union silhouette, placed in the emptiest band) for formats with no voice. The last caption (the CTA) holds to the final frame. Use as the last step of any video ad.

gooseworks-ai/goose-skills

Assemble a cartoon / animated / hand-crafted music-video ad from a config — a sung song carries the whole narrative while N per-bar i2v clips (one recurring animated character, one look pack) are each cut to their BAR window from librosa beat-tracking and hard-concatenated on the bar, VEED-whisper white bold-sans captions in the BOTTOM third (Alignment 2, above the logo bug, no pill) burned from the song's word timings re-spelled against the locked lyrics, a persistent brand logo bug held over the body (suppressed on the end card), and closed on a solid-brand-color PIL end card with the song still playing under it — never AI-rendered text. This is the FREE deterministic assembly stage (cut-to-bar + hard concat + logo bug + captions + end card + song mux); the song, character, keyframes, and clips come from create-music-elevenlabs / create-image-fal / create-video-fal. Use for the cartoon-music-video format.

gooseworks-ai/goose-skills

Assemble a cinematic live-action-style music-video ad from a config — an original sung anthem carries the whole narrative while N 35mm-film-look i2v clips are each cut to their lyric window and hard-concatenated on the beat as a 3-act arc, the anthem muxed at loudnorm I=-14, cinematic lower-third serif captions built from the song's OWN word timings (never Whisper) with the hook line landing on the chorus drop, and closed on a brand end card composited from the real asset — never AI-rendered text. This is the FREE deterministic assembly stage (cut-to-window + hard concat + anthem mux + captions + end card); the anthem, keyframes, and clips come from create-music-elevenlabs / create-image-gpt-image-fal / create-video-fal. Use for the cinematic-music-video format.

gooseworks-ai/goose-skills

Generate a single 4-15s vertical video clip with ByteDance Seedance 2.0 reference-to-video via fal.ai. Multi-image reference (avatar + product + setting), native lip-synced VO + ambient audio (generate-audio on by default), internal multi-cut handling within one render. Routes through the GooseWorks FAL proxy (bills the Ads agent). The default clip atom for AI-creator UGC ads built on the NB2 + Seedance architecture. Validated on beauty-by-earth/video-01.

gooseworks-ai/goose-skills

Mandatory pre-publish review gate for a UGC video render. Transcribes the finished render's AUDIO with Whisper and word-diffs it against the approved spoken script, then gates set_final_render — blocking a render whose generated audio mis-voices a word (e.g. the approved "human-vetted" spoken as "human witted"), drops an approved phrase, or comes back silent. Runnable, gating counterpart to content-goose's review-transcript-integrity atom. Every ugc-video-formats recipe runs this after render and BEFORE set_final_render.

gooseworks-ai/goose-skills

Render pixel-accurate Apple Notes (iPhone, light mode) screenshot mockups from a JSON note spec. Outputs HTML + PNG at the iPhone 16/15 Pro native 1180×2556. Supports paragraphs, images, checklists, dividers, autocorrect underline, smart quotes, and an optional iOS keyboard chrome overlay used by the parent video-ad molecule.

gooseworks-ai/goose-skills

Find speakers, hosts, and guest profiles at conferences and events on Luma. Two modes - free direct scrape for hosts, or Apify-powered search for full guest profiles with LinkedIn/Twitter/bio.

Skills associés