Communitygithub.com

gooseworks-ai/goose-skills

Assemble a cinematic live-action-style music-video ad from a config — an original sung anthem carries the whole narrative while N 35mm-film-look i2v clips are each cut to their lyric window and hard-concatenated on the beat as a 3-act arc, the anthem muxed at loudnorm I=-14, cinematic lower-third serif captions built from the song's OWN word timings (never Whisper) with the hook line landing on the chorus drop, and closed on a brand end card composited from the real asset — never AI-rendered text. This is the FREE deterministic assembly stage (cut-to-window + hard concat + anthem mux + captions + end card); the anthem, keyframes, and clips come from create-music-elevenlabs / create-image-gpt-image-fal / create-video-fal. Use for the cinematic-music-video format.

O que é goose-skills?

goose-skills is a Claude Code agent skill that assemble a cinematic live-action-style music-video ad from a config — an original sung anthem carries the whole narrative while N 35mm-film-look i2v clips are each cut to their lyric window and hard-concatenated on the beat as a 3-act arc, the anthem muxed at loudnorm I=-14, cinematic lower-third serif captions built from the song's OWN word timings (never Whisper) with the hook line landing on the chorus drop, and closed on a brand end card composited from the real asset — never AI-rendered text. This is the FREE deterministic assembly stage (cut-to-window + hard concat + anthem mux + captions + end card); the anthem, keyframes, and clips come from create-music-elevenlabs / create-image-gpt-image-fal / create-video-fal. Use for the cinematic-music-video format.

Funciona com~Claude Code~Codex CLI~Cursor
npx skills add https://github.com/gooseworks-ai/goose-skills/tree/HEAD/skills/ads/capabilities/render-cinematic-music-video

Perguntar na sua IA favorita

Abre um novo chat com esta habilidade de agente já pré-carregada.

Documentação

render-cinematic-music-video

Assemble a cinematic music-video ad from a config: a live-action-STYLE short film where an original sung anthem is the score and every visual beat is a shot-on-film tableau (Kodak Portra grain, light leaks, natural light, handheld imperfection) timed to the lyrics, arranged as a 3-act arc (its shape is the user's story_arc choice) with the hook line on the chorus drop. This capability is the FREE, deterministic assembly — cut-to-window, hard-concat, anthem mux, caption burn, and the brand end card.

scripts/config.example.json is the worked example (Hype and Vice "Game Day Girls", ~28s 1080×1920 9:16, 14 tableaux) — copy its structure, never its creative values; scripts/PIPELINE.md maps every config block to its source step and scripts/README.md documents the free assembly.

Run

This is the FREE, deterministic assembly stage — it spends nothing. The paid inputs are separate capabilities: the sung anthem (create-music-elevenlabs, force_instrumental false — the lyrics ARE the script, returns mp3 + words_timestamps); one 35mm-film keyframe per beat in one look pack (create-image-gpt-image-fal); and one Kling 3.0 i2v clip per beat (create-video-fal). Given the delivered anthem + words.json + one clip per beat + the brand end-card asset, render-cinematic-music-video cuts each clip to its lyric window, hard-concats on the beat, muxes the anthem, burns the cinematic lower-third captions, and overlays the end card → the master. Re-cuts reuse the existing anthem / keyframes / clips and cost $0.

Choices

The creative content this stage assembles is decided upstream by the format's choices, asked of the user before any paid step. The worked example's values are examples, never defaults:

  • song_style — the anthem's genre, mood, BPM and arrangement. The demo used a triumphant 120-BPM indie-pop cinematic anthem.
  • vocalist — who sings it. The demo used a confident-but-warm female lead.
  • cast — who is on screen in the tableaux. The demo used four young college women.
  • setting — place + time of day (drives the look pack's light and palette). The demo used an American college town on game day, autumn, golden hour.
  • story_arc — the 3-act shape of the tableaux and lyrics. The demo used one game day, morning → stadium peak → twilight, with an origin-story wink.

This assembly reads them only through the config (tableaux[] windows + captions, the anthem, captions.accent_words, end_card.*); it hardcodes none of them.

Contract (the free assembly)

  • The sung anthem carries the story — no separate VO. The generated ElevenLabs track IS the bed and the script (force_instrumental false); do not add a spoken voiceover or a second bed.
  • Plan the timeline AROUND the delivered anthem. The anthem is generated first and reshapes/overshoots length; snap every tableau boundary to the lyric-phrase edges in the returned word timings — never trim the anthem to a pre-planned grid.
  • Captions from the anthem's OWN word timings, not Whisper. Derive words.json from the music model's words_timestamps, chunk ~4 words at lyric boundaries, and burn cinematic lower-third serif captions (--placement low) with the hook line accent-treated (bold-italic). Whisper on sung audio returns "🎵 Music Playing 🎵".
  • Land the hook on the chorus drop. The hero tableau (is_hook) is timed so the load-bearing line sits on the chorus drop; accent that line in the captions.
  • One cinematic look pack + hard cuts on the beat. The look pack (named film stock + grain + flares + handheld) plus a continuity anchor holds every clip together; cut each clip to its lyric window and hard-concat (one optional match-cut into the hero reveal) — no dissolves.
  • End card from the real brand asset — never AI-render the wordmark as diffusion text. Composite the lockup (PIL/ffmpeg) from the real asset or use a designed keyframe base; diffusion garbles a wordmark.
  • FFmpeg composite, deterministic, FREE. Normalize each clip to its beat-locked window, hard-concat, mux the anthem (afade in/out + loudnorm I=-14), burn the caption ASS, overlay the end card → a 1080×1920 h264+aac master. No paid calls, no keys.

Individual skills in this repo

This repo contains 7 individual skills — each has its own dedicated page.

gooseworks-ai/goose-skills

Burned-in captions for a finished vertical video, three kinds. transcribe.py gets word timings from the video's own audio through the GooseWorks proxy (fal Whisper, bills the Ads agent, cents); captions.py burns one to three words at a time with Pillow + ffmpeg (no libass needed), either pinned to a split-screen seam (plate 25% above / 75% below) or at a fixed height, in a plate, outline or one-word serif style, with an optional red hook card; plates.py burns per-beat caption blocks (black, one union silhouette, placed in the emptiest band) for formats with no voice. The last caption (the CTA) holds to the final frame. Use as the last step of any video ad.

gooseworks-ai/goose-skills

Assemble a cartoon / animated / hand-crafted music-video ad from a config — a sung song carries the whole narrative while N per-bar i2v clips (one recurring animated character, one look pack) are each cut to their BAR window from librosa beat-tracking and hard-concatenated on the bar, VEED-whisper white bold-sans captions in the BOTTOM third (Alignment 2, above the logo bug, no pill) burned from the song's word timings re-spelled against the locked lyrics, a persistent brand logo bug held over the body (suppressed on the end card), and closed on a solid-brand-color PIL end card with the song still playing under it — never AI-rendered text. This is the FREE deterministic assembly stage (cut-to-bar + hard concat + logo bug + captions + end card + song mux); the song, character, keyframes, and clips come from create-music-elevenlabs / create-image-fal / create-video-fal. Use for the cartoon-music-video format.

gooseworks-ai/goose-skills

Assemble an editorial-motion podcast-clip ad from a config — a real clipped podcast MP3 carries the narrative while N flat limited-palette editorial-illustration keyframes (one look pack) are animated NOT by generative i2v but by DETERMINISTIC ffmpeg ken-burns (zoompan) + hard cuts (no crossfades, which expose geometric drift), each beat snapped to its spoken line, the real audio muxed, Whisper-driven captions burned only mid-sentence, and closed on a PIL brand end card — never AI-rendered text. This is the FREE deterministic assembly stage (ffmpeg ken-burns + hard concat + audio mux + captions + end card); the real audio is clipped from source and the keyframes come from create-image-fal. Use for the editorial-motion-podcast format.

gooseworks-ai/goose-skills

Generate a single 4-15s vertical video clip with ByteDance Seedance 2.0 reference-to-video via fal.ai. Multi-image reference (avatar + product + setting), native lip-synced VO + ambient audio (generate-audio on by default), internal multi-cut handling within one render. Routes through the GooseWorks FAL proxy (bills the Ads agent). The default clip atom for AI-creator UGC ads built on the NB2 + Seedance architecture. Validated on beauty-by-earth/video-01.

gooseworks-ai/goose-skills

Mandatory pre-publish review gate for a UGC video render. Transcribes the finished render's AUDIO with Whisper and word-diffs it against the approved spoken script, then gates set_final_render — blocking a render whose generated audio mis-voices a word (e.g. the approved "human-vetted" spoken as "human witted"), drops an approved phrase, or comes back silent. Runnable, gating counterpart to content-goose's review-transcript-integrity atom. Every ugc-video-formats recipe runs this after render and BEFORE set_final_render.

gooseworks-ai/goose-skills

Render pixel-accurate Apple Notes (iPhone, light mode) screenshot mockups from a JSON note spec. Outputs HTML + PNG at the iPhone 16/15 Pro native 1180×2556. Supports paragraphs, images, checklists, dividers, autocorrect underline, smart quotes, and an optional iOS keyboard chrome overlay used by the parent video-ad molecule.

gooseworks-ai/goose-skills

Find speakers, hosts, and guest profiles at conferences and events on Luma. Two modes - free direct scrape for hosts, or Apify-powered search for full guest profiles with LinkedIn/Twitter/bio.

Habilidades Relacionadas