Communitygithub.com

gooseworks-ai/goose-skills

Assemble a cartoon / animated / hand-crafted music-video ad from a config — a sung song carries the whole narrative while N per-bar i2v clips (one recurring animated character, one look pack) are each cut to their BAR window from librosa beat-tracking and hard-concatenated on the bar, VEED-whisper white bold-sans captions in the BOTTOM third (Alignment 2, above the logo bug, no pill) burned from the song's word timings re-spelled against the locked lyrics, a persistent brand logo bug held over the body (suppressed on the end card), and closed on a solid-brand-color PIL end card with the song still playing under it — never AI-rendered text. This is the FREE deterministic assembly stage (cut-to-bar + hard concat + logo bug + captions + end card + song mux); the song, character, keyframes, and clips come from create-music-elevenlabs / create-image-fal / create-video-fal. Use for the cartoon-music-video format.

O que é goose-skills?

goose-skills is a Claude Code agent skill that assemble a cartoon / animated / hand-crafted music-video ad from a config — a sung song carries the whole narrative while N per-bar i2v clips (one recurring animated character, one look pack) are each cut to their BAR window from librosa beat-tracking and hard-concatenated on the bar, VEED-whisper white bold-sans captions in the BOTTOM third (Alignment 2, above the logo bug, no pill) burned from the song's word timings re-spelled against the locked lyrics, a persistent brand logo bug held over the body (suppressed on the end card), and closed on a solid-brand-color PIL end card with the song still playing under it — never AI-rendered text. This is the FREE deterministic assembly stage (cut-to-bar + hard concat + logo bug + captions + end card + song mux); the song, character, keyframes, and clips come from create-music-elevenlabs / create-image-fal / create-video-fal. Use for the cartoon-music-video format.

Funciona com~Claude Code~Codex CLI~Cursor
npx skills add https://github.com/gooseworks-ai/goose-skills/tree/HEAD/skills/ads/capabilities/render-cartoon-music-video

Perguntar na sua IA favorita

Abre um novo chat com esta habilidade de agente já pré-carregada.

Documentação

render-cartoon-music-video

Assemble a cartoon music-video ad from a config: an animated / illustrated / hand-crafted story (stop-motion felt-and-foam, claymation, 2D toon, cut-out) where a sung song is the entire script and one recurring animated character carries every shot, cut to the bar with the climax line on the drop. This capability is the FREE, deterministic assembly — cut-to-bar, hard-concat, the logo bug + captions burn, and the solid-color end card.

scripts/config.example.json is the worked example (Coinbase "Bet on anything", ~55s 1080×1920 9:16, 16 body bars + a 3s end card — its felt look, character, story and song are that demo's picks, never defaults); scripts/PIPELINE.md maps every config block to its source step and scripts/README.md documents the free assembly.

Choices

The creative calls are the caller's (the video-format recipe asks the user); this assembly never picks them. The demo's value is an example only:

  • Art style / look pack (felt stop-motion, claymation, 2D toon, cut-out…) — demo: felt-and-foam diorama.
  • Character — the ONE recurring animated figure — demo: a gender-neutral felt figure in a brand-blue jersey.
  • Story arc — demo: underdog payoff.
  • Song style + lead vocal — demo: hype dance-pop / EDM-trap, male half-rapped lead.

Run

This is the FREE, deterministic assembly stage — it spends nothing. The paid inputs are separate capabilities: the sung song (create-music-elevenlabs, or a user-supplied mp3 + Whisper word timings) beat-tracked with librosa so the BAR GRID sets the timeline; one locked recurring character + one keyframe per bar in one look pack (create-image-fal, Nano Banana); and one Seedance i2v clip per bar (create-video-fal). Given the song + word-timestamps.json + bars.json + one clip per bar + the brand wordmark SVG, render-cartoon-music-video cuts each clip to its bar window, hard-concats on the bar, burns the logo bug + captions, appends the solid-color end card, and muxes the song under it → the master. Re-cuts reuse the existing song / keyframes / clips and cost $0.

Contract (the free assembly)

  • The sung song carries the narrative — no separate VO. The generated/supplied track IS the bed (no VO to duck under); do not add a spoken voiceover or a second bed.
  • Plan the timeline AROUND the delivered song's BAR GRID. Beat-track the song with librosa (assume 4/4); one bar = one shot (default). Snap every tableau window to the bar boundaries — never trim the song to a pre-planned grid.
  • Captions from the song's word timings, re-spelled against the locked lyrics. Whisper mishears shouted accents ("BETS" → "Hearts"); re-spell the timed tokens against the locked lyric file (never edit lyrics to match Whisper). VEED-whisper white bold sans in the BOTTOM third (Alignment 2, margin_v above the bottom-left logo bug), ~4–5 words per cue, NO background pill; captions STOP at the end-card boundary. Never mid-frame — it covers the character (fixed 2026-07). If the host ffmpeg lacks libass, render the cues as timed PIL PNG overlays (ffmpeg overlay=…:enable='between(t,st,en)') at the same bottom placement.
  • Land the climax line on the drop. The climax tableau is timed so the payoff line sits on the sub-bass drop; accent that line.
  • Persistent logo bug, suppressed on the end card. Burn a white brand wordmark bottom-left over the body bars (cairosvg → PIL from the real SVG), suppressed on the end card where the big wordmark dominates.
  • End card via cairosvg + PIL from the real wordmark — never AI-render brand text. Solid brand-color card + white wordmark + subhead, holding ~3s WITH the song still playing under it (afade-out over the tail — no silent tail). A diffusion model garbles a wordmark.
  • FFmpeg composite, deterministic, FREE. Cut each clip to its bar window, hard-concat, burn the logo bug + captions, append the end card, mux the song over the whole video with a 0.5s afade tail, loudnorm I=-14 → a 1080×1920 h264+aac master. No paid calls, no keys. When the host ffmpeg lacks libass/drawtext, BOTH the captions and the logo bug are timed PIL PNG overlays (overlay=x:y:enable='between(t,st,en)'), not an ASS burn.
  • QC PER SCENE, never just the master — the two failure modes are content, not assembly. The upstream keyframe/i2v steps can (a) drift the character off its look (felt → smooth-3D "man" in the demo) partway through (off the chosen look — in the felt demo, felt → smooth-3D) and (b) hallucinate hands — realistic fingers in a hand macro, a pointing finger on a "tap the phone" shot, or a disembodied hand sliding in from the frame edge. Both hide at thumbnail size. Extract a 2 fps contact sheet + a per-bar FACE crop and HAND crop (across each clip's full duration), confirm every bar holds the ONE chosen look (the demo: matte felt) with the ONE character and no human/floating hands, and re-check the served bytes after publish. A drifted bar means regenerating that bar's KEYFRAME (not re-cutting) — see the recipe STEP 3/6.

Individual skills in this repo

This repo contains 7 individual skills — each has its own dedicated page.

gooseworks-ai/goose-skills

Burned-in captions for a finished vertical video, three kinds. transcribe.py gets word timings from the video's own audio through the GooseWorks proxy (fal Whisper, bills the Ads agent, cents); captions.py burns one to three words at a time with Pillow + ffmpeg (no libass needed), either pinned to a split-screen seam (plate 25% above / 75% below) or at a fixed height, in a plate, outline or one-word serif style, with an optional red hook card; plates.py burns per-beat caption blocks (black, one union silhouette, placed in the emptiest band) for formats with no voice. The last caption (the CTA) holds to the final frame. Use as the last step of any video ad.

gooseworks-ai/goose-skills

Assemble a cinematic live-action-style music-video ad from a config — an original sung anthem carries the whole narrative while N 35mm-film-look i2v clips are each cut to their lyric window and hard-concatenated on the beat as a 3-act arc, the anthem muxed at loudnorm I=-14, cinematic lower-third serif captions built from the song's OWN word timings (never Whisper) with the hook line landing on the chorus drop, and closed on a brand end card composited from the real asset — never AI-rendered text. This is the FREE deterministic assembly stage (cut-to-window + hard concat + anthem mux + captions + end card); the anthem, keyframes, and clips come from create-music-elevenlabs / create-image-gpt-image-fal / create-video-fal. Use for the cinematic-music-video format.

gooseworks-ai/goose-skills

Assemble an editorial-motion podcast-clip ad from a config — a real clipped podcast MP3 carries the narrative while N flat limited-palette editorial-illustration keyframes (one look pack) are animated NOT by generative i2v but by DETERMINISTIC ffmpeg ken-burns (zoompan) + hard cuts (no crossfades, which expose geometric drift), each beat snapped to its spoken line, the real audio muxed, Whisper-driven captions burned only mid-sentence, and closed on a PIL brand end card — never AI-rendered text. This is the FREE deterministic assembly stage (ffmpeg ken-burns + hard concat + audio mux + captions + end card); the real audio is clipped from source and the keyframes come from create-image-fal. Use for the editorial-motion-podcast format.

gooseworks-ai/goose-skills

Generate a single 4-15s vertical video clip with ByteDance Seedance 2.0 reference-to-video via fal.ai. Multi-image reference (avatar + product + setting), native lip-synced VO + ambient audio (generate-audio on by default), internal multi-cut handling within one render. Routes through the GooseWorks FAL proxy (bills the Ads agent). The default clip atom for AI-creator UGC ads built on the NB2 + Seedance architecture. Validated on beauty-by-earth/video-01.

gooseworks-ai/goose-skills

Mandatory pre-publish review gate for a UGC video render. Transcribes the finished render's AUDIO with Whisper and word-diffs it against the approved spoken script, then gates set_final_render — blocking a render whose generated audio mis-voices a word (e.g. the approved "human-vetted" spoken as "human witted"), drops an approved phrase, or comes back silent. Runnable, gating counterpart to content-goose's review-transcript-integrity atom. Every ugc-video-formats recipe runs this after render and BEFORE set_final_render.

gooseworks-ai/goose-skills

Render pixel-accurate Apple Notes (iPhone, light mode) screenshot mockups from a JSON note spec. Outputs HTML + PNG at the iPhone 16/15 Pro native 1180×2556. Supports paragraphs, images, checklists, dividers, autocorrect underline, smart quotes, and an optional iOS keyboard chrome overlay used by the parent video-ad molecule.

gooseworks-ai/goose-skills

Find speakers, hosts, and guest profiles at conferences and events on Luma. Two modes - free direct scrape for hosts, or Apify-powered search for full guest profiles with LinkedIn/Twitter/bio.

Habilidades Relacionadas