Communitygithub.com

freudianflips/breakout-launch-videos

Tool manual under launch-film (voice casting, one take per line, fitting lines, stitching, forced alignment, the vintage chain, the local clone and the free scratch voice). Produce the voice-over of a launch film. Casting across every Higgsfield preset voice, one file per script line fitted into its frame, pace and length rules, word-level timestamps by forced alignment so every key word lands on its visual, and fixes for lines that run long. Use after launch-story has an approved script and before picture timing is locked, for "voice-over", "narrator", "cast a voice", "record the lines", "align the words", "scratch voice".

Qu'est-ce que breakout-launch-videos ?

breakout-launch-videos is a Claude Code agent skill that tool manual under launch-film (voice casting, one take per line, fitting lines, stitching, forced alignment, the vintage chain, the local clone and the free scratch voice). Produce the voice-over of a launch film. Casting across every Higgsfield preset voice, one file per script line fitted into its frame, pace and length rules, word-level timestamps by forced alignment so every key word lands on its visual, and fixes for lines that run long. Use after launch-story has an approved script and before picture timing is locked, for "voice-over", "narrator", "cast a voice", "record the lines", "align the words", "scratch voice".

Compatible avec✓Claude Code~Codex CLI~Cursor
npx skills add https://github.com/freudianflips/breakout-launch-videos/tree/HEAD/skills/launch-voice

Demander à votre IA préférée

Ouvre une nouvelle conversation avec cette compétence d'agent déjà préchargée.

Documentation

Launch voice

Tool manual under launch-film. Start at skills/launch-film/SKILL.md: it holds the core, the hook-plus-body template and the lessons. Where this page differs, launch-film wins.

In a modern launch film the voice leads and the picture illustrates it word by word: when the voice says "tagged", something gets tagged. That only works when every word has a timestamp.

Providers

  • Default: Higgsfield on your own account. text2speech_v2 (engines elevenlabs, minimax, seed_speech, and others per voice) for casting across about 100 preset voices, about 0.3 credits a line. Seed Audio for a narrator you steer with a prompt and audio references. scripts/hf_voice.py wraps the CLI.
  • Free scratch voice: scripts/scratch_voice.py makes placeholder takes with macOS say (or espeak-ng on Linux). Build and time the whole film on it, label every review "scratch voice", then swap the real takes in.
  • Free local clone: scripts/clone_local.py (Chatterbox, MIT) clones your approved narrator from 8 to 10 s of their lines. A stand-in, not a master. Only clone a voice you have the rights to.
  • Review voices inside a full film, never as loose clips. Keep a list of rejected voices so they are never proposed again.

Casting

  1. python3 skills/launch-voice/scripts/hf_voice.py cast "<the film's opening line>" film/vo/audition speaks the line with every preset voice (about 34 credits for about 110 voices), checks each take with whisper, and measures pace and pitch movement. ranking.md lists exact takes first, nearest 2.5 words a second and about 2 semitones of movement (conversational, not announced). --only Name,Name recasts a few.
  2. Take the top 6 or so across genders into the full script (hf_voice.py lines), then fit them (below). A voice that needs more than 12 % speed-up to fit is too slow for the script: drop it or cut words.
  3. Build a full film per surviving voice and let the owner pick. ElevenLabs takes varied least in pace and pitch in testing, so it is the default engine.

Length and pace

  • Target 2.2 to 2.7 spoken words a second. A strong reference launch ran 2.4 (151 words in 63 s).
  • A 20 s film carries about 45 words; 30 s about 70; 60 s about 140. Cut words before you speed the voice up.
  • One line per storyboard frame, 6 to 20 words, phrased so the key noun or verb falls on a beat.

Procedure for the film template

  1. The script lives in film/film.json lines. Record 1 file per line into film/vo/body/line-<id>.wav: real takes with hf_voice.py lines (write the lines as {"lines": [{"line": "02", "text": "..."}]}), or scratch takes with python3 skills/launch-voice/scripts/scratch_voice.py film/film.json (add --hook film/hooks/<name> to voice a hook line too).
  2. Run python3 skills/launch-voice/scripts/prep_lines.py film/film.json (--align force to insist on forced alignment; with uv installed it fetches torch on the first run, about 2 GB). It writes film/vo/body/lines.json with each line's duration and word times: forced alignment when torch is available, else whisper.cpp, else an even estimate flagged estimate.
  3. film/build.py places the lines on the film timeline and every picture cue follows its word. A new take is 1 prep run plus a rebuild.

Procedure for a free-form film

  1. vo/script-lines.json ({"lines": [{"line": "01", "text": "..."}]}), then one file per line: hf_voice.py lines vo/script-lines.json vo/takes/<voice> --voice <Name>. Text, voice, engine and job ids land in lines.json.
  2. Fit: python3 skills/launch-voice/scripts/fit_lines.py vo/placement.plan.json vo/takes/<voice> vo/versions/<voice>. The plan gives each line its earliest start, its breath gap and, where a frame must end, end_by. Pauses over 0.25 s inside a line shrink to 0.22 s; a run of lines that still overflows is sped up as one, capped at 12 %. Then stitch: python3 skills/launch-voice/scripts/stitch.py vo/versions/<voice>/placement.json vo/versions/<voice> vo/versions/<voice>/vo.wav.
  3. Word times per line. For anything a visual must land on, use forced alignment: uv run -q --with torch --with torchaudio --with soundfile python skills/launch-voice/scripts/force_align.py line.wav "the exact words". For a quick read, align_lines.py runs whisper.cpp line by line. Anchors resolve in speaking order (each the first match after the previous one), so any voice re-times the film. A single whisper pass over the whole film drifts near line ends; per-line passes do not.
  4. Hand the word times to launch-motion (anchor each build to its word) and launch-sound (events follow words).

Vintage narrator

For an archival or documentary opening: python3 skills/launch-voice/scripts/vintage.py IN.wav OUT.wav --crackle crackle.wav band-limits to 150 Hz to 5.8 kHz, adds tape saturation, a slow wow and a real vinyl crackle bed (generate 1 with Seed Audio). Seed Audio can also speak a line in a vintage narrator voice directly ("a warm, deep 1960s documentary narrator on an old vinyl record ... exactly these words: ..."); it reads slowly and may rephrase ("gotta" became "got to"), so check with whisper.

Sync rules

  • A key word lands with its visual within 2 frames. Start builds on the word's first frame; if you let a visual lead its word, keep it under 4 frames and do it everywhere.
  • Let the voice breathe at scene changes: 0.25 to 0.5 s between lines.
  • The last line ends at least 1 s before the logo; the logo gets its own near silence (launch-mix drop-out).

When the voice provider is unavailable

Build against the scratch voice, label every review as scratch, and keep the picture anchored to the aligned words, so swapping in the real voice is 1 prep run plus a render. Never switch voice providers or accounts without the owner's go.

Fixes

  • A line runs long by under 5 %: ffmpeg -af atempo=1.04. More than that: rewrite the line.
  • A word is mispronounced: respell it phonetically in the line's say field (for hf_voice.py), never in the script.
  • Whisper mishears a soft ending: use forced alignment or shorten the anchor prefix; do not re-record.
  • Digits: spell them out in the text you align ("twenty four", not "24").
  • The whisper base model mishears product names. Check final words with the small model or larger (WHISPER_MODEL).

Individual skills in this repo

This repo contains 9 individual skills — each has its own dedicated page.

freudianflips/breakout-launch-videos

Build the brand kit a launch film uses from a website. Runs tools/brand-from-url.mjs on the reader's site to extract branding (colours, fonts, logo, name, tagline, copy, product images), then reviews the preview sheet with the reader and fixes any wrong role by hand. Use for "brand from website", "brand kit", "extract branding", "get our colours and fonts", "use our logo", or before the first launch film for a new brand.

freudianflips/breakout-launch-videos

Make or change a product launch film in your own brand. A film is a hook (the story before the cut) on a shared body (the name punching in, hero words cut on the voice, your product at close range, the logo lockup in near silence), built in HyperFrames from brand/brand.json and film/film.json. Holds the core, the template workflow (build, render and mix 1 hook or all of them), the generation recipes, the checks and the lessons. Use for "launch video", "launch film", "product video", "promo", "reveal", "announcement video", "new hook", "change the body", "render all launch videos".

freudianflips/breakout-launch-videos

Tool manual under launch-film (mixdown.py for buses, ducking, loudness and the brickwall). Mix and master the audio of a launch film. Music, SFX and voice-over on separate buses, the bed ducked by the voice's real envelope with a presence carve, the music dropped out for the logo, glue compression, a true-peak-safe limiter, -14 LUFS and -1 dBTP, then laid onto the rendered picture. Use after the score, the sound effects and the voice exist, for "mix", "master", "loudness", "LUFS", "ducking", "mux the audio".

freudianflips/breakout-launch-videos

Tool manual under launch-film (the motion language for launch films built in HyperFrames). Ease tokens, camera doctrine, continuity at every cut, type and cursor craft, and brand-neutral signature moves (logo build, world pass, branch and pulse, waterfall fill, identity morph, macro pull-back, rise) that replace HyperFrames' stock look. Says which HyperFrames blueprints and transitions to keep, transform or avoid. Use whenever a launch film's picture is designed or animated, for "animation", "motion design", "transitions", "camera move", "make it feel like Apple", "less template".

freudianflips/breakout-launch-videos

Tool manual under launch-film (measuring a render and studying a reference film). Review a rendered launch film like a demanding creative director, then drive the v1 to v2 loop. Measures rhythm, static runs, cuts against the beat grid, loudness, true peak and voice pace with a script, then applies an eye checklist for continuity at every cut, legibility, brand, real UI, claims and differentiation. Writes a gap list and a scorecard; nothing ships on numbers alone. Use on every render, and on a reference film to learn from it, for "review the video", "check the film", "what is wrong with this cut", "study this reference".

freudianflips/breakout-launch-videos

Tool manual under launch-film (music fitted to the edit with Higgsfield Sonilo, ACE-Step locally, or the free scratch bed). Score a launch film with music whose sections are the film's acts. Each take is measured for tempo and section energy, stretched onto the film's BPM and offset so its drop lands on the cut. Use after the storyboard's beat sheet exists, for "music", "score", "soundtrack", "background track", "the drop", "beat grid".

freudianflips/breakout-launch-videos

Tool manual under launch-film (the sound palette, the event map, the cue sheet and generated sound kits). Design the sound effects of a launch film. Builds an event map from the storyboard, the voice-over word timings and the rendered timeline, voices every event from 1 consistent palette (the free synthesised, key-locked forge, or real glass, felt, wood, marimba, clicks and air generated on Higgsfield and pitched into the film's key), and renders 1 SFX stem in a shared acoustic space. Use once picture timing and the score are fixed, for "sound effects", "SFX", "sound design", "UI sounds", "sonic logo".

freudianflips/breakout-launch-videos

Tool manual under launch-film (the brief, the story spine, the script, the beat sheet and reference studies). Turn a launch brief into a launch film's story in your own brand, with the one message, the narrative spine, the voice-over script at the right length, a beat sheet on the music's bar grid and a STORYBOARD.md where every shot names its continuity in and out, its signature move and its sound. Includes the reference study and the logo-swap differentiation test. Use first, before voice, score or motion, for "launch video script", "storyboard", "brief", "what should the launch film say", "study this reference video".

freudianflips/breakout-launch-videos

Tool manual under launch-film (backgrounds, plates, Blender glass and 3D shots, shader layers). Build the world a launch film lives in, in your own brand, with generated background plates (cost-gated), Blender for real 3D glass objects and camera moves rendered as alpha layers, WebGL shader layers for atmosphere, and the layer order, depth of field and grade that keep UI legible over living backgrounds. Use when a storyboard needs backgrounds, 3D or atmosphere, for "background", "plate", "3D object", "glass", "Blender", "shader", "atmosphere"; motion of the UI itself is launch-motion.

Skills associés