Launch voice
Tool manual under
launch-film. Start atskills/launch-film/SKILL.md: it holds the core, the hook-plus-body template and the lessons. Where this page differs,launch-filmwins.
In a modern launch film the voice leads and the picture illustrates it word by word: when the voice says "tagged", something gets tagged. That only works when every word has a timestamp.
Providers
- Default: Higgsfield on your own account.
text2speech_v2(engineselevenlabs,minimax,seed_speech, and others per voice) for casting across about 100 preset voices, about 0.3 credits a line. Seed Audio for a narrator you steer with a prompt and audio references.scripts/hf_voice.pywraps the CLI. - Free scratch voice:
scripts/scratch_voice.pymakes placeholder takes with macOSsay(orespeak-ngon Linux). Build and time the whole film on it, label every review "scratch voice", then swap the real takes in. - Free local clone:
scripts/clone_local.py(Chatterbox, MIT) clones your approved narrator from 8 to 10 s of their lines. A stand-in, not a master. Only clone a voice you have the rights to. - Review voices inside a full film, never as loose clips. Keep a list of rejected voices so they are never proposed again.
Casting
python3 skills/launch-voice/scripts/hf_voice.py cast "<the film's opening line>" film/vo/auditionspeaks the line with every preset voice (about 34 credits for about 110 voices), checks each take with whisper, and measures pace and pitch movement.ranking.mdlists exact takes first, nearest 2.5 words a second and about 2 semitones of movement (conversational, not announced).--only Name,Namerecasts a few.- Take the top 6 or so across genders into the full script (
hf_voice.py lines), then fit them (below). A voice that needs more than 12 % speed-up to fit is too slow for the script: drop it or cut words. - Build a full film per surviving voice and let the owner pick. ElevenLabs takes varied least in pace and pitch in testing, so it is the default engine.
Length and pace
- Target 2.2 to 2.7 spoken words a second. A strong reference launch ran 2.4 (151 words in 63 s).
- A 20 s film carries about 45 words; 30 s about 70; 60 s about 140. Cut words before you speed the voice up.
- One line per storyboard frame, 6 to 20 words, phrased so the key noun or verb falls on a beat.
Procedure for the film template
- The script lives in
film/film.jsonlines. Record 1 file per line intofilm/vo/body/line-<id>.wav: real takes withhf_voice.py lines(write the lines as{"lines": [{"line": "02", "text": "..."}]}), or scratch takes withpython3 skills/launch-voice/scripts/scratch_voice.py film/film.json(add--hook film/hooks/<name>to voice a hook line too). - Run
python3 skills/launch-voice/scripts/prep_lines.py film/film.json(--align forceto insist on forced alignment; withuvinstalled it fetches torch on the first run, about 2 GB). It writesfilm/vo/body/lines.jsonwith each line's duration and word times: forced alignment when torch is available, else whisper.cpp, else an even estimate flaggedestimate. film/build.pyplaces the lines on the film timeline and every picture cue follows its word. A new take is 1 prep run plus a rebuild.
Procedure for a free-form film
vo/script-lines.json({"lines": [{"line": "01", "text": "..."}]}), then one file per line:hf_voice.py lines vo/script-lines.json vo/takes/<voice> --voice <Name>. Text, voice, engine and job ids land inlines.json.- Fit:
python3 skills/launch-voice/scripts/fit_lines.py vo/placement.plan.json vo/takes/<voice> vo/versions/<voice>. The plan gives each line its earliest start, its breath gap and, where a frame must end,end_by. Pauses over 0.25 s inside a line shrink to 0.22 s; a run of lines that still overflows is sped up as one, capped at 12 %. Then stitch:python3 skills/launch-voice/scripts/stitch.py vo/versions/<voice>/placement.json vo/versions/<voice> vo/versions/<voice>/vo.wav. - Word times per line. For anything a visual must land on, use forced alignment:
uv run -q --with torch --with torchaudio --with soundfile python skills/launch-voice/scripts/force_align.py line.wav "the exact words". For a quick read,align_lines.pyruns whisper.cpp line by line. Anchors resolve in speaking order (each the first match after the previous one), so any voice re-times the film. A single whisper pass over the whole film drifts near line ends; per-line passes do not. - Hand the word times to
launch-motion(anchor each build to its word) andlaunch-sound(events follow words).
Vintage narrator
For an archival or documentary opening: python3 skills/launch-voice/scripts/vintage.py IN.wav OUT.wav --crackle crackle.wav band-limits to 150 Hz to 5.8 kHz, adds tape saturation, a slow wow and a real vinyl crackle bed (generate 1 with Seed Audio). Seed Audio can also speak a line in a vintage narrator voice directly ("a warm, deep 1960s documentary narrator on an old vinyl record ... exactly these words: ..."); it reads slowly and may rephrase ("gotta" became "got to"), so check with whisper.
Sync rules
- A key word lands with its visual within 2 frames. Start builds on the word's first frame; if you let a visual lead its word, keep it under 4 frames and do it everywhere.
- Let the voice breathe at scene changes: 0.25 to 0.5 s between lines.
- The last line ends at least 1 s before the logo; the logo gets its own near silence (
launch-mixdrop-out).
When the voice provider is unavailable
Build against the scratch voice, label every review as scratch, and keep the picture anchored to the aligned words, so swapping in the real voice is 1 prep run plus a render. Never switch voice providers or accounts without the owner's go.
Fixes
- A line runs long by under 5 %:
ffmpeg -af atempo=1.04. More than that: rewrite the line. - A word is mispronounced: respell it phonetically in the line's
sayfield (forhf_voice.py), never in the script. - Whisper mishears a soft ending: use forced alignment or shorten the anchor prefix; do not re-record.
- Digits: spell them out in the text you align ("twenty four", not "24").
- The whisper base model mishears product names. Check final words with the small model or larger (
WHISPER_MODEL).