Communitygithub.com

cosmix/product-video-kit

Turn a written script into a spoken WAV with Gemini TTS (gemini-3.8-flash-tts) using a prebuilt voice and an optional style direction. Use when the user asks for narration, a voiceover, text-to-speech, an audio version of a text, or a spoken sample of a script. Covers choosing a voice, preparing the transcript, and writing style directions.

product-video-kit 是什么?

product-video-kit is a Claude Code agent skill that turn a written script into a spoken WAV with Gemini TTS (gemini-3.8-flash-tts) using a prebuilt voice and an optional style direction. Use when the user asks for narration, a voiceover, text-to-speech, an audio version of a text, or a spoken sample of a script. Covers choosing a voice, preparing the transcript, and writing style directions.

兼容平台✓Claude Code~Codex CLI~Cursor✓Gemini CLI
npx skills add https://github.com/cosmix/product-video-kit/tree/HEAD/claude-skills/narration

在你喜欢的 AI 中提问

打开一个已预加载此 Agent Skill 的新对话。

文档

Narration

tts.py (next to this file) sends a script to gemini-3.8-flash-tts and writes the returned WAV (24 kHz, mono, 16-bit PCM).

When to use

  • The user wants audio from text: a voiceover, narration, a read-aloud of a document, a demo clip, a spoken announcement.
  • One narrator per file. The API supports two-speaker dialogue, but this script does not; for a dialogue, narrate each speaker's lines separately or extend the script.

Not for: transcription (speech to text), music, or sound effects.

Running it

Needs GEMINI_API_KEY in the environment. uv installs google-genai from the script header on first run.

SKILL=~/.claude/skills/narration
uv run $SKILL/tts.py script.txt --voice Charon --style "calm and warm, speaking slowly" -o intro.wav
uv run $SKILL/tts.py script.txt                  # voice Kore, no style, writes script.wav
echo "Hello there." | uv run $SKILL/tts.py - -o hello.wav
uv run $SKILL/tts.py --list-voices
FlagMeaning
SCRIPTTranscript file, or - for stdin
--voicePrebuilt voice name (case-insensitive) or a voicekey_... id; default Kore
--styleTurn-level delivery direction; default none
-oOutput path; default is the script path with .wav
--modelOverride the model, e.g. gemini-3.8-flash-lite-tts

Play the result to check it (aplay out.wav on Linux, afplay out.wav on macOS, or ffplay -autoexit out.wav), and tell the user the file path.

Choosing a voice

Pick the voice for the persona. Age, gender and accent come from the voice, never from --style.

VoiceTraitVoiceTraitVoiceTrait
ZephyrBrightPuckUpbeatCharonInformative
KoreFirmFenrirExcitableLedaYouthful
OrusFirmAoedeBreezyCallirrhoeEasy-going
AutonoeBrightEnceladusBreathyIapetusClear
UmbrielEasy-goingAlgiebaSmoothDespinaSmooth
ErinomeClearAlgenibGravellyRasalgethiInformative
LaomedeiaUpbeatAchernarSoftAlnilamFirm
SchedarEvenGacruxMaturePulcherrimaForward
AchirdFriendlyZubenelgenubiCasualVindemiatrixGentle
SadachbiaLivelySadaltagerKnowledgeableSulafatWarm

Starting points: explainers and documentation Charon, Sadaltager, Iapetus; product demos Puck, Achird; bedtime or meditation Sulafat, Vindemiatrix, Achernar; announcements Kore, Alnilam; storytelling Gacrux, Algenib.

Writing the script

The model reads the file as a verbatim transcript: every word in it is spoken. Never put directions ("read this cheerfully") in the script; they go in --style.

  • Pacing: punctuation does most of it. Commas and dashes give short breaks, ellipses trail off, paragraph breaks give longer ones. Add <short pause> or <long pause> where timing matters.
  • Emphasis: capitalise the word and back it with punctuation: This is a VERY important point!
  • Momentary sounds: inline tags at the exact point they happen: <breath>, <sigh>, <laugh>, <gasp>, <cough>, <throat-clearing>.
  • Written forms: spell out what should be said aloud. Write twenty twenty-six if 2026 might read wrong, expand abbreviations the listener would not hear as letters, and drop markdown, URLs and code blocks.
  • Language: write in the target language; the model supports over 130 and follows the text.
Welcome back. <short pause> Today we're looking at something small... but it matters.

Every build you run — every single one — starts here. <breath> And most people never look at it.

Writing style directions

--style sets sustained delivery for the whole file. Start with no style at all: most scripts need none, and the voice plus punctuation carries them. Add a style only when a listen shows the delivery is wrong.

A good direction is a few words naming tone, energy and pace:

GoodWhy
calm and relaxedTone and energy, nothing else
cheerful and friendlyTwo compatible attributes
speaking slowly, warmPace plus tone
whispersA single sustained delivery
sarcasticAn attitude the model can hold
speaking rapidly, out of breathPace plus physical state

Avoid:

BadProblem
a 60-year-old British manAge, gender and accent belong to the voice choice
keep the voice steady throughoutThe docs say not to give instructions to hold the voice steady
start excited, then get sad at the endStyle is one sustained delivery; split the script into separate files for a change of mood
pause after the second sentencePoint-in-time events go in the script as tags
A paragraph of scene, backstory and audienceLong style blocks dilute the direction; keep it to a short phrase

When iterating, change one thing at a time (voice, then style, then script punctuation) and listen after each.

Long scripts

The docs give no input length limit. If the audio stops early or drifts, split the script at paragraph boundaries, narrate each part with the same voice and style, and join them:

sox part1.wav part2.wav part3.wav full.wav

Individual skills in this repo

This repo contains 1 individual skill — each has its own dedicated page.

相关技能