Communitygithub.com

Hownameee/short-videos

Generate, revise or troubleshoot local Kokoro narration, pronunciation and scene word timing. Reuse verified unchanged clips; inspect the Python environment before setup changes. Exclude research, visuals, music composition and publishing.

Was ist short-videos?

short-videos is a Claude Code agent skill that generate, revise or troubleshoot local Kokoro narration, pronunciation and scene word timing. Reuse verified unchanged clips; inspect the Python environment before setup changes. Exclude research, visuals, music composition and publishing.

Funktioniert mit~Claude Code~Codex CLI~Cursor
npx skills add https://github.com/Hownameee/short-videos/tree/HEAD/.agent/skills/kokoro-tts

In Ihrer bevorzugten KI fragen

Öffnet einen neuen Chat, in dem dieser Agent-Skill bereits geladen ist.

Dokumentation

Kokoro narration

Use approved text and the shared generator; do not copy a project-specific TTS script or replace predicted timestamps with estimated sentence timing.

.venv-kokoro/bin/python tools/kokoro/generate_audio.py --project projects/NNN-topic

The helper supports matching American/British English voices at 24 kHz. Read setup and language reference only for installation, dependency failures or changing language; inspect interpreter/pyvenv.cfg first. Ask before adding dependencies. Other languages need a supported provider or an audition; never promise native Vietnamese from English phonetics.

Revisions and reuse

Keep stable unique scene IDs. The generator checks text/pronunciation, model, voice, language, speed, sample rate, package and generator/caption code version, WAV hash/duration and word/caption coverage. Only invalid scenes are synthesized; unchanged WAVs and timestamps survive bit-for-bit. Old verified manifests can be adopted once. A fully unchanged run leaves the manifest unchanged too. Generation stages all results before swapping the voice directory; failure preserves the previous set. Do not run simultaneous writers for one project.

Use --force to deliberately regenerate all clips, including after changing locally cached model weights under the same model identifier (weight bytes are not part of the cache key). New model/voice/speed invalidates every affected clip. A music-only revision does not call TTS.

Normalize acronyms/numbers/URLs with the script's pronunciation map. Preserve sourceText versus spoken text; captions follow spoken text unless separately aligned. Say each intended phrase once. Punctuation is not a timing guarantee.

Measure actual WAV durations before scene frames. Shorten, expand or split dense scenes before increasing speed; never trim speech to fit a guessed duration. Keep raw audio and create separate padded/resampled derivatives if needed.

Validate

Run workflow preflight after generation. Require complete ordered in-range words, phrase boundaries, correct sample rate, nonempty decodable audio and measured duration. Timings include actual preceding chunk lengths and attach punctuation. Review missing/repeated/clipped words, pronunciation, pauses, peaks and tail room. Listen when possible; predicted alignment is not independent transcription. Inspect three successive rendered highlights and a later scene/chunk, following defaults. Record unperformed listening.

Individual skills in this repo

This repo contains 11 individual skills — each has its own dedicated page.

Hownameee/short-videos

Router for all Remotion skills

Hownameee/short-videos

Transcribing, displaying and animating captions

Hownameee/short-videos

Structure Remotion markup for interactivity

Hownameee/short-videos

Remotion Map animation knowledge

Hownameee/short-videos

Content, animation and effects best practices

Hownameee/short-videos

Upgrade Remotion, and related packages

Hownameee/short-videos

>- Run Phase 1 of the shared short-video workflow: turn a raw topic or prompt into a sourced research pack and 2–4 distinct video directions, each with a defensible tension hook, promise, payoff, beats, visual approach, duration range, and risks. Use whenever the user asks for video ideas, angles, research, facts, hooks, topic analysis, or choices before scriptwriting. Stop for the user's direction choice and do not silently continue into script or production.

Hownameee/short-videos

Produce an approved short video or correct an existing video's background sound, emotion, captions, mix or export. Generate only changed narration, render and review a current preview, run technical and editorial QC, then export. Exclude script-only and general audio-advice requests.

Hownameee/short-videos

>- Run Phase 2 of the shared short-video workflow: convert an explicitly selected direction and its research into a reviewable beat sheet, narration, flexible scene timing, on-screen text, visual plan, captions, music and SFX cues, pronunciation notes, payoff, takeaway, and contextual CTA. Use whenever a user asks to write, revise, preview, or approve a short-video script. Stop for script approval and do not generate TTS or render video yet.

Hownameee/short-videos

>- Design and review a coherent visual system for short-form video before and during Remotion implementation. Use for video art direction, scene UI, typography, layout, color, diagrams, captions, motion, transitions, platform safe areas, or visual-quality review. Use Remotion skills separately for API, timeline, media, Studio, and rendering behavior; do not use this skill for narration-only or research-only work.

Hownameee/short-videos

Coordinate prompt-to-short-video discovery, scripting and production. Resume existing projects and route soundtrack or mix corrections directly to production. Preserve direction choice and script approval; exclude general audio advice.

Verwandte Skills