Communitygithub.com

Andersonlimahw/lemon-ai-hub

Synchronize narration, captions, music, and SFX with picture — narration writing, script lock, TTS or recorded voiceover, word-level timestamps, caption blocks, shot windows derived from the voice track, dwell budgets, ducking, and loudness targets. Use for narrated explainers, voiceover changes, subtitles, or any video where timing must follow audio.

What is lemon-ai-hub?

lemon-ai-hub is a Codex agent skill that synchronize narration, captions, music, and SFX with picture — narration writing, script lock, TTS or recorded voiceover, word-level timestamps, caption blocks, shot windows derived from the voice track, dwell budgets, ducking, and loudness targets. Use for narrated explainers, voiceover changes, subtitles, or any video where timing must follow audio.

Works with~Claude Code✓Codex CLI~Cursor
npx skills add https://github.com/Andersonlimahw/lemon-ai-hub/tree/HEAD/plugins/motion-movie-expert/skills/voice-video-sync

Ask in your favorite AI

Open a new chat with this agent skill pre-loaded.

Documentation

Voice ↔ Video Sync

Rule zero: the voice track drives the clock. Picture is timed to measured audio, never to guessed durations.

1. Write the narration

  • Open on the viewer's situation, not a definition. One storyline; cause before effect.
  • Convert numbers to felt scale ("the size of a phone book", not "3.2 MB").
  • Metaphors must carry weight; the narrator has a point of view.
  • Do not describe what the picture already shows; say what the picture cannot.
  • End by calling back to the opening. Chapter boundaries hand off: last line of chapter N sets up the first line of N+1.
  • Budget by speaking rate: English ≈ 125–150 words/min (TTS voices vary — measure one sample); Portuguese ≈ 140–160 words/min; Chinese ≈ 4–5.5 characters/s.

Format for the script file (one paragraph = one shot, | splits caption blocks):

# CHAPTER 1 Why maps lie
Every map you have seen | was drawn by someone | with a reason.

Mercator wanted sailors | to hold a straight course.
## gap 0.6

## gap <seconds> asks for a pause of at least that long after the paragraph (dwell for the last visual). Produce it in the audio — SSML <break time="600ms"/> for TTS, or inserted silence when assembling per-paragraph files. The timeline tool never shifts timings; it verifies the pause.

2. Lock the script

Show the full script with word count and estimated duration and get an explicit "approved". After lock, every timing is derived from audio; changing a word re-times all shots after it.

3. Produce the voice

  • TTS (default when no recording): ask the user for a preferred engine/voice. Sending text to a cloud TTS (ElevenLabs, Azure, OpenAI, Google, edge-tts) shares the script with that provider — confirm first. Local options (Kokoro, Piper) keep it on the machine.
  • Recorded VO: accept WAV 48 kHz. Trim silence, keep breaths natural, normalize later in the mix.
  • Keep one file per paragraph or one master file plus timestamps; both work with the timeline tool.

4. Get timestamps

Use word- or segment-level timestamps from the TTS API response, or force-align the audio with the script (WhisperX, whisper-timestamped, stable-ts, or Whisper with word_timestamps=True). Normalize them to this JSON:

{"segments": [{"text": "Every map you have seen", "start": 0.00, "end": 1.42,
  "words": [{"word": "Every", "start": 0.00, "end": 0.21}]}]}

5. Build the timeline

python3 <plugin-root>/scripts/build_timeline.py words.json --script script.txt --fps 30 \
  --out-dir timeline/ --max-chars 42 --min-dwell 1.0

Produces timeline.json (shots with start/end frames, caption blocks, chapter starts), captions.srt, captions.vtt, and warnings for: script words missing from the audio (or audio words missing from the script), captions too long or too fast to read, shots whose last caption leaves less than --min-dwell seconds before the next shot (add a pause in the audio or merge), requested ## gap pauses the audio does not have, and silences longer than 3 s. Words a forced aligner left untimed (common for numbers in WhisperX) are interpolated between their neighbours. Feed timeline.json to the composition; never retype frame numbers by hand.

6. Caption rules

  • ≤42 characters per line, ≤2 lines, ≥1.0 s on screen, ≈17 characters/s reading speed maximum.
  • Break at phrase boundaries, never between article and noun.
  • Keep captions out of the lower platform UI zone in 9:16 and away from on-screen UI text.
  • Burned-in captions for social; sidecar SRT/VTT for YouTube and web players.

7. Mix

LayerTarget
Integrated loudness (final)−14 LUFS for social/web (−16 acceptable); −23 LUFS for broadcast
True peak≤ −1 dBTP
Music under VOduck 8–12 dB while speech is present, 150–300 ms attack/release
SFX≈10 dB under music; on the causing frame, not after

Measure with <plugin-root>/scripts/qc.sh loudness final.mp4 (add -23 for broadcast). When HyperFrames audio skills are installed, hyperframes-audio provides ducking ("voiceover carve"), EQ, and automation inside the composition; otherwise use FFmpeg sidechaincompress and loudnorm (two-pass).

8. Preview checkpoint

Render the first ≤30 s with voice, captions, and music and ask: "Voice, pace, caption size, and rhythm — OK?" Changing pace after this point means re-running TTS and re-timing everything.

Individual skills in this repo

This repo contains 8 individual skills — each has its own dedicated page.

Andersonlimahw/lemon-ai-hub

Create and edit motion videos end to end — product motion films, motion graphics, narrated explainers, captioned clips, UI animation recreations, and recuts. Orchestrates brief, storyboard, choreography, build (HyperFrames or any deterministic renderer), voice/caption sync, and render QC. Use when the user asks to make, animate, storyboard, narrate, caption, edit, or render a video or motion piece; skip for static images and web UI work with no video deliverable.

Andersonlimahw/lemon-ai-hub

Build and render a deterministic video composition — HyperFrames (HTML + GSAP) by default, or a custom seek(t) + Playwright + FFmpeg pipeline. Covers project setup, composition contract, timeline rules, creative direction for frames, lint/check/snapshot, preview, and local render. Use when implementing an approved storyboard as code.

Andersonlimahw/lemon-ai-hub

Motion law for multi-scene videos — seam continuity (axis, direction, speed, phase), carriers, causal motion, springs and timing, camera, transitions, stillness, and loop design. Use when planning or fixing how scenes connect and move, before and during any animation build.

Andersonlimahw/lemon-ai-hub

Short design-led motion graphics where motion is the message — kinetic typography, stat count-up, chart/data-viz reveal, logo sting, lower third or callout, animated map, tweet/news/headline card, webpage or UI walkthrough, and image-plus-data fusion. Usually under 10 s, up to about 30 s, no narration; renders to MP4 or a transparent overlay. Use when the user wants a short motion graphic, animated stat/chart/logo/map/card, or overlay; when the upstream HyperFrames `motion-graphics` skill is installed, use it for exact CLI flags and registry blocks.

Andersonlimahw/lemon-ai-hub

Turn a video idea, product, or script into an approved master prompt and pre-implementation storyboard — intake, concept, 8–12 state sequence, beat grid, transition map, camera plan, audio cue plan, and render architecture. Use before building any new video, or when a video brief is vague.

Andersonlimahw/lemon-ai-hub

Motion tokens, spring presets, and React/Next.js (motion/react) interaction patterns for product UI that appears in a video — rebuilding real product screens faithfully, animating them as seekable scenes, and keeping the live product's motion consistent with the film. Use when a motion video shows or recreates product UI; for web UI work with no video deliverable use design-expert instead.

Andersonlimahw/lemon-ai-hub

Edit existing video and verify any render before delivery — recut, trim, reframe to other aspect ratios, burn in or attach captions, swap or remix audio, loop, and export; plus the QC gates (probe, contact sheet, stills per beat, loudness, loop seam, legibility, facts, PII). Use for editing footage or for the final check of any motion video.

Andersonlimahw/lemon-ai-hub

Agent skill at plugins/video/SKILL.md

Related Skills