Cinematic Story Video Production Skill
This skill guides the end-to-end production of high-fidelity cinematic storytelling videos from novel chapters or narrative scripts. It combines multimodal image generation with character reference anchoring, pure studio voiceover via Gemini TTS, and fluid multi-shot camera direction with HyperFrames AI (GSAP).
Core Philosophy & Production Standards
- Strict Character Consistency: Every character must be visually anchored using a Reference Concept Sheet. Never generate story scenes from blind text prompts alone—always pass character reference image bytes and enforce the visual DNA (facial structure, hair, clothing).
- Pure Studio Voiceover (Gemini TTS): The narrator voice must be generated using
gemini-3.8-flash-tts. NEVER put system prompts, tone instructions, or metadata into the content prompt; pass strictly the pure story narration text, or the TTS model will speak instructions aloud. - No Synthetic Audio Clutter: Avoid procedural/synthetic SFX and generic background music unless explicitly requested by the user. A clean, high-clarity vocal narration gives an immersive, audiobook-meets-film experience.
- No Subtitles: Avoid burning or overlaying subtitles onto the cinematic canvas unless requested. Keep the screen clean with a 2.39:1 anamorphic letterbox to maximize visual impact.
- Dynamic Cinematic Shots: Avoid single static shots per paragraph. For high-intensity or action scenes, break the beat into rapid multi-shot cuts (e.g., Cut 1: extreme detail/reaction, Cut 2: impact/wide, Cut 3: close-up emotion).
- Strict Security (
.env): Never read, cat, print, display, or edit.env. Backend scripts load environment variables programmatically into runtime memory.
Production Pipeline (Step-by-Step)
flowchart TD
A["Chapter Text (chapter_XX.md)"] --> B["1. Scene Breakdown & Storyboard"]
B --> C["2. Character Bible & Reference Sheets"]
C --> D["3. Multimodal Image Generation (gemini-2.5-flash-image)"]
B --> E["4. Pure Studio Voiceover (gemini-3.8-flash-tts)"]
D --> F["5. HyperFrames Composition & GSAP Motion"]
E --> F
F --> G["6. Validation (npx hyperframes lint)"]
G --> H["7. Full Render (npx hyperframes render)"]
Step 1: Story Breakdown & Storyboard Construction
- Read the target novel chapter or story text (e.g.,
chapter_01.md). - Break down the chapter into 8–12 distinct cinematic beats/scenes.
- For each scene, specify:
scene_id: e.g.,scene_01,scene_02, etc.title: Short descriptive title (Thai/English).narration: The exact, pure Thai narration text for that scene.visual_prompt: Detailed cinematic prompt (16:9, lighting, camera angle, atmospheric mood).characters_present: List of character keys (e.g.,["prima", "tawan"]).
- Save the breakdown to
video_project/storyboard.json.
Step 2: Character Bible & Reference Anchoring
- Check
character_references/character_bible.mdto see if existing characters are defined. - If introducing new characters, generate a dedicated Character Reference Sheet first:
- Format: 16:9 widescreen concept sheet showing multiple angles (close-up portrait, 3/4 view, full body) on a neutral cinematic studio background.
- Save to
character_references/<character_key>_reference_sheet.png. - Document exact visual DNA in
character_references/character_bible.md(age, facial features, hairstyle, clothing style, distinctive traits).
See Character Consistency Guide for detailed reference prompt templates.
Step 3: Multimodal Image Generation
- Use
gemini-2.5-flash-imagevia the Google GenAI SDK. - For scenes featuring established characters:
- Load the character reference sheet as binary image bytes.
- Supply the reference image part alongside the text prompt in the multimodal request.
- Explicitly instruct the model: "Using the exact character from the reference image (, ): [Scene description]".
- For action / high-drama scenes:
- Generate multi-shot sub-cuts (e.g.,
scene1_cut1_tire.png,scene1_cut2_impact.png,scene1_cut3_face.png).
- Generate multi-shot sub-cuts (e.g.,
- Save images to
video_project/images/and copy them to the HyperFrames project assets (sample_hyperframe/assets/).
Helper script: scripts/generate_scene_images.py
Step 4: Pure Studio Voiceover (Gemini TTS)
- Use model
gemini-3.8-flash-tts. - Configure audio generation:
response_modalities=["AUDIO"]speech_configwith prebuilt voice (e.g.,PuckorAoede).
- CRITICAL: Pass ONLY the exact story narration text into
contents.# CORRECT: response = client.models.generate_content( model="gemini-3.8-flash-tts", contents=scene["narration"], # Pure text only! config=types.GenerateContentConfig(response_modalities=["AUDIO"], speech_config=...) ) - Convert raw output PCM/WAV to clean 44.1kHz MP3 using FFmpeg.
- Save audio files to
video_project/audio/voice/and copy tosample_hyperframe/assets/.
Helper script: scripts/generate_gemini_tts.py
Step 5: HyperFrames Assembly & Cinematic Motion
- Probe the duration of each scene's voice file using
ffprobe. - Calculate scene start times and durations:
scene_duration = voice_duration + 0.6s(0.2s pre-roll + 0.4s breathing room).
- Generate
sample_hyperframe/index.htmlwith:- Cinematic 2.39:1 widescreen letterboxing overlays (
.letterbox-top,.letterbox-bottom). - Pure image containers with zero subtitles.
- Native
<audio>elements scheduled at their exact timestamps (tl.call(() => audio.play(), ..., voice_start)). - GSAP timeline animation:
- Multi-shot cuts for action beats (display toggling with white flash transitions).
- Slow Ken Burns push-in / pan (
scale: 1.0 -> 1.08-1.12). - Smooth crossfades / cuts between scenes.
- Register the timeline:
window.__timelines = [tl].
- Cinematic 2.39:1 widescreen letterboxing overlays (
- Update
sample_hyperframe/hyperframes.jsonwith the exact total duration:{ "duration": 177.6, "fps": 30, "width": 1920, "height": 1080 }
Helper script: scripts/build_hyperframes_project.py
Detailed animation guide: HyperFrames Motion Guide
Step 6: Verification & Rendering
- Lint Check:
Ensure 0 errors, 0 warnings.cd sample_hyperframe npx hyperframes lint - Render Full Movie:
npx hyperframes render --output ../video_project/output/<chapter_name>_full_movie.mp4 - Quality Verification:
- Confirm file exists and size is reasonable (~150–250 MB for ~3 min Full HD).
- Verify audio-video synchronization and absence of black frames or audio dropouts.