Communitygithub.com

nine-programmer/CineStory-Studio

>- Produce cinematic storytelling videos from fiction chapters or narrative scripts using Gemini multimodal image generation, Gemini TTS studio voiceover, and HyperFrames AI GSAP animations with strict character consistency, fluid camera cuts, and no subtitle/audio clutter. Use whenever the user asks to turn a story, novel chapter, or script into a video, create cinematic animations with HyperFrames, or produce audio-visual storytelling content.

CineStory-Studio とは?

CineStory-Studio is a Gemini CLI agent skill that >- Produce cinematic storytelling videos from fiction chapters or narrative scripts using Gemini multimodal image generation, Gemini TTS studio voiceover, and HyperFrames AI GSAP animations with strict character consistency, fluid camera cuts, and no subtitle/audio clutter. Use whenever the user asks to turn a story, novel chapter, or script into a video, create cinematic animations with HyperFrames, or produce audio-visual storytelling content.

対応~Claude Code~Codex CLI~Cursor✓Gemini CLI
npx skills add https://github.com/nine-programmer/CineStory-Studio/tree/HEAD/.agents/skills/cinematic-story-video

お気に入りのAIに質問する

このエージェントスキルを事前に読み込んだ状態で新しいチャットを開きます。

ドキュメント

Cinematic Story Video Production Skill

This skill guides the end-to-end production of high-fidelity cinematic storytelling videos from novel chapters or narrative scripts. It combines multimodal image generation with character reference anchoring, pure studio voiceover via Gemini TTS, and fluid multi-shot camera direction with HyperFrames AI (GSAP).


Core Philosophy & Production Standards

  1. Strict Character Consistency: Every character must be visually anchored using a Reference Concept Sheet. Never generate story scenes from blind text prompts alone—always pass character reference image bytes and enforce the visual DNA (facial structure, hair, clothing).
  2. Pure Studio Voiceover (Gemini TTS): The narrator voice must be generated using gemini-3.8-flash-tts. NEVER put system prompts, tone instructions, or metadata into the content prompt; pass strictly the pure story narration text, or the TTS model will speak instructions aloud.
  3. No Synthetic Audio Clutter: Avoid procedural/synthetic SFX and generic background music unless explicitly requested by the user. A clean, high-clarity vocal narration gives an immersive, audiobook-meets-film experience.
  4. No Subtitles: Avoid burning or overlaying subtitles onto the cinematic canvas unless requested. Keep the screen clean with a 2.39:1 anamorphic letterbox to maximize visual impact.
  5. Dynamic Cinematic Shots: Avoid single static shots per paragraph. For high-intensity or action scenes, break the beat into rapid multi-shot cuts (e.g., Cut 1: extreme detail/reaction, Cut 2: impact/wide, Cut 3: close-up emotion).
  6. Strict Security (.env): Never read, cat, print, display, or edit .env. Backend scripts load environment variables programmatically into runtime memory.

Production Pipeline (Step-by-Step)

flowchart TD
    A["Chapter Text (chapter_XX.md)"] --> B["1. Scene Breakdown & Storyboard"]
    B --> C["2. Character Bible & Reference Sheets"]
    C --> D["3. Multimodal Image Generation (gemini-2.5-flash-image)"]
    B --> E["4. Pure Studio Voiceover (gemini-3.8-flash-tts)"]
    D --> F["5. HyperFrames Composition & GSAP Motion"]
    E --> F
    F --> G["6. Validation (npx hyperframes lint)"]
    G --> H["7. Full Render (npx hyperframes render)"]

Step 1: Story Breakdown & Storyboard Construction

  1. Read the target novel chapter or story text (e.g., chapter_01.md).
  2. Break down the chapter into 8–12 distinct cinematic beats/scenes.
  3. For each scene, specify:
    • scene_id: e.g., scene_01, scene_02, etc.
    • title: Short descriptive title (Thai/English).
    • narration: The exact, pure Thai narration text for that scene.
    • visual_prompt: Detailed cinematic prompt (16:9, lighting, camera angle, atmospheric mood).
    • characters_present: List of character keys (e.g., ["prima", "tawan"]).
  4. Save the breakdown to video_project/storyboard.json.

Step 2: Character Bible & Reference Anchoring

  1. Check character_references/character_bible.md to see if existing characters are defined.
  2. If introducing new characters, generate a dedicated Character Reference Sheet first:
    • Format: 16:9 widescreen concept sheet showing multiple angles (close-up portrait, 3/4 view, full body) on a neutral cinematic studio background.
    • Save to character_references/<character_key>_reference_sheet.png.
    • Document exact visual DNA in character_references/character_bible.md (age, facial features, hairstyle, clothing style, distinctive traits).

See Character Consistency Guide for detailed reference prompt templates.


Step 3: Multimodal Image Generation

  1. Use gemini-2.5-flash-image via the Google GenAI SDK.
  2. For scenes featuring established characters:
    • Load the character reference sheet as binary image bytes.
    • Supply the reference image part alongside the text prompt in the multimodal request.
    • Explicitly instruct the model: "Using the exact character from the reference image (, ): [Scene description]".
  3. For action / high-drama scenes:
    • Generate multi-shot sub-cuts (e.g., scene1_cut1_tire.png, scene1_cut2_impact.png, scene1_cut3_face.png).
  4. Save images to video_project/images/ and copy them to the HyperFrames project assets (sample_hyperframe/assets/).

Helper script: scripts/generate_scene_images.py


Step 4: Pure Studio Voiceover (Gemini TTS)

  1. Use model gemini-3.8-flash-tts.
  2. Configure audio generation:
    • response_modalities=["AUDIO"]
    • speech_config with prebuilt voice (e.g., Puck or Aoede).
  3. CRITICAL: Pass ONLY the exact story narration text into contents.
    # CORRECT:
    response = client.models.generate_content(
        model="gemini-3.8-flash-tts",
        contents=scene["narration"],  # Pure text only!
        config=types.GenerateContentConfig(response_modalities=["AUDIO"], speech_config=...)
    )
    
  4. Convert raw output PCM/WAV to clean 44.1kHz MP3 using FFmpeg.
  5. Save audio files to video_project/audio/voice/ and copy to sample_hyperframe/assets/.

Helper script: scripts/generate_gemini_tts.py


Step 5: HyperFrames Assembly & Cinematic Motion

  1. Probe the duration of each scene's voice file using ffprobe.
  2. Calculate scene start times and durations:
    • scene_duration = voice_duration + 0.6s (0.2s pre-roll + 0.4s breathing room).
  3. Generate sample_hyperframe/index.html with:
    • Cinematic 2.39:1 widescreen letterboxing overlays (.letterbox-top, .letterbox-bottom).
    • Pure image containers with zero subtitles.
    • Native <audio> elements scheduled at their exact timestamps (tl.call(() => audio.play(), ..., voice_start)).
    • GSAP timeline animation:
      • Multi-shot cuts for action beats (display toggling with white flash transitions).
      • Slow Ken Burns push-in / pan (scale: 1.0 -> 1.08-1.12).
      • Smooth crossfades / cuts between scenes.
    • Register the timeline: window.__timelines = [tl].
  4. Update sample_hyperframe/hyperframes.json with the exact total duration:
    { "duration": 177.6, "fps": 30, "width": 1920, "height": 1080 }
    

Helper script: scripts/build_hyperframes_project.py Detailed animation guide: HyperFrames Motion Guide


Step 6: Verification & Rendering

  1. Lint Check:
    cd sample_hyperframe
    npx hyperframes lint
    
    Ensure 0 errors, 0 warnings.
  2. Render Full Movie:
    npx hyperframes render --output ../video_project/output/<chapter_name>_full_movie.mp4
    
  3. Quality Verification:
    • Confirm file exists and size is reasonable (~150–250 MB for ~3 min Full HD).
    • Verify audio-video synchronization and absence of black frames or audio dropouts.

関連スキル