Communitygithub.com

delorenj/skills

Generate professional voiceovers using ElevenLabs AI. Use when the user needs to create voiceovers for videos, audio narration, or text-to-speech content. Supports multiple voices with character presets (narrator, salesperson, expert) for natural delivery. Includes single scene regeneration for fine-tuning.

skills とは?

skills is a Claude Code agent skill that generate professional voiceovers using ElevenLabs AI. Use when the user needs to create voiceovers for videos, audio narration, or text-to-speech content. Supports multiple voices with character presets (narrator, salesperson, expert) for natural delivery. Includes single scene regeneration for fine-tuning.

対応~Claude Code~Codex CLI~Cursor
npx skills add https://github.com/delorenj/skills/tree/HEAD/elevenlabs-remotion

お気に入りのAIに質問する

このエージェントスキルを事前に読み込んだ状態で新しいチャットを開きます。

ドキュメント

ElevenLabs Voiceover Generation

Generate professional AI voiceovers for Remotion videos using ElevenLabs API.

Prerequisites

  • ELEVEN_API_KEY in .env.local

Quick Start

# Generate voiceover from text
node ~/.agents/skills/elevenlabs-remotion/generate.js --text "Your text here" --output public/audio/voiceover.mp3

# Generate with narrator style (more natural)
node ~/.agents/skills/elevenlabs-remotion/generate.js --text "Your text" --character narrator --output voiceover.mp3

# Generate scenes with request stitching
node ~/.agents/skills/elevenlabs-remotion/generate.js --scenes remotion/scenes.json --output-dir public/audio/project/

# Regenerate a single scene
node ~/.agents/skills/elevenlabs-remotion/generate.js --scenes scenes.json --scene scene2 --new-text "Updated text"

# List available voices and character presets
node ~/.agents/skills/elevenlabs-remotion/generate.js --list-voices
node ~/.agents/skills/elevenlabs-remotion/generate.js --list-characters

Character Presets

Use character presets for more natural voiceovers instead of literal screen text reading:

CharacterDescriptionBest For
literalReads text exactly as writtenScreen text, quotes
narratorProfessional storyteller, smooth, engagingExplainers, documentaries
salespersonEnthusiastic, persuasive, energeticMarketing, ads
expertAuthoritative, confident, knowledgeableLegal content, tutorials
conversationalCasual, friendly, naturalSocial media, casual content
dramaticIntense, emotional, impactfulHooks, problem statements
calmSoothing, reassuring, gentleTrust-building, conclusions
# Use narrator style globally
node ~/.agents/skills/elevenlabs-remotion/generate.js --scenes scenes.json --character narrator --output-dir public/audio/

# Or set per-scene in scenes.json
{
  "scenes": [
    { "id": "scene1", "text": "Problem statement", "character": "dramatic" },
    { "id": "scene2", "text": "Solution", "character": "calm" }
  ]
}

Scene-Based Generation with Request Stitching

Generate multiple scenes with consistent prosody using ElevenLabs request stitching:

scenes.json Format

{
  "name": "product-demo",
  "voice": "George",
  "character": "narrator",
  "scenes": [
    {
      "id": "scene1",
      "text": "Generic text-to-speech sounds robotic. Your brand deserves better.",
      "duration": 4.5,
      "character": "dramatic"
    },
    {
      "id": "scene2",
      "text": "With voice cloning, you can use your own voice for unlimited content.",
      "duration": 5.5
    },
    {
      "id": "scene3",
      "text": "Record a short sample. Clone it. Create professional voiceovers in minutes.",
      "duration": 6,
      "delay": 0.3
    }
  ]
}

Generate All Scenes

node ~/.agents/skills/elevenlabs-remotion/generate.js \
  --scenes remotion/product-demo-scenes.json \
  --output-dir public/audio/product-demo/

This creates:

  • product-demo-scene1.mp3 through sceneN.mp3
  • product-demo-combined.mp3 (all scenes stitched)
  • product-demo-info.json (metadata with durations)

Single Scene Regeneration

If a scene starts too early, has wrong timing, or needs different text:

# Regenerate scene2 with new text
node ~/.agents/skills/elevenlabs-remotion/generate.js \
  --scenes remotion/scenes.json \
  --scene scene2 \
  --new-text "Updated scene 2 text" \
  --output-dir public/audio/project/

# Regenerate scene3 with different character
node ~/.agents/skills/elevenlabs-remotion/generate.js \
  --scenes remotion/scenes.json \
  --scene scene3 \
  --character salesperson \
  --output-dir public/audio/project/

# Just regenerate (same text, same character)
node ~/.agents/skills/elevenlabs-remotion/generate.js \
  --scenes remotion/scenes.json \
  --scene scene1 \
  --output-dir public/audio/project/

# Embed a thumbnail into an MP4 video
node ~/.agents/skills/elevenlabs-remotion/generate.js \
  --embed-thumbnail public/videos/my-video.mp4 \
  --thumbnail public/videos/my-thumbnail.png \
  --output public/videos/my-video-with-thumb.mp4

The tool automatically:

  • Uses request stitching from previous scenes for consistent prosody
  • Updates the info.json file with new metadata
  • Updates scenes.json if --new-text is provided

Thumbnail Embedding

Embed a thumbnail image into MP4 videos so platforms like Twitter, YouTube, and video players display your custom thumbnail instead of the first frame.

Embed Thumbnail into Video

# Basic usage - outputs to video-thumb.mp4
node ~/.agents/skills/elevenlabs-remotion/generate.js \
  --embed-thumbnail public/videos/promo.mp4 \
  --thumbnail public/videos/thumbnail.png

# Custom output path
node ~/.agents/skills/elevenlabs-remotion/generate.js \
  --embed-thumbnail public/videos/promo.mp4 \
  --thumbnail public/videos/thumbnail.png \
  --output public/videos/promo-final.mp4

Workflow with Remotion

# 1. Render your video
npx remotion render MyVideo public/videos/my-video.mp4

# 2. Render your thumbnail (use Still composition)
npx remotion still MyVideoThumbnail public/videos/my-thumbnail.png

# 3. Embed the thumbnail
node ~/.agents/skills/elevenlabs-remotion/generate.js \
  --embed-thumbnail public/videos/my-video.mp4 \
  --thumbnail public/videos/my-thumbnail.png \
  --output public/videos/my-video-final.mp4

Supported Formats

  • Video: MP4 (H.264/H.265)
  • Thumbnail: PNG, JPG, JPEG

The embedding uses ffmpeg's -disposition:v:1 attached_pic flag to set the thumbnail as an attached picture, which most video players and platforms recognize.

Timing Validation

The skill automatically validates timing after generation using ffprobe:

What It Checks

CheckThresholdDescription
Duration mismatch>15%Warns if actual differs from expected duration
Leading silence>200msAudio starts late (voiceover delayed)
Trailing silence>500msUnnecessary silence at end
Speaking rate2-4.5 wpsOptimal ~3 words/second

Validate Existing Audio

# Validate all scenes in a project
node ~/.agents/skills/elevenlabs-remotion/generate.js --validate public/audio/product-demo/

Output example:

🔍 Validating product-demo (6 scenes)

❌ scene1: 3.00s (expected: 4.5s)
   ❌ Audio 1.50s shorter than expected
   👍 8 words @ 3.1 words/sec
⚠️ scene2: 6.35s (expected: 5.5s)
   ⚠️ Leading silence: 235ms (may start late)
   🐢 10 words @ 1.8 words/sec
✅ scene4: 4.36s (expected: 4s)
   👍 9 words @ 2.3 words/sec

📊 Total duration: 30.80s (expected: 30.00s)

Updated info.json

After validation, the info.json includes actual measurements:

{
  "scenes": [
    {
      "id": "scene1",
      "duration": 4.5,
      "actualDuration": 3.0,
      "leadingSilence": 0.05,
      "wordsPerSecond": 3.1
    }
  ]
}

Use actualDuration in your Remotion composition for precise sync.

Options

OptionDescriptionDefault
--text, -tText to convert to speechRequired (or --file/--scenes)
--file, -fRead text from file-
--output, -oOutput file pathoutput.mp3
--output-dirOutput directory for scenespublic/audio
--voice, -vVoice name or IDGeorge
--model, -mModel IDeleven_multilingual_v2
--character, -cCharacter presetliteral
--scenesJSON file with scenes-
--sceneRegenerate single scene ID-
--new-textNew text for scene regen-
--validateValidate existing audio dir-
--skip-validationSkip auto-validationfalse
--embed-thumbnailVideo file to embed thumbnail into-
--thumbnailThumbnail image file (PNG/JPG)-
--stabilityVoice stability (0-1)varies by character
--similarityVoice similarity (0-1)varies by character
--styleStyle exaggeration (0-1)varies by character
--no-combinedSkip combined filefalse

Recommended Voices

VoiceStyleBest For
GeorgeWarm, captivating BritishNarration, explainers
AntoniProfessional, warmLegal content, tutorials
ArnoldAuthoritative, deepCorporate, serious topics
JoshFriendly, conversationalMarketing, casual content

Integration with Remotion

After generating scene voiceovers, use them in your composition:

import { Audio, Sequence, staticFile } from "remotion";

// Use individual scene audio files for precise sync
const SCENE_DURATIONS = {
  scene1: 4.5,  // From info.json
  scene2: 5.5,
  scene3: 8.0,
};

export const VideoWithVoiceover: React.FC = () => {
  const { fps } = useVideoConfig();

  const scene1Frames = Math.round(SCENE_DURATIONS.scene1 * fps);
  const scene2Frames = Math.round(SCENE_DURATIONS.scene2 * fps);

  return (
    <>
      <Sequence from={0} durationInFrames={scene1Frames}>
        <Audio src={staticFile("audio/project/project-scene1.mp3")} />
        <Scene1Visual />
      </Sequence>

      <Sequence from={scene1Frames} durationInFrames={scene2Frames}>
        <Audio src={staticFile("audio/project/project-scene2.mp3")} />
        <Scene2Visual />
      </Sequence>
    </>
  );
};

Tips for Best Results

  1. Use character presets: Don't read screen text literally - use narrator or expert for natural flow
  2. Punctuation matters: Use periods for pauses, commas for brief breaks
  3. Numbers: Write out numbers ("five hundred" not "500") for natural speech
  4. Abbreviations: Write full words ("twenty-four hours" not "24h")
  5. Scene-by-scene: Different scenes can have different characters (dramatic intro, calm CTA)
  6. Fine-tune: Use --scene to regenerate individual scenes without redoing everything
  7. Request stitching: Keeps voice consistent across all scenes

Workflow Example

# 1. Create scenes.json with your script
# 2. Generate all scenes with narrator style
node ~/.agents/skills/elevenlabs-remotion/generate.js \
  --scenes remotion/my-video-scenes.json \
  --character narrator \
  --output-dir public/audio/my-video/

# 3. Preview in Remotion, notice scene2 starts too early
# 4. Regenerate just scene2 with updated text
node ~/.agents/skills/elevenlabs-remotion/generate.js \
  --scenes remotion/my-video-scenes.json \
  --scene scene2 \
  --new-text "Slightly longer text to fill the visual timing" \
  --output-dir public/audio/my-video/

# 5. Update video composition with new duration from info.json
# 6. Repeat until timing is perfect

Individual skills in this repo

This repo contains 14 individual skills — each has its own dedicated page.

delorenj/skills

Plan and generate terminal ASCII animations/screensaver-style output (FPS, refresh rules, loop policy, low-flicker guidance), with a static poster frame and an optional local demo script.

delorenj/skills

ASCII video: convert video/audio to colored ASCII MP4/GIF.

delorenj/skills

Plan, build and quality-check a premium short commercial for a real business (roofing, property, ecommerce, B2B software) using AI-generated footage, code-built motion (HTML/GSAP), selective Three.js and an independent-critic "Gauntlet" loop. Use when asked to make a launch-style / SaaS-style video, business ad, explainer, sample reel or pitch video, or to review and improve one. Encodes motion principles distilled from 28 professional launch films, a quality bar, audio rules and a business-offer playbook.

delorenj/skills

CSS animation adapter patterns for HyperFrames. Use when authoring CSS keyframes, animation-delay based timing, animation-fill-mode, animation-play-state, or CSS-only motion that HyperFrames must seek deterministically during preview and rendering.

delorenj/skills

HyperFrames CLI dev loop — `npx hyperframes` for scaffolding (init), validation (lint, inspect), preview, render, and environment troubleshooting (doctor, browser, info, upgrade). Use when running any of these commands or troubleshooting the HyperFrames build/render environment. For asset preprocessing commands (`tts`, `transcribe`, `remove-background`), invoke the `hyperframes-media` skill instead.

delorenj/skills

Asset preprocessing for HyperFrames compositions — text-to-speech narration (Kokoro), audio/video transcription (Whisper), and background removal for transparent overlays (u2net). Use when generating voiceover from text, transcribing speech for captions, removing the background from a video or image to use as a transparent overlay, choosing a TTS voice or whisper model, or chaining these (TTS → transcribe → captions). Each command downloads its own model on first run.

delorenj/skills

Install and wire registry blocks and components into HyperFrames compositions. Use when running hyperframes add, installing a block or component, wiring an installed item into index.html, or working with hyperframes.json. Covers the add command, install locations, block sub-composition wiring, component snippet merging, and registry discovery.

delorenj/skills

Create video compositions, animations, title cards, overlays, captions, voiceovers, audio-reactive visuals, and scene transitions in HyperFrames HTML. Use when asked to build any HTML-based video content, add captions or subtitles synced to audio, generate text-to-speech narration, create audio-reactive animation (beat sync, glow, pulse driven by music), add animated text highlighting (marker sweeps, hand-drawn circles, burst lines, scribble, sketchout), or add transitions between scenes (crossfades, wipes, reveals, shader transitions). Covers composition authoring, timing, media, and the full video production workflow. For dev-loop CLI commands (init, lint, inspect, preview, render) see the hyperframes-cli skill; for asset preprocessing commands (tts, transcribe, remove-background) see the hyperframes-media skill.

delorenj/skills

Manim CE animations: 3Blue1Brown math/algo videos.

delorenj/skills

Join a Google Meet or Zoom call as a video meeting agent via PikaStreaming. Trigger: user drops a Google Meet or Zoom link, or asks to join a meeting.

delorenj/skills

Translate an existing Remotion (React-based) video composition into a HyperFrames HTML composition. Use ONLY when the user explicitly asks to port, convert, migrate, translate, or rewrite a Remotion composition as HyperFrames (e.g. "port my Remotion project to HyperFrames"). Do NOT use when (a) authoring a NEW HyperFrames composition (even if A/B-testing a Remotion video); (b) Remotion is mentioned in passing; (c) Remotion code is shared as reference, not for translation; (d) the user wants "the same video as my Remotion one" without explicitly asking to migrate the source — treat as a fresh HyperFrames build. When in doubt, default to the `hyperframes` skill. Detects unsupported patterns (useState, useEffect side effects, async calculateMetadata, third-party React component libraries, `@remotion/lambda`) and recommends the runtime interop escape hatch instead of a lossy translation.

delorenj/skills

Generate speech using the self-hosted voxxy (vox) TTS service at https://vox.delo.sh. Use when the user asks to speak, say, narrate, synthesize speech, clone a voice, create a voice, add or register a voice, pipe TTS, or control voice qualities by description (e.g. "a young woman with a cheerful voice"). Handles HTTP API usage, voice profile management, description-based voice design, cloning, MCP registration, and integration patterns for new platforms.

delorenj/skills

Capture a website and create a HyperFrames video from it. Use when: (1) a user provides a URL and wants a video, (2) someone says "capture this site", "turn this into a video", "make a promo from my site", (3) the user wants a social ad, product tour, or any video based on an existing website, (4) the user shares a link and asks for any kind of video content. Even if the user just pastes a URL — this is the skill to use.

delorenj/skills

YouTube transcripts to summaries, threads, blogs.

関連スキル