Communitygithub.com

assafkip/claude-video-editor

Build videos from stills when there is nothing to cut — narration first (ElevenLabs v3 with audio tags), then still keyframes from ANY source (AI image tools, screenshots, photos, HyperFrames-rendered HTML), animated deterministically with the animate-stills recipe (HyperFrames camera moves + GSAP atmosphere). Use when the user wants an explainer, documentary-style video, or any video built from generated or collected visuals instead of recorded footage.

Qu'est-ce que claude-video-editor ?

claude-video-editor is a Claude Code agent skill that build videos from stills when there is nothing to cut — narration first (ElevenLabs v3 with audio tags), then still keyframes from ANY source (AI image tools, screenshots, photos, HyperFrames-rendered HTML), animated deterministically with the animate-stills recipe (HyperFrames camera moves + GSAP atmosphere). Use when the user wants an explainer, documentary-style video, or any video built from generated or collected visuals instead of recorded footage.

Compatible avec✓Claude Code~Codex CLI~Cursor✓Gemini CLI
npx skills add https://github.com/assafkip/claude-video-editor/tree/HEAD/skills/generate-footage

Demander à votre IA préférée

Ouvre une nouvelle conversation avec cette compétence d'agent déjà préchargée.

Documentation

generate-footage

The plugin's other entry point. video-use edits footage you HAVE. This skill builds videos from footage you DON'T have yet: narration plus still keyframes, composed and animated by HyperFrames, rendered deterministically.

Pipeline (proven in production; see docs/rca/ for why the motion step is deterministic):

script (with v3 audio tags)
  → eleven_v3 narration            (scripts/tts_v3.py)
  → transcribe                     (npx hyperframes transcribe narration.mp3)
  → lock scene windows to sentence boundaries (word-level timestamps)
  → STILLS FROM ANYWHERE           (one keyframe per scene — see below)
  → animate-stills composition     (recipes/04-animate-stills.md, build_scenes.py)
  → lint → render → motion + audio QA (scripts/motion_check.py, contact sheet)

Order matters: voice FIRST, visuals second

Generate and transcribe the narration BEFORE building any composition or collecting keyframes. The spoken sentence boundaries are the scene windows. v3 reads are non-deterministic in pacing — a 120s estimate can come out 100s or 166s — so visuals built first always get re-timed.

Voiceover — ElevenLabs v3 with audio tags

scripts/tts_v3.py — stdlib-only, no pip deps. Key from $ELEVENLABS_API_KEY or the plugin's skills/video-use/.env.

python3 scripts/tts_v3.py --probe                      # verify eleven_v3 + list voices
python3 scripts/tts_v3.py --voice <id> --stability 0.0 --script SCRIPT.md --out narration.mp3
  • Audio tags go in the script text: [excited], [whispers], [softly], [dramatic tone], [warmly]. v3 performs them; v2 ignores them.
  • Stability: 0.0 = Creative (max expressiveness), 0.5 = Natural, 1.0 = Robust. A "flat" read is fixed by lowering stability and re-tagging the script, not by louder words.
  • Delivery levers: CAPS on power words, ... for dramatic pauses, contrast (drop to [whispers] right before a big line).
  • Audition 2-3 voices on the first sentence before committing. The first voice is never the pick.
  • v3 may paraphrase slightly. Transcribe and READ the result; regenerate if a key line drifted.

Stills from anywhere

A keyframe is just a PNG. The pipeline does not care where it came from. One still per scene, named kf-NN.png, listed in the scene manifest (recipes/04-animate-stills.md documents the schema). Sources that work:

  • AI image tools — any text-to-image tool you already have (Gemini, GPT image, Midjourney, local SD, or a hosted actor — see the optional section at the end). Generate from per-scene prompts with one locked style suffix.
  • Screenshots and photos — product UI captures, archive material, scans, real photography. The Ken Burns treatment was invented for this.
  • HyperFrames-rendered HTML — build a styled HTML frame and snapshot it; the engine produces its own stills.

Craft rules that hold regardless of source:

  • Lock ONE style in DESIGN.md and apply it to every frame. Drift across frames reads as different shots once there's motion and grading.
  • Verify the file type. Some tools serve JPEGs with .png names; check magic bytes before composing (file kf-01.png).
  • QA every frame (read the image). One off-style frame is a regenerate, not a keep. Check for: text artifacts, watermarks, anachronisms, faces in close-up (drift risk), wrong palette.
  • No named IP, no real living person's likeness in generated frames for anything posted publicly.

Animation — deterministic, in-repo

Stills become motion with the animate-stills recipe (recipes/04-animate-stills.md): a manifest-driven HyperFrames composition applying camera language (push-ins, drifts, holds) and GSAP atmosphere (mist, rain, light shifts, vignette) over each still, scene windows locked to the transcript. scripts/build_scenes.py generates the scene layers from the manifest; scripts/motion_check.py proves real motion via pixel-diff in the rendered output. Renders identically every run, costs nothing, works offline.

Requires Node >= 22 for the HyperFrames CLI.

Captions and assembly

  • Captions burned in (muted autoplay on Reddit/X), phrase-level, timed from the transcript words.
  • White flash (0.28s) at story turns; soft crossfades elsewhere; dark vignette grade unifies stills from mixed sources and makes captions read.
  • Verify motion in the render: pixel-diff two frames 4s apart inside one scene (scripts/motion_check.py — mean gray diff > 3 = real motion).
  • Audio QA: ffmpeg -af volumedetect — documentary speech sits near -26 dB mean.

Optional: hosted generative tools

These are NOT dependencies. They are one way among several to source stills or clips, and they can disappear without notice — plan accordingly.

  • Apify text-to-image actors (e.g. akash9078/ai-image-generator, ~$0.01/image) worked well for batch keyframe generation. Watch for silent failures (ok:true with a null URL) and mislabeled file types.
  • Hosted image-to-video actors are where this pipeline previously kept its motion step. On 2026-06-11 the actor danitn11/wan22-lightning-image-to-video (2 total users) died upstream mid-production and every run failed instantly; the full analysis is in docs/rca/rca-generate-footage-animation-2026-06-11.md. Generative image-to-video now lives in a separate companion repo with a pluggable BYO-key backend; this plugin's motion step is deterministic and local.

Individual skills in this repo

This repo contains 9 individual skills — each has its own dedicated page.

assafkip/claude-video-editor

GSAP animation reference for HyperFrames. Covers gsap.to(), from(), fromTo(), easing, stagger, defaults, timelines (gsap.timeline(), position parameter, labels, nesting, playback), and performance (transforms, will-change, quickTo). Use when writing GSAP animations in HyperFrames compositions.

assafkip/claude-video-editor

HyperFrames CLI tool — hyperframes init, lint, preview, render, transcribe, tts, doctor, browser, info, upgrade, compositions, docs, benchmark. Use when scaffolding a project, linting or validating compositions, previewing in the studio, rendering to video, transcribing audio, generating TTS, or troubleshooting the HyperFrames environment.

assafkip/claude-video-editor

Install and wire registry blocks and components into HyperFrames compositions. Use when running hyperframes add, installing a block or component, wiring an installed item into index.html, or working with hyperframes.json. Covers the add command, install locations, block sub-composition wiring, component snippet merging, and registry discovery.

assafkip/claude-video-editor

Create video compositions, animations, title cards, overlays, captions, voiceovers, audio-reactive visuals, and scene transitions in HyperFrames HTML. Use when asked to build any HTML-based video content, add captions or subtitles synced to audio, generate text-to-speech narration, create audio-reactive animation (beat sync, glow, pulse driven by music), add animated text highlighting (marker sweeps, hand-drawn circles, burst lines, scribble, sketchout), or add transitions between scenes (crossfades, wipes, reveals, shader transitions). Covers composition authoring, timing, media, and the full video production workflow. For CLI commands (init, lint, preview, render, transcribe, tts) see the hyperframes-cli skill.

assafkip/claude-video-editor

Beginner-friendly end-to-end video creator for HyperFrames. Use when the user says "make a video", "create a video", "new video", "build a video", "video from scratch", "I want to make a video", "help me create a video", or when someone who's never used HyperFrames before arrives with a concept, script, or rough idea and wants a finished MP4. Interviews the user in one pass, then builds the full video with mandatory preview and visual-verification gates.

assafkip/claude-video-editor

Build and iterate short-form vertical (9:16) videos in Hyperframes — TikTok/Reels/Shorts style. Use when Nate says "short-form video", "vertical video", "TikTok/Reels/Shorts", "make a short", "talking-head + motion graphics", or when the target is a 1080x1920 composition with face video + synced scene overlays + karaoke captions. Encodes the full May Shorts 19 playbook: face-mode choreography, audio-synced scene timing, karaoke captions, and the 10-rule quality checklist.

assafkip/claude-video-editor

Edit any video by conversation — the orchestrator for the bundled video pipeline. Use when the user wants to edit, cut, caption, grade, or add motion graphics to footage; drops a folder of clips and says "edit this"; or asks to make a promo/teaser/tutorial/short. Routes to video-use (the cut engine) and HyperFrames (motion-graphics overlays), with a motion-craft library for taste.

assafkip/claude-video-editor

Production pipeline for mathematical and technical animations using Manim Community Edition. Creates 3Blue1Brown-style explainer videos, algorithm visualizations, equation derivations, architecture diagrams, and data stories. Use when users request: animated explanations, math animations, concept visualizations, algorithm walkthroughs, technical explainers, 3Blue1Brown style videos, or any programmatic animation with geometric/mathematical content.

assafkip/claude-video-editor

Capture a website and create a HyperFrames video from it. Use when: (1) a user provides a URL and wants a video, (2) someone says "capture this site", "turn this into a video", "make a promo from my site", (3) the user wants a social ad, product tour, or any video based on an existing website, (4) the user shares a link and asks for any kind of video content. Even if the user just pastes a URL — this is the skill to use.

Skills associés