video-studio
One conversational pipeline for editing any video. Three layers compose; this skill routes between them. The LLM never watches the video — it reads it (transcript + on-demand timeline PNGs) and edits with word-boundary precision.
The three layers (what's bundled here)
-
The cut engine —
video-use(skills/video-use/) Transcribe → pack → reason → EDL → render → self-eval. Owns: filler-word removal, dead-space trimming, color grade, 30ms audio fades, subtitle burn, and therender.pycomposite that lays overlays onto the cut. This is the spine of every edit. Readskills/video-use/SKILL.mdfor the 12 production rules. -
The motion layer — HyperFrames (
skills/hyperframes/,hyperframes-cli/,hyperframes-registry/,gsap/,website-to-hyperframes/) HTML + GSAP compositions rendered through headless Chrome to.movwith alpha. Use for: lower-third captions, title cards, callouts, audio-reactive bits, scene transitions, kinetic text. video-use spawns these as overlays and composites them. -
The craft layer —
make-a-video,short-form-video,docs/MOTION_PHILOSOPHY.mdStoryboard, brand system, pacing, motion taste. Reach here for how it should feel, not how to render.
The standard loop
- Inventory. List the sources in the folder. Don't transcribe unprompted — wait for the user to point at footage and say what they want.
- Transcribe + pack (video-use): one transcript call per source → one
takes_packed.md. - Strategy → confirm. Propose the cut (beats, order, what gets cut). Wait for OK. Never touch the cut without approval.
- EDL. Write the cut list (
ranges= source seconds,overlays= caption .movs in output seconds). Seetemplates/project/for the shape. - Motion (HyperFrames, when a beat needs it): author one composition, render per
variant
--format mov(alpha — webm comes out opaque). - Composite (video-use):
render.py <edl> -o final.mp4 --no-subtitles --no-loudnorm(drop--no-loudnormonly if the source has real audio). - Self-eval → persist. The engine checks every cut boundary before you show
anything. Session memory lives in the project's
project.md.
Hard rules (non-negotiable; taste is free)
- Captions render
--format mov(ProRes 4444 yuva). webm = opaque, no alpha. - Silent source needs a stereo silent track before render.py (its audio path
expects one) and
--no-loudnorm. - Speed-ramps are pre-rendered
setptsclips, referenced as their own EDL sources. No native EDL speed. - Outputs live in
<videos>/edit/— keep the skill dirs clean.
Setup
Run ./install.sh once (checks ffmpeg/node/chrome, sets up video-use's venv, wires the
skills, prompts for the ElevenLabs key). Then cd into a folder of footage and start a
session. See the repo README.md.