Communitygithub.com

edit-video: Talking-Head-Rohmaterial zum fertigen YouTube-Video

Agent Skill, der eine rohe Talking-Head-Aufnahme zu fertigen YouTube-16:9- und Instagram-9:16-Videos schneidet: Transkription, sauberer Schnitt, Motion Graphics, B-Roll, Soundeffekte, optionale Musik, wortgenaue Untertitel, Master auf -14 LUFS und Abschlussprüfung. Englisch, Hindi, Telugu und gemischte Sprache.

Was ist edit-video: Talking-Head-Rohmaterial zum fertigen YouTube-Video?

Schritt 1 des Rezepts YouTube-Episode (der Schnitt), aus SrigadaAkshayKumar/ai-video-editor (veröffentlicht am 2026-10-09; die README zeigt 16:9-Kontaktbögen fertiger Schnitte und beschreibt eine gesichtslose Kanal-Pipeline, die täglich auf GitHub Actions läuft). Du legst eine Rohaufnahme in inbox/talking/ und sagst „schneide video-1, schnell und knackig“. Der Agent übersetzt deine Stilwörter in Einstellungen, transkribiert (Whisper für Englisch, ElevenLabs für Telugu und Hindi), entfernt Wiederholungen, Versprecher und Pausen, schreibt einen Bildplan mit sechs Layouts und 28 Motion-Graphic-Typen plus B-Roll, rendert das Bild mit Remotion, setzt wortgenaue HyperFrames-Untertitel, mischt Effekte und optionale CC0-Musik mit ffmpeg, mastert auf -14 LUFS und prüft Längen, Lautheit, eingefrorene Frames und jeden Schnittpunkt. Jede Entscheidung liegt als JSON vor, jeder Schritt lässt sich einzeln wiederholen.

Funktioniert mit~Claude Code~Codex CLI~Cursor
npx skills add https://github.com/SrigadaAkshayKumar/ai-video-editor/tree/HEAD/.claude/skills/edit-video

In Ihrer bevorzugten KI fragen

Öffnet einen neuen Chat, in dem dieser Agent-Skill bereits geladen ist.

Dokumentation

edit-video — talking head → final-16x9.mp4 + final-9x16.mp4

You are the editor. Tools in tools/ do the mechanical work; you make every editorial decision and keep it in files (work/edl.json, work/visual-plan.json, work/audio-plan.json) so any stage can be redone. Read docs/lessons.md once per session — every rule there was paid for with a broken render.

Architecture (why it is shaped like this)

raw ─ transcribe ─ cut ─┬─ work/cleancut.mp4 ──────────────┐
                        │                                    │ (dialogue)
       visual-plan.json ┴─ Remotion `Doc` (muted) ×2 formats  │
                             └─ HyperFrames captions pass ×2  │
       audio-plan.json ─────────────── tools/mix.mjs ◄───────┘ → work/mix.wav (once)
                                                     finalize: picture + -14 LUFS master + intro/outro
  • Remotion draws the whole picture: the a-roll in one of six stage layouts (full / dock left / dock right / band / corner / takeover), 28 motion-graphic types, b-roll cutaways, camera moves, chapter rail.
  • HyperFrames draws only the word-level captions (browser shaping renders Telugu/Devanagari correctly).
  • ffmpeg does ALL audio: no <audio> in any render — embedding it corrupted dialogue repeatedly in the template this repo inherits from.
  • Both formats share one clean-cut timeline, one visual plan and one mix.

Stages

#StageToolSkill / notes
0Intakenpm run new -- <inbox-name> --brief "…" [--lang te|hi|en] [--style glass|broadcast] [--keyterms "A,B"]see Intake below
1Transcribenpm run transcribe -- <p>English → whisper (free) · Telugu/Hindi → ElevenLabs
2Clean cutnpm run cut -- <p> [--dry-run]clean-cut skill. Most important stage.
3Visual planwrite work/visual-plan.json; npm run visuals -- <p> --check; --stills t1,t2motion-graphics skill
4B-rollnpm run broll -- search|getinside motion-graphics
5Render picturenpm run visuals -- <p>~1-2 min per minute of video per format
6Captionsnpm run scaffold -- <p> (+ npm run captions -- <p> --style …)captions skill
7Audio planwrite work/audio-plan.json; npm run sfx -- list; npm run bgm -- search|get; npm run mix -- <p>sfx, bgm skills
8Finalizenpm run finalize -- <p>lint, render, master, intro/outro, output/credits.md
9Verifynpm run verify -- <p> → Read both work/verify-*/sheet.jpgnever skip
10Packageyoutube-metadata, youtube-thumbnail skillsoffer at the end

npm run status shows progress. After a re-cut: node tools/retime.mjs <p> carries both plans onto the new timeline (old clean → raw → new clean) and lists anything whose moment was cut.

Running it

  1. Intake. The user drops the clip in inbox/talking/ and names it in the prompt ("edit video-1 …", possibly written "/video-1" — strip the slash). npm run inbox lists what is waiting. npm run new -- video-1 --brief "<the user's style words, verbatim>" (+ --lang if they said it, --keyterms for names/brands). The prompt's style instructions are the brief — map them to settings before planning (see "Style brief → settings"). Ask nothing the prompt already answers. Brand/channel: none by default — this repo edits for many clients. Only when the user says whose video it is, load config/brands/<slug>.json (or create it from what they tell you) and pass --brand, plus --intro/--outro clips from inbox/brand/ if they ask for them.
  2. Transcribe, then read work/transcript.md end to end. Check the header's engine and language.
  3. Clean cut — follow clean-cut exactly, including the boundary-audit and fragile-word output.
  4. Plan visuals from work/cleancut.transcript.md (clean-cut times). Write the plan, --check it (it validates every component's required fields — the template's #1 silent defect), fix, then render --stills at 6-10 interesting times and Read them in both formats before the full render. Framing: if the speaker is off-centre in 9:16, set "framing": {"9:16": "40% 45%"}.
  5. Render picture, scaffold captions, plan audio and mix, finalize, verify.
  6. Report: both output paths; raw → clean duration and what was cut (by reason); overlay/b-roll/SFX counts; music used; anything fragile or uncertain; credits that MUST go in the description; and say plainly that you cannot hear the mix — ask the user to listen for the joins, bed level and ducking. Offer /youtube-metadata and /youtube-thumbnail.

Style brief → settings

Translate the user's words into concrete choices and state them in one line before planning:

They saySet
fast / energetic / reels / punchykeepPauseMax 0.25, caption pop 2-3 words, a beat every 2-4s, camera punches, sub-band SFX on structure
calm / professional / documentary / long-formkeepPauseMax 0.4-0.5, caption clean 4-6 words, a beat every 6-10s, fewer SFX, soft bed
minimal / clean / no-fussfew overlays (text marks only where useful), no marquee/glitch, clean captions or none
bold / premium / news / editorial--style broadcast, high-contrast accents
storytelling / emotionalkaraoke captions, piano/cinematic bed, a music gap on the key line
"no music" / "no SFX" / "no captions" / "no b-roll"honour exactly; skip that stage
(music not mentioned)no BGM — the user says at the start of the prompt whether to add music; default is none. When added: CC0 only (--cc0) and SFX must stay clearly audible over it
a colour ("use yellow", a hex)caption --accent, overlay section accents/brand

Anything unusual in the brief overrides these defaults. Keep the brief in project.json.brief so a re-edit later follows the same style.

Review gates

Default: run straight through and report. If the user asks for checkpoints (or project.json.review is "stages"), stop after the clean cut, after the visual plan (show stills), and after the audio plan. A requested change is the new bar for that stage — revise and re-show the same stage.

Rules

  • All plan times are on the clean-cut timeline; EDL times are raw.
  • Overlay text is short English by default (native script lives in the captions) unless the brief says otherwise.
  • Never hand-edit generated files (compositions/captions.html, work/props-*.json) — regenerate.
  • Keys in .env: ELEVENLABS_API_KEY (Telugu/Hindi), PEXELS_API_KEY or PIXABAY_API_KEY (b-roll). Missing one → stop and say exactly which, and that it goes in C:\AI-edits\.env.
  • Windows: write JSON with the Write tool (PowerShell mangles inline JSON; Out-File/Set-Content add a BOM — the tools tolerate it in JSON, but never write code files that way).
  • Cleanup once the user confirms the final: render/, work/visuals-*.mp4, work/captioned-*, work/stills-*, work/verify-* are regenerable. Keep output/, work/*.json, assets/.

Individual skills in this repo

This repo contains 9 individual skills — each has its own dedicated page.

SrigadaAkshayKumar/ai-video-editor

Turn a finished AI-edits video (projects/<p>, or several parts of one long video) into YouTube Shorts. Reuses the edit's word timestamps, reads the whole transcript, picks self-contained dialog moments, adds text hooks + comment prompts, renders 9:16 clips with Remotion. Use when the user asks to clip/make shorts from an edited project or a long video.

SrigadaAkshayKumar/ai-video-editor

Choose, fetch and place background music for an edit — Openverse search (commercial-use CC, or CC0-only), music_cues in work/audio-plan.json with fades, offsets and a deliberate silence gap; tools/mix.mjs conditions every bed, sets it to one measured house level under the voice and sidechain-ducks it. Use when adding, swapping or re-levelling music.

SrigadaAkshayKumar/ai-video-editor

Generate and style word-synced captions (Telugu, Hindi, English, code-mixed) as the HyperFrames captions pass over each format's Remotion picture, fix caption text, black captions out under on-screen text, and check caption content at every cut join. Use when creating, restyling, correcting or repositioning captions in any edit.

SrigadaAkshayKumar/ai-video-editor

Editorial rules for turning a raw talking video or voiceover into a clean cut — removing retakes, repetitions, false starts, stutters, fillers, mispronounced words, contradictions/inconsistencies, off-script chatter and dead air — by writing work/edl.json for tools/cut.mjs. Telugu, Hindi, English and code-mixed speech. Use for stage 2 of edit-video or whenever the user asks to tighten/re-cut a video.

SrigadaAkshayKumar/ai-video-editor

Turns a bare voiceover (no presenter on camera) into a finished faceless documentary/explainer for YouTube 16:9 and Instagram 9:16 — clean-cuts the narration, DIRECTS it (story spine, tension curve, hooks, involvement beats), researches the claims and screenshots the real sources, plans a scene track (stock b-roll, stills, news screenshots, pure-graphic beats) with motion graphics, designs music + SFX, renders, captions and credits it. Use when the user gives a voiceover/narration file (or says "faceless video", "make a video from this VO", "documentary").

SrigadaAkshayKumar/ai-video-editor

Plan the picture of an edit in work/visual-plan.json — motion graphics (28 component types across six stage layouts that dock/shrink/dim the video), b-roll cutaways from Pexels/Pixabay, camera moves, a-roll framing for 9:16 — validate it, QA it with stills in both formats, and render it with Remotion. Use for any titles, charts, stats, lower thirds, diagrams, b-roll or zooms in a talking-head or faceless edit.

SrigadaAkshayKumar/ai-video-editor

Design the sound effects of an edit in work/audio-plan.json sfx_events — which effect, where, how often — using the measured SFX library (repo assets/sfx, HyperFrames' bundled set, CC0 fetches from Openverse) and the spectral rule that sub-band cues carry structure while vocal-band cues are rationed. Mixed and levelled by tools/mix.mjs. Use when adding or adjusting SFX in any edit.

SrigadaAkshayKumar/ai-video-editor

Write the upload package for a video this repo edited — YouTube title options, description with chapters, tags, pinned comment, thumbnail text, and the Instagram Reels caption + hashtags for the 9:16 cut — optimised for whichever channel/client the user says it is for (config/brands/ profiles, only when named). Use when the user asks what to title a video, for a description, tags, chapters, "the YouTube metadata", an Instagram caption, or SEO/CTR advice on a finished render.

SrigadaAkshayKumar/ai-video-editor

Design and build a high-CTR thumbnail for a video this repo edited — hook selection, the real face frame pulled from the footage, background cut-out, type set with real fonts (Telugu/Devanagari included), the YouTube 1280x720 thumbnail and the Instagram 1080x1920 cover, an optional image-generation prompt for the background layer, and a feed-size legibility check. Use when the user asks for a thumbnail, thumbnail ideas, a cover image, or how to make the thumbnail more clickable.

Verwandte Skills