Communitygithub.com

SrigadaAkshayKumar/ai-video-editor

Plan the picture of an edit in work/visual-plan.json — motion graphics (28 component types across six stage layouts that dock/shrink/dim the video), b-roll cutaways from Pexels/Pixabay, camera moves, a-roll framing for 9:16 — validate it, QA it with stills in both formats, and render it with Remotion. Use for any titles, charts, stats, lower thirds, diagrams, b-roll or zooms in a talking-head or faceless edit.

Was ist ai-video-editor?

ai-video-editor is a Claude Code agent skill that plan the picture of an edit in work/visual-plan.json — motion graphics (28 component types across six stage layouts that dock/shrink/dim the video), b-roll cutaways from Pexels/Pixabay, camera moves, a-roll framing for 9:16 — validate it, QA it with stills in both formats, and render it with Remotion. Use for any titles, charts, stats, lower thirds, diagrams, b-roll or zooms in a talking-head or faceless edit.

Funktioniert mit~Claude Code~Codex CLI~Cursor
npx skills add https://github.com/SrigadaAkshayKumar/ai-video-editor/tree/HEAD/.claude/skills/motion-graphics

In Ihrer bevorzugten KI fragen

Öffnet einen neuen Chat, in dem dieser Agent-Skill bereits geladen ist.

Dokumentation

motion-graphics — work/visual-plan.json

Graphics serve the speech: every beat shows what is being said at that moment. Variety is a hard requirement — the template's first visual system was rejected as "very monotonous: a few text visuals, a few screen partitions with graphics on the left, that's it." The edit must keep changing shape: how the frame is divided, what kind of object appears, how it arrives.

Plan shape (clean-cut times)

{
  "style": "glass",                         // or "broadcast" (flat editorial). Default: project style
  "framing": { "9:16": "45% 40%" },         // a-roll crop per format (talking head)
  "rail": true,                             // chapter rail from chapter_open titles
  "overlays": [ { "start": 41.5, "duration": 3.8, "type": "chapter_open", "text": "What Is Swayam?", "eyebrow": "Section 01", "section": 1 } ],
  "brolls": [ { "start": 88.0, "duration": 3.4, "src": "assets/broll/empty-office.mp4", "media_start": 1, "motion": "pan_right", "grade": "cool", "transition": "fade", "reason": "over 'onboarding is not happening'" } ],
  "camera_moves": [ { "start": 250.2, "duration": 2.4, "kind": "punch_in", "scale": 1.12 } ],
  "caption_blackouts": [ { "start": 300, "end": 318 } ],
  "scenes": [ … faceless only, see edit-video-faceless … ]
}

Any item can carry "only": "9:16" (or "16:9") to appear in one format.

Six stage layouts — picked by the component, never set directly

LayoutFrame (16:9)Frame (9:16)Components
fullvideo untouchedsamekeyword_chip, lower_third, side_note, annotation, marquee_strip, speed_hint, sample_answer
dockvideo card one side, graphic column the other (side = graphic side)video card top, graphic strip belowbar_chart, line_chart, donut_chart, progress_ring, step_progress, timeline, bullet_list, comparison, term_card
bandvideo top 2/3, wide strip belowvideo top 42%, strip belowfact_band
cornergraphic owns frame, video small cardcard top-rightnumber_roll, stat_trio, flow_diagram, matrix_grid, checklist
takeovervideo dims + blurs behind full-frame typesamechapter_open, big_statement, word_swap, quote_pull, poll_prompt, title_slide

Alternate dock side between consecutive dock beats. Over a whole video no layout carries more than ~a third of the reframing beats.

Component vocabulary — the exact fields each one reads (validated by --check)

Takeover (full stops of the edit)

  • title_slide — text (title), eyebrow, optional value (big ghosted year), style_hint (subtitle). Readable at frame 0; plays over the first line.
  • chapter_open — text (2-5 words), eyebrow ("Section 02"), optional railLabel, numeral ("" hides the ghost number). ≥3s; nothing else in its window.
  • big_statement — text (<~8 words), eyebrow = ONE word of text to highlight.
  • word_swap — from, to, optional text caption. Hinge moments only ("Certificate → Skill").
  • quote_pull — text (the quote), eyebrow (attribution).
  • poll_prompt — text (question), items [{label, value 0-100 = visual weight only}], eyebrow.

Corner (the graphic IS the content)

  • number_roll — numeric value, prefix/suffix/decimals, text label. A phrase, not a figure? use big_statement.
  • stat_trio — items [{label, sublabel (prefix e.g. "₹"), value (number)}] ×2-3.
  • flow_diagram — items [{label, sublabel}] ×2-5, text heading. Reach for it whenever a process/pipeline is described.
  • matrix_grid — items [{label}], text. Sets: tools, roles, vendors.
  • checklist — items [{label, state: "yes"|"no"}], text. "Do this, not that".

Dock (all need side)

  • bar_chart / line_chart / donut_chart — items [{label, value}], text.
  • progress_ring — value 0-100 (not current), text.
  • step_progress — items [{label}], current_step (number), text.
  • timeline — items [{label, sublabel}], optional current index, text.
  • bullet_list — items [{label}], text.
  • comparison — left and right {label, text} — NOT items (blank cards otherwise).
  • term_card — text (term), eyebrow, items[0].label (definition).

Band — fact_band — items [{label, sublabel}] ×2-3.

Full-bleed marks (never the centre where the face is; side left/right)

  • keyword_chip — text (1-3 words) · lower_third — text, eyebrow · side_note — text
  • annotation — text, shape: underline|circle|bracket (cheap and lively — use it often)
  • marquee_strip — text (≤2 per video: the loudest object) · speed_hint — optional text
  • sample_answer — text with [placeholders] and \n lines; holds still to be screenshotted; pair with a caption_blackouts span (the card already says it).

Rules

  • Overlay text is short English (native-script words live in the captions), ~15 chars/second, ≥1.5s on screen. Text-only types ≤ about half the plan; if content has a shape, draw the shape.
  • A beat every ~4-5s (talking head) / ~3-4s (faceless); sustained graphics count for their length. Don't put the same component twice in a row (the motion repeats). Overlays don't overlap each other.
  • A full-bleed mark never overlaps a reframing window (±0.62s) — the frame is mid-move.
  • section (0-4) picks the accent colour; it follows the running chapter if omitted.
  • 9:16: the same plan renders portrait-aware (cards top, graphics between card and captions, everything inside Instagram's safe zone). Long text wraps more — keep it short, check the stills.

B-roll cutaways (talking head)

  • npm run broll -- search <p> "literal visual phrase" --format 16:9|9:16 → Read the thumbnails (stock search returns wrong subjects/regions/watermarks) → npm run broll -- get <p> provider:id --name slug. English, concrete, filmable: "analyst studying financial charts", not "margins".
  • 2.0-4.5s each; cut away to something the speaker literally just named; ≥8s apart, roughly one per 25-40s; never during takeover/corner/chapter_open (inside dock/band is fine — it plays in the card); don't reuse a query. motion ken_in|ken_out|pan_left|pan_right|static, grade neutral|cool|warm|noir|hot (carries the mood), transition fade|cut|whip|flash.

Camera moves

punch_in (hard emphasis), drift_in (long explanation), pull_out (after a chapter opener), whip (~0.8s into a new section). Scale 1.06-1.15, ~one per 20-30s of full-bleed time; moves that collide with a reframing window are dropped automatically.

QA loop (don't skip — a full render is minutes)

  1. npm run visuals -- <p> --check until 0 errors (it names the exact missing/mis-named field).
  2. npm run visuals -- <p> --stills 3.2,41.8,… at one moment per component family → Read work/stills-16x9/*.png and work/stills-9x16/*.png. Look for: empty components, text over the caption band, collisions, a face covered for long, off-centre 9:16 framing.
  3. npm run visuals -- <p> (both formats), then the captions pass.
  4. New component types go in remotion/src/doc/overlays/ + the registry + RULES in tools/render-visuals.mjs; check them with npm run gallery (one still per type, both formats).

Individual skills in this repo

This repo contains 9 individual skills — each has its own dedicated page.

SrigadaAkshayKumar/ai-video-editor

Turn a finished AI-edits video (projects/<p>, or several parts of one long video) into YouTube Shorts. Reuses the edit's word timestamps, reads the whole transcript, picks self-contained dialog moments, adds text hooks + comment prompts, renders 9:16 clips with Remotion. Use when the user asks to clip/make shorts from an edited project or a long video.

SrigadaAkshayKumar/ai-video-editor

Choose, fetch and place background music for an edit — Openverse search (commercial-use CC, or CC0-only), music_cues in work/audio-plan.json with fades, offsets and a deliberate silence gap; tools/mix.mjs conditions every bed, sets it to one measured house level under the voice and sidechain-ducks it. Use when adding, swapping or re-levelling music.

SrigadaAkshayKumar/ai-video-editor

Generate and style word-synced captions (Telugu, Hindi, English, code-mixed) as the HyperFrames captions pass over each format's Remotion picture, fix caption text, black captions out under on-screen text, and check caption content at every cut join. Use when creating, restyling, correcting or repositioning captions in any edit.

SrigadaAkshayKumar/ai-video-editor

Editorial rules for turning a raw talking video or voiceover into a clean cut — removing retakes, repetitions, false starts, stutters, fillers, mispronounced words, contradictions/inconsistencies, off-script chatter and dead air — by writing work/edl.json for tools/cut.mjs. Telugu, Hindi, English and code-mixed speech. Use for stage 2 of edit-video or whenever the user asks to tighten/re-cut a video.

SrigadaAkshayKumar/ai-video-editor

Turns a bare voiceover (no presenter on camera) into a finished faceless documentary/explainer for YouTube 16:9 and Instagram 9:16 — clean-cuts the narration, DIRECTS it (story spine, tension curve, hooks, involvement beats), researches the claims and screenshots the real sources, plans a scene track (stock b-roll, stills, news screenshots, pure-graphic beats) with motion graphics, designs music + SFX, renders, captions and credits it. Use when the user gives a voiceover/narration file (or says "faceless video", "make a video from this VO", "documentary").

SrigadaAkshayKumar/ai-video-editor

End-to-end editor for TALKING-HEAD videos (a person on camera) in this repo. Use when the user hands over a raw video with a speaker and wants a finished edit, or asks to continue/resume/redo any stage of an existing talking-head project — clean cut, motion graphics, b-roll, SFX, music, captions, final render, thumbnail. Delivers YouTube 16:9 and Instagram 9:16. Telugu, Hindi, English and code-mixed speech. For a voiceover with no presenter use edit-video-faceless.

SrigadaAkshayKumar/ai-video-editor

Design the sound effects of an edit in work/audio-plan.json sfx_events — which effect, where, how often — using the measured SFX library (repo assets/sfx, HyperFrames' bundled set, CC0 fetches from Openverse) and the spectral rule that sub-band cues carry structure while vocal-band cues are rationed. Mixed and levelled by tools/mix.mjs. Use when adding or adjusting SFX in any edit.

SrigadaAkshayKumar/ai-video-editor

Write the upload package for a video this repo edited — YouTube title options, description with chapters, tags, pinned comment, thumbnail text, and the Instagram Reels caption + hashtags for the 9:16 cut — optimised for whichever channel/client the user says it is for (config/brands/ profiles, only when named). Use when the user asks what to title a video, for a description, tags, chapters, "the YouTube metadata", an Instagram caption, or SEO/CTR advice on a finished render.

SrigadaAkshayKumar/ai-video-editor

Design and build a high-CTR thumbnail for a video this repo edited — hook selection, the real face frame pulled from the footage, background cut-out, type set with real fonts (Telugu/Devanagari included), the YouTube 1280x720 thumbnail and the Instagram 1080x1920 cover, an optional image-generation prompt for the background layer, and a feed-size legibility check. Use when the user asks for a thumbnail, thumbnail ideas, a cover image, or how to make the thumbnail more clickable.

Verwandte Skills