Communitygithub.com

SrigadaAkshayKumar/ai-video-editor

Generate and style word-synced captions (Telugu, Hindi, English, code-mixed) as the HyperFrames captions pass over each format's Remotion picture, fix caption text, black captions out under on-screen text, and check caption content at every cut join. Use when creating, restyling, correcting or repositioning captions in any edit.

Was ist ai-video-editor?

ai-video-editor is a Claude Code agent skill that generate and style word-synced captions (Telugu, Hindi, English, code-mixed) as the HyperFrames captions pass over each format's Remotion picture, fix caption text, black captions out under on-screen text, and check caption content at every cut join. Use when creating, restyling, correcting or repositioning captions in any edit.

Funktioniert mit~Claude Code~Codex CLI~Cursor
npx skills add https://github.com/SrigadaAkshayKumar/ai-video-editor/tree/HEAD/.claude/skills/captions

In Ihrer bevorzugten KI fragen

Öffnet einen neuen Chat, in dem dieser Agent-Skill bereits geladen ist.

Dokumentation

captions

npm run captions -- <p> [--format 9:16] [--style pop|clean|karaoke] [--max-words N] [--position bottom|middle] [--accent "#hex"] [--uppercase]

Writes edit-<fmt>/compositions/captions.html for every format (or just --format) from work/cleancut.words.json (already on the clean-cut timeline). style, accent and uppercase are shared by all formats (one brand look); max-words and position are per format — pass them with --format. Everything is remembered in project.json → captions, so re-running with no flags keeps it. The scaffold already mounts the captions (<div id="captions" … data-track-index="4">).

Defaults: bottom band in both formats — 16:9 5 words, 64px · 9:16 3 words, 76px at y≈1348, inside Instagram's safe zone. The Remotion graphics keep clear of exactly this band (CAPTION_SAFE_Y in remotion/src/doc/theme.ts), so captions and graphics never collide by construction.

npm run scaffold -- <p> builds the captions pass over work/visuals-<fmt>.mp4 and generates the captions; re-run npm run captions alone to restyle.

Choosing a style

ContentStyleWords/group
Shorts / reels, energetic, hookspop (white, thick outline, active word in accent, word bounce)2–3
Educational / interview / long-formclean (dark pill, accent on active word)4–6
Story / calm / music-ledkaraoke (dim → white fill as words are spoken)3–5

Accent = the brand colour from the brief, else #FFD400. Use --uppercase only for English-only videos (no effect on Indic scripts).

Text corrections

Never hand-edit the generated HTML. Fix words at the source:

  • Add "textFixes": { "<raw word id>": "Correct" } to work/edl.json and re-run npm run cut then npm run scaffold (refresh) — raw ids are the src field in work/cleancut.words.json.
  • Common fixes: brand/person names, English words the ASR wrote in Telugu/Devanagari script when the user wants them in Latin (or vice versa), casing at the start of a sentence after a cut, numbers (ఇరవై → 20 when it reads better).

Scripts and fonts

Telugu (Noto Sans Telugu), Devanagari (Noto Sans Devanagari) and Latin (Poppins) fonts ship as woff2 in each edit-<fmt>/assets/fonts/ and are declared in both index.html and the captions template — the renderer has no system fonts. If you add a new font anywhere, it needs its own @font-face with a local woff2, or lint fails with font_family_without_font_face.

Indic words are long: in 9:16 keep --max-words ≤ 3; the generator also caps characters per group.

Blackouts

When a card already shows the words (sample_answer, a full-screen quote being read), add "caption_blackouts": [{ "start", "end" }] to work/visual-plan.json and re-run npm run captions. Words inside are dropped and any hole ends a caption group, so no line spans the card (the template's "word before the hole + word after it" garbage line).

Check the joins

npm run verify -- <p> grabs a frame just after every cut. Read the sheet: each caption must be a real fragment of one sentence. A splice of two different takes in one line is a content bug no linter sees. One caption group is visible at a time; groups hold through short pauses and never overlap.

npm run captions marks the stage done.

Individual skills in this repo

This repo contains 9 individual skills — each has its own dedicated page.

SrigadaAkshayKumar/ai-video-editor

Turn a finished AI-edits video (projects/<p>, or several parts of one long video) into YouTube Shorts. Reuses the edit's word timestamps, reads the whole transcript, picks self-contained dialog moments, adds text hooks + comment prompts, renders 9:16 clips with Remotion. Use when the user asks to clip/make shorts from an edited project or a long video.

SrigadaAkshayKumar/ai-video-editor

Choose, fetch and place background music for an edit — Openverse search (commercial-use CC, or CC0-only), music_cues in work/audio-plan.json with fades, offsets and a deliberate silence gap; tools/mix.mjs conditions every bed, sets it to one measured house level under the voice and sidechain-ducks it. Use when adding, swapping or re-levelling music.

SrigadaAkshayKumar/ai-video-editor

Editorial rules for turning a raw talking video or voiceover into a clean cut — removing retakes, repetitions, false starts, stutters, fillers, mispronounced words, contradictions/inconsistencies, off-script chatter and dead air — by writing work/edl.json for tools/cut.mjs. Telugu, Hindi, English and code-mixed speech. Use for stage 2 of edit-video or whenever the user asks to tighten/re-cut a video.

SrigadaAkshayKumar/ai-video-editor

Turns a bare voiceover (no presenter on camera) into a finished faceless documentary/explainer for YouTube 16:9 and Instagram 9:16 — clean-cuts the narration, DIRECTS it (story spine, tension curve, hooks, involvement beats), researches the claims and screenshots the real sources, plans a scene track (stock b-roll, stills, news screenshots, pure-graphic beats) with motion graphics, designs music + SFX, renders, captions and credits it. Use when the user gives a voiceover/narration file (or says "faceless video", "make a video from this VO", "documentary").

SrigadaAkshayKumar/ai-video-editor

End-to-end editor for TALKING-HEAD videos (a person on camera) in this repo. Use when the user hands over a raw video with a speaker and wants a finished edit, or asks to continue/resume/redo any stage of an existing talking-head project — clean cut, motion graphics, b-roll, SFX, music, captions, final render, thumbnail. Delivers YouTube 16:9 and Instagram 9:16. Telugu, Hindi, English and code-mixed speech. For a voiceover with no presenter use edit-video-faceless.

SrigadaAkshayKumar/ai-video-editor

Plan the picture of an edit in work/visual-plan.json — motion graphics (28 component types across six stage layouts that dock/shrink/dim the video), b-roll cutaways from Pexels/Pixabay, camera moves, a-roll framing for 9:16 — validate it, QA it with stills in both formats, and render it with Remotion. Use for any titles, charts, stats, lower thirds, diagrams, b-roll or zooms in a talking-head or faceless edit.

SrigadaAkshayKumar/ai-video-editor

Design the sound effects of an edit in work/audio-plan.json sfx_events — which effect, where, how often — using the measured SFX library (repo assets/sfx, HyperFrames' bundled set, CC0 fetches from Openverse) and the spectral rule that sub-band cues carry structure while vocal-band cues are rationed. Mixed and levelled by tools/mix.mjs. Use when adding or adjusting SFX in any edit.

SrigadaAkshayKumar/ai-video-editor

Write the upload package for a video this repo edited — YouTube title options, description with chapters, tags, pinned comment, thumbnail text, and the Instagram Reels caption + hashtags for the 9:16 cut — optimised for whichever channel/client the user says it is for (config/brands/ profiles, only when named). Use when the user asks what to title a video, for a description, tags, chapters, "the YouTube metadata", an Instagram caption, or SEO/CTR advice on a finished render.

SrigadaAkshayKumar/ai-video-editor

Design and build a high-CTR thumbnail for a video this repo edited — hook selection, the real face frame pulled from the footage, background cut-out, type set with real fonts (Telugu/Devanagari included), the YouTube 1280x720 thumbnail and the Instagram 1080x1920 cover, an optional image-generation prompt for the background layer, and a feed-size legibility check. Use when the user asks for a thumbnail, thumbnail ideas, a cover image, or how to make the thumbnail more clickable.

Verwandte Skills