Communitygithub.com

SrigadaAkshayKumar/ai-video-editor

Design and build a high-CTR thumbnail for a video this repo edited — hook selection, the real face frame pulled from the footage, background cut-out, type set with real fonts (Telugu/Devanagari included), the YouTube 1280x720 thumbnail and the Instagram 1080x1920 cover, an optional image-generation prompt for the background layer, and a feed-size legibility check. Use when the user asks for a thumbnail, thumbnail ideas, a cover image, or how to make the thumbnail more clickable.

What is ai-video-editor?

ai-video-editor is a Claude Code agent skill that design and build a high-CTR thumbnail for a video this repo edited — hook selection, the real face frame pulled from the footage, background cut-out, type set with real fonts (Telugu/Devanagari included), the YouTube 1280x720 thumbnail and the Instagram 1080x1920 cover, an optional image-generation prompt for the background layer, and a feed-size legibility check. Use when the user asks for a thumbnail, thumbnail ideas, a cover image, or how to make the thumbnail more clickable.

Works with~Claude Code~Codex CLI~Cursor
npx skills add https://github.com/SrigadaAkshayKumar/ai-video-editor/tree/HEAD/.claude/skills/youtube-thumbnail

Ask in your favorite AI

Open a new chat with this agent skill pre-loaded.

Documentation

Thumbnail (and Reels cover)

The rule that shapes everything

An image model cannot produce the creator's real face, and cannot set type (it mangles Latin and can't shape Telugu conjuncts at all). A generated "Indian man with glasses" is a stranger, and a stranger on a personal channel destroys recognition. So every thumbnail is layers:

  1. the face — a real frame from the footage, cut out;
  2. the background/graphic — generated, a frame, or a plain accent gradient;
  3. the type — set by tools/thumbnail.mjs render with real fonts. Faceless videos: no face layer; a graphic/stock background + type.

1. Read the video first (never ask what it is about)

work/cleancut.transcript.md, work/visual-plan.json (big_statements are already the punchlines — the best one is usually the thumbnail line; chapter_opens are the sections), the title if the youtube-metadata skill ran (the thumbnail completes the title, never repeats it).

2. One hook

The single moment with the biggest gap between what the viewer assumes and what the video says: the counter-intuitive answer, the number they didn't know, the binary they can't resolve without watching. For student/early-career audiences: money, jobs and wasted effort; fear of wasted effort beats promise of reward. Lead with the risk — but only one the video actually delivers.

3. Face frame (talking head)

node tools/thumbnail.mjs frames <p> --auto            # punch-in moments + big_statement starts
node tools/thumbnail.mjs frames <p> --at 41.2,88.0    # or specific clean-cut times

Frames come from the RAW footage at full quality (mapped through work/timeline.json). Read them; pick eyes to camera, brows up or furrowed, mouth mid-word, hands in frame if gesturing. Say which and why. node tools/thumbnail.mjs cutout <p> work/thumb/face-41.2.png → work/thumb/cutout.png (Read it: hair and hand edges intact?).

4. Background (optional generated layer)

Give the user an image-generation prompt for the BACKGROUND ONLY, written as composition and lighting, not a scene: "dark navy studio backdrop, single warm rim light from upper right, soft amber glow on the left third, fine dot-grid texture, deep vignette, empty negative space on the right two thirds for a subject cut-out, cinematic, high contrast, 16:9 1280x720". Negative prompt always: no text, no letters, no words, no watermark, no logos, no people, no faces, no hands, cluttered, low contrast, busy. ("no people" matters — the generated layer must not add a second person.) Or skip it: the renderer draws an accent gradient.

5. Render

node tools/thumbnail.mjs render <p> --big "NOT ENOUGH" --small "Certificates alone" \
     --native "ఉద్యోగం రాదు" --cutout work/thumb/cutout.png [--background work/thumb/bg.png] \
     --accent "#ffcc00" --side right

→ output/thumbnail-16x9.png (1280x720) and output/cover-9x16.png (1080x1920 Reels cover). Layout rules baked in: face one third, text the other; two type sizes; bottom-right clear for the duration stamp. Text 3-4 words; one word huge. A native-script line against an English title signals the language instantly. Accent = the section accent of the hook's chapter, or the brand's colour if the user named a brand (config/brands/<slug>.json). No brand → no channel styling assumed.

6. Check before handing over

Read the feed-size previews (work/thumb/preview-120-*.png): big word readable? face recognisable? Does it duplicate the title? Would a stranger get it? Does the video deliver what it implies?

Don't

No fabricated logos, certificates, salary slips or documents; no trademarks by default (flag, let the user decide); no reflexive red arrow + shocked face — credibility over alarm; no emoji.

Individual skills in this repo

This repo contains 9 individual skills — each has its own dedicated page.

SrigadaAkshayKumar/ai-video-editor

Turn a finished AI-edits video (projects/<p>, or several parts of one long video) into YouTube Shorts. Reuses the edit's word timestamps, reads the whole transcript, picks self-contained dialog moments, adds text hooks + comment prompts, renders 9:16 clips with Remotion. Use when the user asks to clip/make shorts from an edited project or a long video.

SrigadaAkshayKumar/ai-video-editor

Choose, fetch and place background music for an edit — Openverse search (commercial-use CC, or CC0-only), music_cues in work/audio-plan.json with fades, offsets and a deliberate silence gap; tools/mix.mjs conditions every bed, sets it to one measured house level under the voice and sidechain-ducks it. Use when adding, swapping or re-levelling music.

SrigadaAkshayKumar/ai-video-editor

Generate and style word-synced captions (Telugu, Hindi, English, code-mixed) as the HyperFrames captions pass over each format's Remotion picture, fix caption text, black captions out under on-screen text, and check caption content at every cut join. Use when creating, restyling, correcting or repositioning captions in any edit.

SrigadaAkshayKumar/ai-video-editor

Editorial rules for turning a raw talking video or voiceover into a clean cut — removing retakes, repetitions, false starts, stutters, fillers, mispronounced words, contradictions/inconsistencies, off-script chatter and dead air — by writing work/edl.json for tools/cut.mjs. Telugu, Hindi, English and code-mixed speech. Use for stage 2 of edit-video or whenever the user asks to tighten/re-cut a video.

SrigadaAkshayKumar/ai-video-editor

Turns a bare voiceover (no presenter on camera) into a finished faceless documentary/explainer for YouTube 16:9 and Instagram 9:16 — clean-cuts the narration, DIRECTS it (story spine, tension curve, hooks, involvement beats), researches the claims and screenshots the real sources, plans a scene track (stock b-roll, stills, news screenshots, pure-graphic beats) with motion graphics, designs music + SFX, renders, captions and credits it. Use when the user gives a voiceover/narration file (or says "faceless video", "make a video from this VO", "documentary").

SrigadaAkshayKumar/ai-video-editor

End-to-end editor for TALKING-HEAD videos (a person on camera) in this repo. Use when the user hands over a raw video with a speaker and wants a finished edit, or asks to continue/resume/redo any stage of an existing talking-head project — clean cut, motion graphics, b-roll, SFX, music, captions, final render, thumbnail. Delivers YouTube 16:9 and Instagram 9:16. Telugu, Hindi, English and code-mixed speech. For a voiceover with no presenter use edit-video-faceless.

SrigadaAkshayKumar/ai-video-editor

Plan the picture of an edit in work/visual-plan.json — motion graphics (28 component types across six stage layouts that dock/shrink/dim the video), b-roll cutaways from Pexels/Pixabay, camera moves, a-roll framing for 9:16 — validate it, QA it with stills in both formats, and render it with Remotion. Use for any titles, charts, stats, lower thirds, diagrams, b-roll or zooms in a talking-head or faceless edit.

SrigadaAkshayKumar/ai-video-editor

Design the sound effects of an edit in work/audio-plan.json sfx_events — which effect, where, how often — using the measured SFX library (repo assets/sfx, HyperFrames' bundled set, CC0 fetches from Openverse) and the spectral rule that sub-band cues carry structure while vocal-band cues are rationed. Mixed and levelled by tools/mix.mjs. Use when adding or adjusting SFX in any edit.

SrigadaAkshayKumar/ai-video-editor

Write the upload package for a video this repo edited — YouTube title options, description with chapters, tags, pinned comment, thumbnail text, and the Instagram Reels caption + hashtags for the 9:16 cut — optimised for whichever channel/client the user says it is for (config/brands/ profiles, only when named). Use when the user asks what to title a video, for a description, tags, chapters, "the YouTube metadata", an Instagram caption, or SEO/CTR advice on a finished render.

Related Skills