Communitygithub.com

hoodini/ai-agents-skills

Edit any video into a captioned showcase — transcribe (any language, defaults to large-v3), present a transcript_review.txt for the user to fix mishears BEFORE rendering, then build a HyperFrames composition with liquid-glass caption pills, liquid blob background, liquid morph wipes, optional behind-subject text via background removal, and render the final video. Use whenever the user provides a video file and asks to edit it, caption it, add subtitles, fix existing captions, make a reel/promo/captioned tutorial, or "do the same" pattern as a prior captioned video. Supports English, Hebrew, and any Whisper-supported language. **Renders both 16:9 (YouTube / horizontal) and 9:16 (TikTok / Instagram Reels / YouTube Shorts) from the SAME 16:9 source** — vertical mode uses a centered footage strip with a blurred backdrop + liquid blobs and a vertical-tuned caption pill, no need to re-shoot. THE PIPELINE PAUSES FOR USER APPROVAL on the transcript before final render — this is the support mechanism for getting ca...

O que é ai-agents-skills?

ai-agents-skills is a Claude Code agent skill that edit any video into a captioned showcase — transcribe (any language, defaults to large-v3), present a transcript_review.txt for the user to fix mishears BEFORE rendering, then build a HyperFrames composition with liquid-glass caption pills, liquid blob background, liquid morph wipes, optional behind-subject text via background removal, and render the final video. Use whenever the user provides a video file and asks to edit it, caption it, add subtitles, fix existing captions, make a reel/promo/captioned tutorial, or "do the same" pattern as a prior captioned video. Supports English, Hebrew, and any Whisper-supported language. **Renders both 16:9 (YouTube / horizontal) and 9:16 (TikTok / Instagram Reels / YouTube Shorts) from the SAME 16:9 source** — vertical mode uses a centered footage strip with a blurred backdrop + liquid blobs and a vertical-tuned caption pill, no need to re-shoot. THE PIPELINE PAUSES FOR USER APPROVAL on the transcript before final render — this is the support mechanism for getting ca...

Funciona com✓Claude Code~Codex CLI~Cursor
npx skills add https://github.com/hoodini/ai-agents-skills/tree/HEAD/skills/video-edit

Perguntar na sua IA favorita

Abre um novo chat com esta habilidade de agente já pré-carregada.

Documentação

Video Edit — Captioned Showcase Pipeline

End-to-end captioned video editor on top of HyperFrames. The user gives you a video; you orchestrate transcribe → review → render and ALWAYS pause for transcript approval before the long render.

Where this skill sits in the YUV.AI pyramid

video-edit is in the middle tier of the YUV.AI skills pyramid alongside yuv-design-system, yuv-decks, yuv-viral-video, parallax-landing-page, and video-to-landing-page. The top-tier orchestrator yuv-pilot routes here whenever the user wants a captioned showcase, tutorial, or talking-head edit with subtitles.

This is the more general video sibling to yuv-viral-video. The split:

  • yuv-viral-video — opinionated YUV.AI viral-short pipeline (MrBeast pacing, signature editorial style)
  • video-edit — general captioned editor with transcript-review-before-render (Hebrew + English + any Whisper language)

For YUV.AI-branded captioned video, pair this skill with yuv-design-system (Neon mode for type/palette decisions). For generic / third-party captioned video, this skill works standalone.

When to invoke

  • A path to a video file (mp4/mov/mkv) + a request to "edit", "caption", "add subtitles", "make a reel/promo", "do the same"
  • "Fix the captions / Hebrew misspells" — re-enter at the review step on an existing project
  • Any captioned tutorial / talking-head / promo build

Save location

Default: ~/Documents/yuv-projects/videos/<slug>/ — always save captioned video projects here so renders are findable. The <slug> is short, derived from the topic or source filename.

mkdir -p ~/Documents/yuv-projects/videos
cd ~/Documents/yuv-projects/videos
# Initialize the project here.

Final render lands at ~/Documents/yuv-projects/videos/<slug>/renders/<name>_FINAL.mp4. Tell the user where the video lives at the end of the render.


Workflow (12 steps)

  1. Probe the source — ffprobe for dimensions, fps, duration, audio.

  2. Scaffold — cd ~/Documents/yuv-projects/videos && npx hyperframes init <slug> --video <path> --non-interactive. Rename the copied video to source.mp4.

  3. Extract audio — ffmpeg -i source.mp4 -vn -ac 1 -ar 16000 audio.wav.

  4. Transcribe — copy references/transcribe.py into the project. Default model large-v3 (best Hebrew). CUDA usually fails on Windows (missing cuDNN); the script falls back to CPU int8. Force language="he" for Hebrew, language="en" for English; otherwise auto-detect.

  5. Apply known corrections — copy references/corrections-hebrew.md content into a corrections.json at the project root (keys = wrong token, values = correct token).

  6. 🛑 STOP — start the review server and let the user approve in a webapp. First apply known corrections: copy references/make_review.py into the project and run python make_review.py. It applies corrections.json to transcript.json.

    Then spawn the review server as a background task (it blocks until the user clicks "Approve & Render" in the browser):

    python "$HOME/.claude/skills/video-edit/references/serve_review.py" .
    # On Windows: python "C:\Users\<you>\.claude\skills\video-edit\references\serve_review.py" .
    

    The server prints a line like REVIEW_URL=http://localhost:PORT/. Grab that URL from the background-task output (or read stdout) and send the user:

    👉 Review your transcript here: http://localhost:PORT/ When you click Approve & Render, I'll continue automatically.

    The agent does not need a "continue" message — when the user clicks the button, the server writes transcript_review.txt to the project dir AND exits with code 0. The agent's background-task notification fires, and the pipeline resumes from step 8.

    Fallback if no browser / no server: open the editor as a static file (start "" "$HOME/.claude/skills/video-edit/transcript-editor/index.html"), ask the user to pick the project folder, edit, save transcript_review.txt back into the project, and reply "continue". The editor supports both modes.

  7. (Optional) Background removal — see step 7 below; can run in parallel with the user's review.

  8. After approval, run python references/apply_review.py. It re-tokenises edited lines and redistributes word timings back into transcript.json so caption sync still works.

  9. (Optional) Background removal — if any talking-head segment needs behind-subject text, extract the segment as outro.mp4 (or intro.mp4) and run npx hyperframes remove-background <clip>.mp4 -o <name>_subject.webm --quality best. CPU only on most setups (~3–8 min for a ~15s 1440p clip).

  10. Re-encode source with dense keyframes — multi-worker render seeks freeze on sparse keyframes. Always run:

    ffmpeg -y -i source.mp4 -c:v libx264 -preset medium -crf 18 -r 30 -g 30 -keyint_min 30 -sc_threshold 0 -pix_fmt yuv420p -movflags +faststart -c:a copy footage.mp4
    
  11. Re-load the (edited) transcript and generate the body sub-composition via references/gen_body.py. The generator emits the full compositions/components/caption-body.html with editorial + matrix alternating in liquid-glass pills, anchored lower-left-of-centre (clears bottom-right webcam PiPs).

  12. Wire the host index.html from references/host-template.html. Layer order (z-index, NOT track-index):

    • z0: footage .cam-bg
    • z1: liquid blob background (compositions/liquid-blobs.html, mix-blend-mode: screen, full duration)
    • z2: parallax behind-subject caption (intro and/or outro, when bg-removal used)
    • z3: subject cut-out .cam-out / .cam-sub (with matching data-media-start)
    • z6: body captions
    • z46: progress bar + flash + liquid morph wipe
  13. Lint — npx hyperframes lint. Must be 0 errors. Common fixes: GSAP/CSS transform conflict on the wipe element (use xPercent/yPercent or remove the CSS transform); overlapping tweens on the same property (add overwrite: "auto").

  14. Render — npx hyperframes render --quality standard --fps 30 --output renders/<name>_FINAL.mp4. Standard is the right delivery target — high roughly doubles render time. Verify with 6–8 spot-check frames from across the timeline before reporting done.

Vertical (9:16) output for TikTok / Reels / Shorts

When the user asks for vertical / portrait / TikTok / Reels / 9:16 output (from a 16:9 source):

  1. Clone the project to a sibling folder: cp -r project/ project-vertical/.
  2. Replace its index.html with references/host-template-vertical.html (1080×1920 canvas, blurred-bg backdrop with liquid blobs, the 16:9 footage as a centered horizontal strip, captions below).
  3. Replace its gen_body.py with references/gen_body_vertical.py (centered pill, larger fonts, narrower max-width), then re-run it to emit compositions/components/caption-body.html.
  4. Drop the behind-subject cut-out + parallax sub-compositions (the cutout is aligned for 16:9; not worth re-aligning for v1). The vertical comp uses the blurred-source backdrop + blobs for atmosphere instead.
  5. Update data-duration to the actual video duration. Update the brand-chip text in index.html (YUV.AI by default).
  6. Lint + render — same commands. Output is 1080×1920. Drop straight onto TikTok / IG Reels / YT Shorts.

To deliver both 16:9 and 9:16 in one go, run two render commands (in parallel projects). The transcript_review.txt approval applies to both — same captions, two compositions.

Critical rules

  • Never render the final without explicit transcript approval. The review step is the whole point.
  • For Hebrew: large-v3 + language="he" + direction: rtl + Rubik (700 + 900 for editorial dual-weight emphasis).
  • Caption pills always need an opaque dark backing — bare light text vanishes on white app UI.
  • Centre caption pills horizontally but shift the centre x-coord left (e.g. left: 720px) when the footage has a bottom-right webcam PiP.
  • The behind-subject cut-out clip MUST carry data-media-start matching its data-start (or matching the offset from the source if the clip was extracted), or the cut-out plays from frame 0 and desyncs.
  • The remove-background webm keeps the original RGB and writes only the alpha mask — ffprobe reports yuv420p, which looks like "no alpha". Confirm via TAG:ALPHA_MODE=1 or composite over a solid colour.
  • Outro/end cards with burned-in text — do NOT caption over them; they collide.

File references

FilePurpose
transcript-editor/index.htmlInteractive browser editor — video preview, RTL editing, dictionary apply, optional WebLLM AI suggestions, saves transcript_review.txt
references/setup.mdPrerequisites + install commands for Node / Python / FFmpeg / faster-whisper
references/transcribe.pyfaster-whisper transcribe with CPU fallback + word timestamps
references/serve_review.pyLocal review server — auto-loads editor, blocks until user clicks Approve & Render, then writes transcript_review.txt and exits (signals the agent)
references/make_review.pyApply corrections + emit transcript_review.txt (file-mode fallback)
references/apply_review.pyParse edited review file, redistribute word timings, update transcript.json
references/gen_body.pyCaption-body generator (editorial + matrix in liquid-glass pills)
references/host-template.html16:9 host composition with liquid effects + transition wipe
references/host-template-vertical.html9:16 host (1080×1920) — TikTok / Reels / Shorts layout: blurred bg, centered 16:9 footage strip, captions below, brand chip top-right
references/gen_body_vertical.pyCaption-body generator tuned for vertical (centered pill, larger fonts, narrower max-width)
references/liquid-blobs.htmlFull-duration drifting blob layer
references/caption-parallax-outro.htmlBehind-subject caption template (English; clone for other languages)
references/corrections-hebrew.mdKnown Hebrew Whisper mishears
references/transcript-review-workflow.mdThe pause/approve step in detail

Individual skills in this repo

This repo contains 9 individual skills — each has its own dedicated page.

hoodini/ai-agents-skills

Edit any selfie or screen-share footage into a viral short-form video in YUV.AI's signature style — Apple-style liquid-glass cards (real CSS backdrop-filter), dark-mode polish, MrBeast-paced cuts, video-title karaoke captions, premium GSAP motion graphics, no fake content, never covering the speaker's face. Hebrew is rendered in Rubik Black, English in Anton uppercase. Always renders BOTH 9:16 and 16:9 and always saves with _V<N> suffix for backups. Trigger when the user drops a path to an .mp4/.mov/.mkv and says "edit this", "make it viral", "turn this into a short", or any Hebrew equivalent (ערוך סרטון, סרטון ויראלי, להפוך לוויראלי, ריל, שורט). The pipeline is the COMBINATION of two skills: video-use (transcription + word-snapped cuts + base extraction) and hyperframes (HTML/CSS/GSAP visual composition + render). Do NOT use for podcast-only audio edits.

hoodini/ai-agents-skills

Yuval Avidani's YUV.AI brand and design system. Apply ONLY when YUV.AI-branded output is requested — presentations, decks, keynotes, portfolio, brand site, profile, speaker bio, brand assets, or any prompt mentioning YUV.AI / "my brand" / "my deck" / "my site" / "for me". Do NOT auto-trigger on generic "build a game / web app / dashboard / landing page" without YUV.AI context — those use whatever palette fits. Three modes; NEON (hot pink #FF1464 + neon cyan #00E5FF + white, DEFAULT for YUV.AI web, apps, games, dashboards, social, general visual work); DECKS (Fly High purple/yellow/grey, presentations and slides ONLY); Warm Editorial (pink/yellow/bone for Hope, Marcus, bigcats.ai, practical.yuv.ai). Universal Fly High throughline across all modes; flight/progress motifs (HUD strips, dials, "Let's Fly High" tagline, phoenix mark). Anton + Inter (EN), Rubik + Assistant (HE), letter-spacing 0 default. Bundled brand assets, canonical socials, credentials. Project brand wins if specified.

hoodini/ai-agents-skills

Yuval's all-in-one AI video pipeline. Turns an idea/script into a finished, on-brand MP4 by orchestrating HyperFrames (HTML→deterministic video render), Lottie (branded motion graphics), ManimCE (math / neural-network / concept animations), and a transcribe→approve caption flow — all wrapped in the YUV.AI Neon Phoenix brand via a frame.md. Use whenever Yuval wants to make, edit, or explain something as a video: promo, explainer, launch, social reel, "make a video about X", "explain X as a video", "neural network animation", "turn this into a video", captioned tutorial, 16:9 or 9:16. Triggers: video, explainer, promo, reel, manim, lottie, hyperframes, animation, "make a video", "explain ... as a video", מצגת וידאו, סרטון, הסבר וידאו. Routes each beat to the right engine, wraps in brand, self-verifies, and renders.

hoodini/ai-agents-skills

Build a premium cinematic landing page with mouse-scrub video hero and brand-driven narrative-arc sections. Use whenever the user provides a hero video plus a product / subject / brand and wants a landing page, promo site, product showcase, marketing page, or storytelling site. Works for any language (RTL or LTR — Hebrew, English, Arabic, Spanish, French, Japanese, etc.) and any subject (food, tech, animals, fashion, services, SaaS, wildlife campaigns, music releases, books, real estate). The signature effect is mouse-driven video scrubbing — the hero video lives across the entire page as a fixed backdrop, and moving the mouse left-right scrubs the video timeline so the subject responds to the cursor. Below the hero, 4-5 fully-opaque sections each carry their own brand identity (color, typography emphasis, layout pattern) and walk the viewer through a narrative arc (e.g. longing → joy → nostalgia → contemplation → action). Triggers on phrases like "build a landing page from this video", "cinematic landing ...

hoodini/ai-agents-skills

Build a scroll-driven cinematic landing page from a short video. The user provides a 5–15 second video (often AI-generated); this skill extracts every frame at HD JPEG quality, then produces a single-hero HTML page where the user's scroll gesture scrubs the frames in place (the page itself never scrolls) and 5 dramatic text overlays crossfade in/out — Google Anton headlines, Caveat handwritten accents, locked body, virtual scroll. Use this skill whenever the user wants to "turn this video into a landing page", "make a scroll-scrub landing page", "build a parallax hero from this clip", "add a new landing page to the parasites showcase", "do the same as github/lion/hope for this new video", or any variant that pairs a short clip with dramatic scroll-triggered storytelling. Trigger even if the user only says "use my video for a landing page" — that is this skill.

hoodini/ai-agents-skills

Turn any video into a cinematic scroll-driven landing page — Apple-style hero where scrolling progresses the visible frame through the video. Use when the user provides a video file and asks for "a landing page from this video", "scroll-frame website", "Apple-style scroll site", "hero that scrubs the video", "like the GitHub Copilot landing", or any equivalent. Extracts N evenly-spaced frames via ffmpeg, builds a self-contained HTML page with a sticky hero + JS scroll listener that swaps the visible frame as you scroll, plus headline, sections and CTA below. For YUV.AI projects, applies the yuv-design-system skill in Neon mode (pink/cyan/white, default for YUV.AI web) — Decks (purple/yellow) is reserved for slides only. For generic / non-YUV.AI projects, picks an appropriate palette per the source video. Output is one folder with `index.html` and a `frames/` directory — drop on any static host.

hoodini/github-trending

Fetch and display GitHub trending repositories and developers. Use when building dashboards showing trending repos, discovering popular projects, or tracking GitHub trends. Triggers on GitHub trending, trending repos, popular repositories, GitHub discover.

hoodini/mongodb

Work with MongoDB databases using best practices. Use when designing schemas, writing queries, building aggregation pipelines, or optimizing performance. Triggers on MongoDB, Mongoose, NoSQL, aggregation pipeline, document database, MongoDB Atlas.

hoodini/owasp-security

Implement secure coding practices following OWASP Top 10. Use when preventing security vulnerabilities, implementing authentication, securing APIs, or conducting security reviews. Triggers on OWASP, security, XSS, SQL injection, CSRF, authentication security, secure coding, vulnerability.

Habilidades Relacionadas