Communitygithub.com

oyi77/1ai-skills

Add captions/subtitles to talking-head or launch videos using HyperFrames. Use when the user wants embedded captions, verbatim caption rails, or cinematic text behind the subject. Covers the drop/rail/embed caption model, the overlay law (captions are NOT a reserved band), matte occlusion for embedded climaxes, and word-timestamp alignment from TTS output.

O que é 1ai-skills?

1ai-skills is a Claude Code agent skill that add captions/subtitles to talking-head or launch videos using HyperFrames. Use when the user wants embedded captions, verbatim caption rails, or cinematic text behind the subject. Covers the drop/rail/embed caption model, the overlay law (captions are NOT a reserved band), matte occlusion for embedded climaxes, and word-timestamp alignment from TTS output.

Funciona com~Claude Code~Codex CLI~Cursor
npx skills add https://github.com/oyi77/1ai-skills/tree/HEAD/content/hyperframes-embedded-captions

Perguntar na sua IA favorita

Abre um novo chat com esta habilidade de agente já pré-carregada.

Documentação

Overview

This skill adds burned-in captions to an existing video via a HyperFrames overlay composition. Use it when your source footage needs accessible, styled subtitles without a separate subtitle track. It renders the captioned MP4 deterministically from the HTML overlay.

HyperFrames — Embedded Captions

Add captions/subtitles to talking-head or launch videos. The caption model (drop/rail/embed) and the overlay law (captions are NOT a reserved band).

When to Use

  • User wants captions/subtitles on a talking-head video
  • User wants a verbatim caption rail (standard lower-third)
  • User wants cinematic embedded text behind the subject (the climax beat)
  • User wants to add accessibility to a video
  • User wants to burn in captions for social media (no subtitle track)

When NOT to Use

  • User wants a full video workflow (use hyperframes-product-launch-video)
  • User wants a full explainer (use hyperframes-faceless-explainer)
  • User wants captions as a sidecar file (SRT/VTT) — this skill burns them in

The Caption Model — Drop / Rail / Embed

WhatHow it's shown
dropfiller — um/uh, stuttersnot shown
railthe default — ordinary spoken contentclean lower-third subtitle, in front
embeda promoted peak — the headline beatone big word behind the subject (matte occlusion)

The rail carries most of the text; embed is the scarce, earned peak — ≤1 per beat, never two adjacent, spaced ≥ a beat apart.

The Overlay Law

A caption line is composited ON TOP of the film as an overlay; it is NOT a reserved zone.

  • Center the composition on the TRUE vertical center (y = H/2). Do not shift content up to "make room" for captions.
  • Content may extend to the canvas bottom.
  • Avoid parking critical small readable text (URL, legal line) in the bottom ~80px center span where the caption line sits.
  • No machine keep-out gate — judge legibility visually.

Workflow

1 · Generate voiceover + word timestamps

# Using TTS with word-level timestamps
node scripts/heygen-tts.mjs ./vo-spoken.txt -o voiceover.mp3 --words vo-words.json --voice <voice_id>

2 · Align captions to display tokens

node scripts/align-captions.mjs --tokens script-tokens.json --words vo-words.json --out captions.json

captions.json is the caption-rail input (display spelling, spoken timing).

3 · Build the caption rail

const LINES = /* contents of captions.json */ [
  { id: 0, end: 2.74, w: [["This", 0.0], ["week,", 0.30], ...] },
  ...
];

4 · Verify caption presence

Sample 3-4 frames across the VO's spoken window and confirm the caption rail renders visible text on each. If any frame in a spoken interval is missing captions, the build ships uncaptioned — treat as a red gate.

Anti-Rationalization Table

RationalizationReality
"I'll center at 42% to leave room for captions"True center (y = H/2) — captions are overlay
"I'll embed every word"Embed is scarce — rail carries most text
"I'll skip word timestamps"Audio is the clock — all beat times come from word timestamps
"I'll use phonetic spelling in captions"Captions always render the display layer
"I'll guess the pronunciation"Ask the user, then grow the lexicon

Verification

  • Caption rail present on 3+ sampled frames
  • Display spelling matches script (not phonetic)
  • Word timestamps align with audio
  • Composition centered on true center (y = H/2)
  • No critical text in caption overlay zone

Individual skills in this repo

This repo contains 9 individual skills — each has its own dedicated page.

oyi77/1ai-skills

Turn a weekly changelog .md into a finished branded changelog video using HyperFrames. Use when the user provides a changelog/digest markdown and wants the weekly video, or says 'changelog video'. Self-contained — fonts, background, lexicon, and scripts ship in the skill. Produces square 1080, ~45-60s videos with animated brand mock-UIs.

oyi77/1ai-skills

The transition technique catalog for HyperFrames videos — five velocity-matched SEAMS (zoom-through, inverse zoom-through, cut-the-curve, waterfall cut, rack-focus blur-cut) plus in-scene techniques (waterfall entry, nudge curve). Covers partial-travel velocity matching, Z scale-sign rule, size-scaled blur, the 10/65/25 slide ratio, and the 2-3 transition budget. Read before authoring any transition, text-beat handoff, or kinetic text entry.

oyi77/1ai-skills

Turn arbitrary text — an article, notes, a topic, a brief — into a faceless explainer video using HyperFrames. Use when the user wants a topic explainer, concept breakdown, how-to, or listicle with no product/website to capture. Every visual is invented per scene (typography, abstract graphics, diagrams, data-viz). Not a video built from a website (use hyperframes-product-launch-video).

oyi77/1ai-skills

The fallback HyperFrames workflow for anything that doesn't fit other skills — longer multi-scene pieces, brand/sizzle reels, title cards, static loops, freeform compositions. Input- and length-agnostic. Also the home of companion mode (co-create with the full toolbox). Use when no other HyperFrames skill matches, or when the user wants a freeform video.

oyi77/1ai-skills

Create short, unnarrated, design-led motion graphics (~under 10s) using HyperFrames. Use when the user wants kinetic type, stat/chart hits, logo stings, lower-thirds, animated tweets/headlines, or transparent overlay motion graphics. MP4 or transparent overlay output. Teaches seek-safe keyframe authoring and the 2-3 transition budget rule.

oyi77/1ai-skills

Turn a music track into a beat-synced HyperFrames video — lyric, slideshow, or kinetic promo. Use when the user provides an audio file, a video to pull audio from, or a mood brief for generated music. Music drives pacing. Teaches beat-map extraction, section-based scene planning, and audio-reactive visuals.

oyi77/1ai-skills

Create product launch videos from a URL, brief, or script using HyperFrames (HTML→MP4). Use when the user wants a product launch video, site tour, social clip featuring a product's own visuals, or marketing/promoting a product. Up to ~3 min (sweet spot 30-90s). Teaches the HyperFrames production loop: plan → write HTML → wire seekable animations → add media → lint → preview → render.

oyi77/1ai-skills

The Video Production Manager — manages ALL video production under 1ai-content. Delegates to HyperFrames, Remotion, AI video, and FFmpeg based on the task. Has memory via 1ai-hub brain, follows 1ai-rules, and reports to Content Director. Use when the user says 'make a video', 'create video', 'edit video', 'video strategy', or needs any video production work.

oyi77/1ai-skills

Agent skill at content/video/gen/SKILL.md

Habilidades Relacionadas