Communitygithub.com

NeverSight/learn-skills.dev

Use this skill when creating complete videos from scratch - product demos, explainers, social clips, or announcements. Orchestrates the full workflow: deep interview, script generation, visual verification, Remotion project build, audio design, narration, and 4K rendering. Triggers on "make me a video", "create a video about", video production, and end-to-end video creation requests.

learn-skills.dev 是什么?

learn-skills.dev is a Claude Code agent skill that use this skill when creating complete videos from scratch - product demos, explainers, social clips, or announcements. Orchestrates the full workflow: deep interview, script generation, visual verification, Remotion project build, audio design, narration, and 4K rendering. Triggers on "make me a video", "create a video about", video production, and end-to-end video creation requests.

兼容平台✓Claude Code✓Codex CLI~Cursor✓Gemini CLI
npx skills add https://github.com/NeverSight/learn-skills.dev/tree/HEAD/data/skills-md/absolutelyskilled/absolutelyskilled/video-creator

在你喜欢的 AI 中提问

打开一个已预加载此 Agent Skill 的新对话。

文档

When this skill is activated, always start your first response with the :movie_camera: emoji.

Video Creator

This is the orchestrator skill for end-to-end video creation. It coordinates a complete 7-step workflow that takes a user from "I need a video" to a finished 4K render. Each step delegates to companion skills (remotion-video, video-scriptwriting, video-audio-design, video-analyzer) while this skill manages sequencing, approval gates, and handoffs between stages.

You do NOT need to know Remotion internals or audio engineering - those are handled by companion skills. This skill's job is to run the process, ask the right questions, and make sure nothing gets skipped.


When to use this skill

Trigger this skill when the user:

  • Says "make me a video", "create a video about", or "I need a video for"
  • Wants a product demo, explainer, social clip, or announcement video
  • Asks for end-to-end video production from concept to render
  • Needs help turning an idea into a finished video file
  • Mentions wanting a programmatic video with Remotion
  • Has a reference video and wants to create something similar

Do NOT trigger this skill for:

  • Editing an existing video file (use video-analyzer or ffmpeg skills)
  • Writing a script only without producing a video (use video-scriptwriting)
  • Audio-only production like podcasts (use video-audio-design for music/SFX)
  • Remotion coding questions without a production context (use remotion-video)
  • Analyzing or reviewing a video (use video-analyzer)

Key principles

  1. Visual-first: Get visuals approved before spending money on audio. Narration costs real dollars via ElevenLabs - never generate it until the visual layer is locked.

  2. Interview exhaustively: Ask up to 30 questions to capture all context about the product, audience, goals, tone, assets, and visual preferences. Incomplete context leads to expensive re-work.

  3. Structured handoff: The video-script.yaml file is the single source of truth between all steps. Every scene, frame count, narration line, and SFX cue lives in this file.

  4. User approval gates: Explicit approval is required at each major step. Never auto-advance. The user must say "approved" or "looks good" before the next step begins.

  5. 4K default: Always render at 3840x2160 unless the user specifies otherwise. Downscaling is easy; upscaling is not.


The 7-Step Workflow

Step 0: Ensure Remotion Project Exists (prerequisite)

Before anything else, the user must have a Remotion project set up. This is the workspace where all video code will be written. Check if you are already inside a Remotion project (look for remotion.config.ts or @remotion/cli in package.json). If not, scaffold one:

npx create-video@latest

This creates a new folder with the starter project. Then install dependencies:

cd <project-name>
npm install

Do NOT proceed to Step 1 until a Remotion project exists and dependencies are installed. All subsequent steps write code into this project.

If the user already has a Remotion project, cd into it and continue.


Step 1: Deep Interview

Gather all context needed to write a complete script. Ask questions one at a time using an interactive conversational approach.

Question categories (aim for 15-30 questions total):

CategoryExample Questions
Product/subjectWhat does the product do? What is the core value prop?
AudienceWho is the target viewer? Technical level?
Video goalsWhat should the viewer do after watching?
Tone/stylePlayful or professional? Fast or slow-paced?
AssetsDo you have logos, screenshots, brand colors, fonts?
ContentWhat are the key features or messages to cover?
Visual preferencesAny reference videos? Preferred animation style?
DurationHow long should the video be? Where will it be published?

Rules for Step 1:

  • Ask one question at a time - do not dump a questionnaire
  • If the user provides a reference video, analyze it FIRST using the video-analyzer skill before asking questions
  • Summarize what you know after every 5-8 questions
  • Do not proceed until you have enough context for every scene
  • End with: "I have enough context to write the script. Ready to proceed?"

Exit criteria: User confirms you have enough context.


Step 2: Generate Script (YAML)

Use patterns from the video-scriptwriting skill to produce a structured YAML script.

Frame count formula:

frames = duration_seconds * 30

All Remotion compositions use 30fps by default.

Generate a video-script.yaml file with this structure:

title: "Product Demo - Acme Widget"
fps: 30
resolution: { width: 3840, height: 2160 }
total_duration_seconds: 60
scenes:
  - id: scene-01
    title: "Hook"
    duration_seconds: 5
    frames: 150
    narration: "Ever wished your widgets could think for themselves?"
    visuals: "Dark background, glowing Acme logo fades in, particles"
    animation_notes: "Logo scale 0->1 with spring, particles emit from center"
    music_cue: "Ambient synth pad, low energy"
    sfx: "Subtle whoosh on logo reveal"
    transition_out: "crossfade"
  - id: scene-02
    # ... continue for all scenes

Rules for Step 2:

  • Every scene must have duration_seconds, frames, narration, visuals, animation_notes, music_cue, and sfx fields
  • Frame counts must equal duration_seconds * 30
  • Total scene durations must sum to total_duration_seconds
  • Present the full script to the user for review
  • Iterate on feedback until the user explicitly approves

Exit criteria: User approves the script.


Step 3: Visual Verification

Build a minimal Remotion project with visuals only - no audio layer yet.

What to build:

  • Remotion compositions for each scene (visuals, animations, typography, colors)
  • Scene transitions matching the script
  • Correct frame counts per scene
  • Basic layout and spacing at 4K resolution

What NOT to build yet:

  • Audio integration
  • Narration sync
  • Volume ducking
  • Final render pipeline

Process:

  1. Verify the Remotion project is set up (done in Step 0)
  2. Build visual compositions for each scene inside the project
  3. Launch Remotion Studio: npx remotion studio
  4. Tell the user to preview at http://localhost:3000
  5. Iterate on visual feedback (colors, timing, animations, transitions)
  6. Get EXPLICIT visual approval before proceeding

Exit criteria: User explicitly approves the visuals.


Step 4: Build Full Remotion Project

Now that visuals are approved, flesh out the complete project.

Tasks:

  • Finalize all compositions with polished animations
  • Wire up scene-to-scene transitions
  • Add Zod schemas for parametrization (titles, colors, durations)
  • Ensure frame counts match the script exactly
  • Organize project structure cleanly:
src/
  compositions/
    Scene01Hook.tsx
    Scene02Feature.tsx
    ...
  components/
    AnimatedLogo.tsx
    TransitionWipe.tsx
    ...
  Root.tsx
  index.ts
  video-script.yaml

Rules for Step 4:

  • Each scene should be its own composition file
  • Shared components go in components/
  • All magic numbers should be replaced with Zod schema props
  • Frame counts must still match video-script.yaml

Exit criteria: Full project builds without errors and matches approved visuals.


Step 5: Add Background Audio + SFX

Use patterns from the video-audio-design skill.

Tasks:

  1. Source or select background music (user-provided or from documented sources)
  2. Place SFX at trigger points from the script (clicks, typing, whooshes, etc.)
  3. Set base volume levels for music and SFX
  4. Implement ducking infrastructure - prepare volume curves that will lower music during narration segments in Step 6
  5. Preview audio-visual sync in Remotion Studio

Volume guidelines:

LayerBase VolumeDuring Narration
Background music0.3-0.50.1-0.15
SFX0.5-0.80.4-0.6
NarrationN/A (Step 6)1.0

Exit criteria: User approves audio-visual sync.


Step 6: Add Narration (Deferred - costs money)

This step involves paid API calls to ElevenLabs. Always confirm before proceeding.

Setup:

  1. Check if user has an ElevenLabs API key
  2. If not, guide them through signup at https://elevenlabs.io
  3. Ask voice preference questions:
    • Gender preference?
    • Age range (young, middle, mature)?
    • Accent preference?
    • Energy level (calm, moderate, energetic)?
    • Warmth (warm and friendly, neutral, authoritative)?

Process:

  1. Generate narration audio for each scene via ElevenLabs API
  2. Calculate exact audio durations from generated files
  3. Adjust frame counts if narration is longer/shorter than planned
  4. Update video-script.yaml with actual audio durations
  5. Sync narration with visual timing
  6. Activate volume ducking on background music during narration segments
  7. Preview complete audio mix in Remotion Studio

Rules for Step 6:

  • Always confirm costs before making API calls
  • If user wants to skip narration, that is fine - proceed to Step 7 without it
  • If user prefers a different TTS provider, support that
  • Re-sync all timing if audio durations differ from script estimates

Exit criteria: User approves narration sync (or explicitly skips this step).


Step 7: Final Preview + 4K Render

Process:

  1. Launch full preview with all layers in Remotion Studio
  2. Tell user to review the complete video (visuals + music + SFX + narration)
  3. Get final approval
  4. Render at 4K:
npx remotion render src/index.ts Main out/video.mp4 \
  --width 3840 --height 2160

Format guidance:

FormatUse Case
MP4 (H.264)Web, social media, general sharing
MP4 (H.265)Smaller file size, modern device playback
ProResFurther editing in Premiere, Final Cut, DaVinci
WebM (VP9)Web embedding with transparency support

Exit criteria: Rendered video file delivered to user.


Orchestration rules

  1. Each step requires explicit user approval before advancing to the next step
  2. Explain what you are doing and why at each step
  3. If the user provides a reference video, run video-analyzer FIRST before starting Step 1
  4. Always create a video-script.yaml as the single source of truth
  5. The Remotion project structure must be clean and well-organized
  6. If the user wants to skip narration (Step 6), proceed directly to Step 7
  7. If the user wants a different TTS provider, adapt Step 6 accordingly
  8. Never generate narration audio before visual approval (Step 3 complete)
  9. Never skip visual verification - it prevents expensive re-work
  10. Always default to 4K (3840x2160) unless the user specifies otherwise

Supported video types

TypeDurationScenesPer Scene
Product demo30-120s6-155-10s
Explainer60-180s8-205-12s
Social clip15-60s3-83-8s
Announcement15-45s3-64-8s

Anti-patterns / common mistakes

Anti-patternWhy it failsCorrect approach
Generating narration before visual approvalWastes money if visuals changeComplete Steps 1-4 first, then narration
Skipping the interviewScript lacks context, scenes feel genericAsk 15-30 targeted questions
No YAML script fileNo source of truth, timing driftsAlways generate video-script.yaml
Auto-advancing without approvalUser loses control, re-work is expensiveWait for explicit "approved" at each gate
Hardcoding values in RemotionCannot parametrize or reuseUse Zod schemas for all configurable values
Rendering at 1080p by defaultCannot upscale laterAlways default to 4K (3840x2160)
Building audio and visuals simultaneouslyMakes iteration painfulVisual-first, audio after approval

Gotchas

  1. Frame count math matters. 60 seconds at 30fps = 1800 frames exactly. If scene durations do not sum correctly, the video will be too short or too long. Always verify: sum(scene.duration_seconds) == total_duration_seconds.

  2. Remotion Studio needs a running dev server. After building compositions, you must run npx remotion studio and tell the user to open http://localhost:3000. Do not assume they can preview without this.

  3. ElevenLabs narration durations are unpredictable. A 5-second script line might produce 4.2s or 6.1s of audio. Always measure actual durations after generation and adjust frame counts accordingly.

  4. Volume ducking requires knowing narration timing. The ducking curves for background music depend on knowing exactly when narration starts and stops in each scene. This is why Step 5 sets up ducking infrastructure but Step 6 activates it.

  5. 4K rendering is slow. A 60-second video at 3840x2160 can take 10-30 minutes depending on complexity. Warn the user about render time before starting, and suggest a 1080p test render first if they want a quick check.


References

  • references/workflow-checklist.md - Detailed checklist for each of the 7 steps with sub-tasks, expected outputs, and approval criteria
  • references/project-templates.md - Starter Remotion project structures for each video type with file trees and component scaffolds
  • references/troubleshooting.md - Common issues and fixes for Remotion rendering, audio sync, FFmpeg, ElevenLabs API, and performance

Companion check

On first activation of this skill in a conversation: check which companion skills are installed by running ls ~/.claude/skills/ ~/.agent/skills/ ~/.agents/skills/ .claude/skills/ .agent/skills/ .agents/skills/ 2>/dev/null. Compare the results against the recommended_skills field in this file's frontmatter. For any that are missing, mention them once and offer to install:

npx skills add AbsolutelySkilled/AbsolutelySkilled --skill <name>

Skip entirely if recommended_skills is empty or all companions are already installed.

Individual skills in this repo

This repo contains 20 individual skills — each has its own dedicated page.

NeverSight/learn-skills.dev

Use when generating or modifying Remotion video code, creating demo videos, or working with the demo-video/ directory

NeverSight/learn-skills.dev

Create AI avatar and talking head videos via inference.sh CLI. Recommended: P-Video-Avatar (fastest, cheapest, built-in TTS). Also: OmniHuman, Fabric, PixVerse. Audio: Inworld TTS-2 (100+ languages, emotion steering for characters), ElevenLabs, Kokoro. Capabilities: audio-driven avatars, text-to-avatar, lipsync videos, talking head generation, virtual presenters, UGC content. Use for: AI presenters, explainer videos, virtual influencers, dubbing, marketing videos, UGC ads, gaming avatars, NPC dialogue. Triggers: ai avatar, talking head, lipsync, avatar video, virtual presenter, ai spokesperson, audio driven video, heygen alternative, synthesia alternative, talking avatar, lip sync, video avatar, ai presenter, digital human, ugc, ugc video, ugc ad, avatar ugc

NeverSight/learn-skills.dev

Create AI marketing videos for ads, promos, product launches, and brand content. Models: Veo, Seedance, Wan, FLUX for visuals, Kokoro for voiceover. Types: product demos, testimonials, explainers, social ads, brand videos. Use for: Facebook ads, YouTube ads, product launches, brand awareness. Triggers: marketing video, ad video, promo video, commercial, brand video, product video, explainer video, ad creative, video ad, facebook ad video, youtube ad, instagram ad, tiktok ad, promotional video, launch video

NeverSight/learn-skills.dev

Generate AI videos with Google Veo, Seedance 2.0, HappyHorse, Wan, Grok and 40+ models via inference.sh CLI. Models: Veo 3.1, Veo 3, Seedance 2.0, HappyHorse 1.0, Wan 2.5, Grok Imagine Video, OmniHuman, Fabric, HunyuanVideo. Capabilities: text-to-video, image-to-video, reference-to-video, video editing, lipsync, avatar animation, video upscaling, foley sound. Use for: social media videos, marketing content, explainer videos, product demos, AI avatars. Triggers: video generation, ai video, text to video, image to video, veo, animate image, video from image, ai animation, video generator, generate video, t2v, i2v, ai video maker, create video with ai, runway alternative, pika alternative, sora alternative, kling alternative, seedance, happyhorse

NeverSight/learn-skills.dev

ElevenLabs automatic dubbing - translate and dub audio/video into 29 languages while preserving speaker voice via inference.sh CLI. Capabilities: auto speaker detection, voice-preserving translation, video dubbing, audio localization. Use for: content localization, video translation, multilingual content, international distribution. Triggers: dubbing, dub video, translate audio, video translation, audio translation, localize content, elevenlabs dubbing, eleven labs dub, multilingual dub, voice translation, auto dub, language dub, content localization

NeverSight/learn-skills.dev

Explainer video production guide: scripting, voiceover, visuals, and assembly. Covers script formulas, pacing rules, scene planning, and multi-tool pipelines. Use for: product demos, how-it-works videos, onboarding videos, social explainers. Triggers: explainer video, how to make explainer, product video, demo video, video production, video script, animated explainer, product demo video, tutorial video, onboarding video, walkthrough video, video pipeline

NeverSight/learn-skills.dev

Still-to-video conversion guide: model selection, motion prompting, and camera movement. Covers Wan 2.5 i2v, Seedance, Fabric, Grok Video with when to use each. Use for: animating images, creating video from stills, adding motion, product animations. Triggers: image to video, i2v, animate image, still to video, add motion to image, image animation, photo to video, animate still, wan i2v, image2video, bring image to life, animate photo, motion from image

NeverSight/learn-skills.dev

Generate talking head avatar videos with Pruna P-Video-Avatar via inference.sh CLI. Turn a portrait image into a realistic speaking video with built-in TTS. 18x faster and 6x cheaper than competitors. Models: P-Video-Avatar, P-Image (for portrait generation). Capabilities: text-to-avatar, audio-driven avatars, 30 voices, 10 languages, 720p/1080p, built-in TTS, dynamic backgrounds, full-body control. Use for: AI presenters, product demos, explainer videos, virtual influencers, marketing, education, multilingual content, UGC, gaming avatars. Triggers: avatar video, talking head, ai avatar, p-video-avatar, pruna avatar, video avatar, ai presenter, digital human, virtual presenter, lipsync, talking avatar, ai spokesperson, heygen alternative, synthesia alternative, veed alternative, fabric alternative, omnihuman alternative

NeverSight/learn-skills.dev

Generate videos with Pruna P-Video and WAN models via inference.sh CLI. Models: P-Video, WAN-T2V, WAN-I2V. Capabilities: text-to-video, image-to-video, audio support, 720p/1080p, fast inference. Pruna optimizes models for speed without quality loss. Triggers: pruna video, p-video, pruna ai video, fast video generation, optimized video, wan t2v, wan i2v, economic video generation, cheap video generation, pruna text to video, pruna image to video

NeverSight/learn-skills.dev

Render videos from React/Remotion component code via inference.sh. Pass TSX code, get MP4. Supports all Remotion APIs: useCurrentFrame, useVideoConfig, spring, interpolate, AbsoluteFill, Sequence. Configurable resolution, FPS, duration, codec. Use for: programmatic video generation, animated graphics, motion design, data-driven videos, React animations to video. Triggers: remotion, render video from code, tsx to video, react video, programmatic video, remotion render, code to video, animated video, motion graphics code, react animation video

NeverSight/learn-skills.dev

Video ad creation with exact platform-specific specs for TikTok, Instagram, YouTube, Facebook, LinkedIn. Covers dimensions, duration limits, AIDA framework, and caption requirements. Use for: video ads, social media ads, paid media creative, video marketing, ad production. Triggers: video ad, social media ad, tiktok ad, instagram ad, youtube ad, facebook ad, linkedin ad, video creative, ad specs, paid media, video marketing, ad production, reels ad, stories ad, pre roll, bumper ad

NeverSight/learn-skills.dev

Best practices and techniques for writing effective AI video generation prompts. Covers: Veo, Seedance, Wan, Grok, Kling, Runway, Pika, Sora prompting strategies. Learn: shot types, camera movements, lighting, pacing, style keywords, negative prompts. Use for: improving video quality, getting consistent results, professional video prompts. Triggers: video prompt, how to prompt video, veo prompts, video generation tips, better ai video, video prompt engineering, video prompt guide, video prompt template, ai video tips, video prompt best practices, video prompt examples, cinematography prompts

NeverSight/learn-skills.dev

YouTube thumbnail design with specific dimensions, contrast rules, and mobile preview optimization. Covers safe zones, text placement, face expression psychology, and A/B testing. Use for: YouTube thumbnails, video cover images, click-through optimization. Triggers: youtube thumbnail, thumbnail design, video thumbnail, click through rate, ctr optimization, youtube cover, video cover image, thumbnail maker, thumbnail tips, youtube design, video preview image

NeverSight/learn-skills.dev

Landing page conversion optimization with layout rules, hero section design, and CTA psychology. Covers above-the-fold formula, social proof placement, mobile design, and F-pattern reading. Use for: startup landing pages, product pages, SaaS marketing, conversion optimization. Triggers: landing page, hero section, above the fold, conversion optimization, landing page design, cta button, hero image, landing page layout, saas landing page, product page design, conversion rate, landing page best practices

NeverSight/learn-skills.dev

Configure and use the hosted YouTube Data MCP end-to-end with minimal user input. Use when users want the agent to verify Node.js and `npx`, configure MCP server config (Windows/macOS, Cursor/Codex/OpenClaw/OpenCode), request API key at setup time, run post-install capability discovery (`tools/list` and `get_patch_notes`), and then strongly recommend helper skill and Python setup for full local document and spreadsheet workflows.

NeverSight/learn-skills.dev

Creates 120fps GPU-accelerated animations with Motion.dev (Framer Motion successor) for React, Next.js, Svelte, and Astro projects. Use when user requests animation, motion, scroll effects, parallax, hero animations, gestures, drag interactions, spring physics, whileHover effects, whileInView animations, animated UI, micro-interactions, page transitions, or layout animations. Generates production TypeScript/JSX code with accessibility (prefers-reduced-motion) and performance validation (≥60fps). Supports entrance animations, gesture interactions (hover/tap/drag), scroll-based reveals, and layout transitions using spring physics and natural timing. Do NOT use for CSS-only transitions (use native CSS), static sites without JavaScript, Vue animations (use motion-v variant instead), or SVG/Canvas complex animations (GSAP better suited).

NeverSight/learn-skills.dev

Static artifact craft skill for self-contained HTML/CSS/JS documents: docs, sheets, dashboards, explainers, slides, tools, and landing pages. Use when the user asks for a durable, openable, shareable web deliverable they'll keep or hand off — a report, a dashboard, a slide deck, a data table, a page. Local folder first, temporary public link via tunnel (localhost.run), optional durable publish to Surge, GitHub Pages, or Cloudflare. Not for quick look renders, inline snippets, or throwaway scratch. Not for SPA frameworks, backend APIs, database apps, or production product UI.

NeverSight/learn-skills.dev

FFmpeg commands for video/audio conversion, trimming, compression, and processing. Use when user mentions "ffmpeg", "convert video", "compress video", "extract audio", "trim video", "gif from video", "video codec", "transcode", "screen recording", "merge videos", "video to mp4", "reduce file size", or any media processing task.

NeverSight/learn-skills.dev

Vim keybindings, motions, text objects, and operators for efficient text editing. Use when user asks about "vim commands", "vim motions", "text objects", "vim keybindings", "vim cheat sheet", "learn vim", "vim in VS Code", or any Vim editing tasks.

NeverSight/learn-skills.dev

Grok Build ONLY. Turn a 2D character still into smooth animation sprites via image_gen/image_edit base → image_to_video (6s/10s run-in-place) → ffmpeg frames → magenta chroma-key → dense sampled sprites (strip/grid/GIF). Use when the user wants video-to-sprite, motion capture from generated video, smoother run/walk cycles from dense frames, or runs /video2dsprite. Do NOT use on Codex/Claude — only Grok Build has image_to_video. Prefer generate2dsprite for crisp pixel sheets without video.

相关技能