Communitygithub.com

NeverSight/learn-skills.dev

AI 辅助视频编辑工作流,用于剪切、结构化和增强真实素材。覆盖从原始录制到 FFmpeg、Remotion、ElevenLabs、fal.ai 以及在 Descript 或 CapCut 中最终润色的完整流水线。当用户想要编辑视频、剪切素材、制作 Vlog 或构建视频内容时使用。

Was ist learn-skills.dev?

learn-skills.dev is a Claude Code agent skill that aI 辅助视频编辑工作流,用于剪切、结构化和增强真实素材。覆盖从原始录制到 FFmpeg、Remotion、ElevenLabs、fal.ai 以及在 Descript 或 CapCut 中最终润色的完整流水线。当用户想要编辑视频、剪切素材、制作 Vlog 或构建视频内容时使用。.

Funktioniert mit✓Claude Code✓Codex CLI~Cursor
npx skills add https://github.com/NeverSight/learn-skills.dev/tree/HEAD/data/skills-md/aaione/everything-claude-code-zh/video-editing

In Ihrer bevorzugten KI fragen

Öffnet einen neuen Chat, in dem dieser Agent-Skill bereits geladen ist.

Vorschau

Aus dem Skill-README

Dokumentation

视频编辑

AI 辅助编辑真实素材。不是从提示词生成。快速编辑现有视频。

何时激活

  • 用户想要编辑、剪切或结构化视频素材
  • 将长录制转换为短视频内容
  • 从原始素材构建 Vlog、教程或演示视频
  • 为现有视频添加叠加层、字幕、音乐或旁白
  • 为不同平台(YouTube、TikTok、Instagram)重新构架视频
  • 用户说"编辑视频"、"剪切这段素材"、"制作 Vlog"或"视频工作流"

核心论点

AI 视频编辑在你停止要求它创建整个视频,开始用它来压缩、结构化和增强真实素材时才有用。价值不在于生成。价值在于压缩。

流水线

Screen Studio / 原始素材
  → Claude / Codex
  → FFmpeg
  → Remotion
  → ElevenLabs / fal.ai
  → Descript 或 CapCut

每一层都有特定的工作。不要跳过层级。不要试图让一个工具做所有事情。

第 1 层:采集(Screen Studio / 原始素材)

收集源材料:

  • Screen Studio:用于应用演示、编码会话、浏览器工作流的精美屏幕录制
  • 原始摄像头素材:Vlog 素材、访谈、活动录制
  • 通过 VideoDB 的桌面捕获:带实时上下文的会话录制(参见 videodb 技能)

输出:准备好整理的原始文件。

第 2 层:整理(Claude / Codex)

使用 Claude Code 或 Codex 来:

  • 转录和标记:生成转录文本,识别主题和话题
  • 规划结构:决定保留什么、剪切什么、什么顺序合适
  • 识别死区:查找停顿、跑题、重复拍摄
  • 生成编辑决策列表:剪切的时间戳、保留的片段
  • 搭建 FFmpeg 和 Remotion 代码:生成命令和合成
示例提示:
"这是一段 4 小时录制的转录文本。找出 8 个最强片段
用于 24 分钟的 Vlog。给我每个片段的 FFmpeg 剪切命令。"

这一层关乎结构,而非最终创意品味。

第 3 层:确定性剪切(FFmpeg)

FFmpeg 处理枯燥但关键的工作:分割、裁剪、拼接和预处理。

按时间戳提取片段

ffmpeg -i raw.mp4 -ss 00:12:30 -to 00:15:45 -c copy segment_01.mp4

从编辑决策列表批量剪切

#!/bin/bash
# cuts.txt: 开始,结束,标签
while IFS=, read -r start end label; do
  ffmpeg -i raw.mp4 -ss "$start" -to "$end" -c copy "segments/${label}.mp4"
done < cuts.txt

拼接片段

# 创建文件列表
for f in segments/*.mp4; do echo "file '$f'"; done > concat.txt
ffmpeg -f concat -safe 0 -i concat.txt -c copy assembled.mp4

创建代理文件以加快编辑

ffmpeg -i raw.mp4 -vf "scale=960:-2" -c:v libx264 -preset ultrafast -crf 28 proxy.mp4

提取音频用于转录

ffmpeg -i raw.mp4 -vn -acodec pcm_s16le -ar 16000 audio.wav

归一化音频电平

ffmpeg -i segment.mp4 -af loudnorm=I=-16:TP=-1.5:LRA=11 -c:v copy normalized.mp4

第 4 层:可编程合成(Remotion)

Remotion 将编辑问题转化为可组合的代码。用于传统编辑器难以处理的事情:

何时使用 Remotion

  • 叠加层:文本、图片、品牌、字幕条
  • 数据可视化:图表、统计、动画数字
  • 动态图形:转场、说明动画
  • 可组合场景:跨视频可复用的模板
  • 产品演示:带注释的截图、UI 高亮

基本 Remotion 合成

import { AbsoluteFill, Sequence, Video, useCurrentFrame } from "remotion";

export const VlogComposition: React.FC = () => {
  const frame = useCurrentFrame();

  return (
    <AbsoluteFill>
      {/* 主要素材 */}
      <Sequence from={0} durationInFrames={300}>
        <Video src="/segments/intro.mp4" />
      </Sequence>

      {/* 标题叠加 */}
      <Sequence from={30} durationInFrames={90}>
        <AbsoluteFill style={{
          justifyContent: "center",
          alignItems: "center",
        }}>
          <h1 style={{
            fontSize: 72,
            color: "white",
            textShadow: "2px 2px 8px rgba(0,0,0,0.8)",
          }}>
            AI 编辑技术栈
          </h1>
        </AbsoluteFill>
      </Sequence>

      {/* 下一个片段 */}
      <Sequence from={300} durationInFrames={450}>
        <Video src="/segments/demo.mp4" />
      </Sequence>
    </AbsoluteFill>
  );
};

渲染输出

npx remotion render src/index.ts VlogComposition output.mp4

详见 Remotion 文档获取详细模式和 API 参考。

第 5 层:生成资产(ElevenLabs / fal.ai)

只生成你需要的内容。不要生成整个视频。

使用 ElevenLabs 的旁白

import os
import requests

resp = requests.post(
    f"https://api.elevenlabs.io/v1/text-to-speech/{voice_id}",
    headers={
        "xi-api-key": os.environ["ELEVENLABS_API_KEY"],
        "Content-Type": "application/json"
    },
    json={
        "text": "你的旁白文本",
        "model_id": "eleven_turbo_v2_5",
        "voice_settings": {"stability": 0.5, "similarity_boost": 0.75}
    }
)
with open("voiceover.mp3", "wb") as f:
    f.write(resp.content)

使用 fal.ai 的音乐和音效

使用 fal-ai-media 技能进行:

  • 背景音乐生成
  • 音效(用于视频转音频的 ThinkSound 模型)
  • 转场音效

使用 fal.ai 的生成视觉

用于不存在的插入镜头、缩略图或 B-roll:

generate(app_id: "fal-ai/nano-banana-pro", input_data: {
  "prompt": "科技 Vlog 专业缩略图,深色背景,屏幕上的代码",
  "image_size": "landscape_16_9"
})

VideoDB 生成音频

如果 VideoDB 已配置:

voiceover = coll.generate_voice(text="此处为旁白", voice="alloy")
music = coll.generate_music(prompt="编程 Vlog 的 lo-fi 背景音乐", duration=120)
sfx = coll.generate_sound_effect(prompt="微弱的嗖嗖转场音效")

第 6 层:最终润色(Descript / CapCut)

最后一层是人的工作。使用传统编辑器进行:

  • 节奏:调整感觉太快或太慢的剪切
  • 字幕:自动生成,然后手动校对
  • 调色:基本校正和氛围
  • 最终音频混音:平衡人声、音乐和音效电平
  • 导出:特定平台的格式和质量设置

品味在这里体现。AI 清除重复性工作。你做最终决定。

社交媒体重新构架

不同平台需要不同的宽高比:

平台宽高比分辨率
YouTube16:91920x1080
TikTok / Reels9:161080x1920
Instagram 动态1:11080x1080
X / Twitter16:9 或 1:11280x720 或 720x720

使用 FFmpeg 重新构架

# 16:9 转 9:16(居中裁剪)
ffmpeg -i input.mp4 -vf "crop=ih*9/16:ih,scale=1080:1920" vertical.mp4

# 16:9 转 1:1(居中裁剪)
ffmpeg -i input.mp4 -vf "crop=ih:ih,scale=1080:1080" square.mp4

使用 VideoDB 重新构架

from videodb import ReframeMode

# 智能重新构架(AI 引导的主体跟踪)
reframed = video.reframe(start=0, end=60, target="vertical", mode=ReframeMode.smart)

场景检测和自动剪切

FFmpeg 场景检测

# 检测场景变化(阈值 0.3 = 中等灵敏度)
ffmpeg -i input.mp4 -vf "select='gt(scene,0.3)',showinfo" -vsync vfr -f null - 2>&1 | grep showinfo

静音检测用于自动剪切

# 查找静音片段(适用于剪切空白时段)
ffmpeg -i input.mp4 -af silencedetect=noise=-30dB:d=2 -f null - 2>&1 | grep silence

高亮提取

使用 Claude 分析转录文本 + 场景时间戳:

"根据这个带时间戳的转录文本和这些场景变化点,
识别 5 个最适合社交媒体的 30 秒片段。"

每个工具最擅长的领域

工具优势劣势
Claude / Codex整理、规划、代码生成不是创意品味层
FFmpeg确定性剪切、批处理、格式转换无可视化编辑界面
Remotion可编程叠加层、可组合场景、可复用模板非开发者有学习曲线
Screen Studio即刻生成精美屏幕录制仅限屏幕捕获
ElevenLabs语音、旁白、音乐、音效不是工作流中心
Descript / CapCut最终节奏、字幕、润色手动操作,不可自动化

关键原则

  1. 编辑,而非生成。 此工作流用于剪切真实素材,而非从提示词创建。
  2. 结构先于风格。 在第 2 层把故事理顺,再触碰任何视觉内容。
  3. FFmpeg 是骨干。 枯燥但关键。长素材在这里变得可管理。
  4. Remotion 用于可重复性。 如果你将多次执行相同操作,把它做成 Remotion 组件。
  5. 选择性生成。 只对不存在的资产使用 AI 生成,不要用于所有内容。
  6. 品味是最后一层。 AI 清除重复性工作。你做最终创意决策。

相关技能

  • fal-ai-media — AI 图像、视频和音频生成
  • videodb — 服务端视频处理、索引和流媒体
  • content-engine — 平台原生内容分发

Individual skills in this repo

This repo contains 20 individual skills — each has its own dedicated page.

NeverSight/learn-skills.dev

Use when generating or modifying Remotion video code, creating demo videos, or working with the demo-video/ directory

NeverSight/learn-skills.dev

Create AI avatar and talking head videos via inference.sh CLI. Recommended: P-Video-Avatar (fastest, cheapest, built-in TTS). Also: OmniHuman, Fabric, PixVerse. Audio: Inworld TTS-2 (100+ languages, emotion steering for characters), ElevenLabs, Kokoro. Capabilities: audio-driven avatars, text-to-avatar, lipsync videos, talking head generation, virtual presenters, UGC content. Use for: AI presenters, explainer videos, virtual influencers, dubbing, marketing videos, UGC ads, gaming avatars, NPC dialogue. Triggers: ai avatar, talking head, lipsync, avatar video, virtual presenter, ai spokesperson, audio driven video, heygen alternative, synthesia alternative, talking avatar, lip sync, video avatar, ai presenter, digital human, ugc, ugc video, ugc ad, avatar ugc

NeverSight/learn-skills.dev

Create AI marketing videos for ads, promos, product launches, and brand content. Models: Veo, Seedance, Wan, FLUX for visuals, Kokoro for voiceover. Types: product demos, testimonials, explainers, social ads, brand videos. Use for: Facebook ads, YouTube ads, product launches, brand awareness. Triggers: marketing video, ad video, promo video, commercial, brand video, product video, explainer video, ad creative, video ad, facebook ad video, youtube ad, instagram ad, tiktok ad, promotional video, launch video

NeverSight/learn-skills.dev

Generate AI videos with Google Veo, Seedance 2.0, HappyHorse, Wan, Grok and 40+ models via inference.sh CLI. Models: Veo 3.1, Veo 3, Seedance 2.0, HappyHorse 1.0, Wan 2.5, Grok Imagine Video, OmniHuman, Fabric, HunyuanVideo. Capabilities: text-to-video, image-to-video, reference-to-video, video editing, lipsync, avatar animation, video upscaling, foley sound. Use for: social media videos, marketing content, explainer videos, product demos, AI avatars. Triggers: video generation, ai video, text to video, image to video, veo, animate image, video from image, ai animation, video generator, generate video, t2v, i2v, ai video maker, create video with ai, runway alternative, pika alternative, sora alternative, kling alternative, seedance, happyhorse

NeverSight/learn-skills.dev

ElevenLabs automatic dubbing - translate and dub audio/video into 29 languages while preserving speaker voice via inference.sh CLI. Capabilities: auto speaker detection, voice-preserving translation, video dubbing, audio localization. Use for: content localization, video translation, multilingual content, international distribution. Triggers: dubbing, dub video, translate audio, video translation, audio translation, localize content, elevenlabs dubbing, eleven labs dub, multilingual dub, voice translation, auto dub, language dub, content localization

NeverSight/learn-skills.dev

Explainer video production guide: scripting, voiceover, visuals, and assembly. Covers script formulas, pacing rules, scene planning, and multi-tool pipelines. Use for: product demos, how-it-works videos, onboarding videos, social explainers. Triggers: explainer video, how to make explainer, product video, demo video, video production, video script, animated explainer, product demo video, tutorial video, onboarding video, walkthrough video, video pipeline

NeverSight/learn-skills.dev

Still-to-video conversion guide: model selection, motion prompting, and camera movement. Covers Wan 2.5 i2v, Seedance, Fabric, Grok Video with when to use each. Use for: animating images, creating video from stills, adding motion, product animations. Triggers: image to video, i2v, animate image, still to video, add motion to image, image animation, photo to video, animate still, wan i2v, image2video, bring image to life, animate photo, motion from image

NeverSight/learn-skills.dev

Generate talking head avatar videos with Pruna P-Video-Avatar via inference.sh CLI. Turn a portrait image into a realistic speaking video with built-in TTS. 18x faster and 6x cheaper than competitors. Models: P-Video-Avatar, P-Image (for portrait generation). Capabilities: text-to-avatar, audio-driven avatars, 30 voices, 10 languages, 720p/1080p, built-in TTS, dynamic backgrounds, full-body control. Use for: AI presenters, product demos, explainer videos, virtual influencers, marketing, education, multilingual content, UGC, gaming avatars. Triggers: avatar video, talking head, ai avatar, p-video-avatar, pruna avatar, video avatar, ai presenter, digital human, virtual presenter, lipsync, talking avatar, ai spokesperson, heygen alternative, synthesia alternative, veed alternative, fabric alternative, omnihuman alternative

NeverSight/learn-skills.dev

Generate videos with Pruna P-Video and WAN models via inference.sh CLI. Models: P-Video, WAN-T2V, WAN-I2V. Capabilities: text-to-video, image-to-video, audio support, 720p/1080p, fast inference. Pruna optimizes models for speed without quality loss. Triggers: pruna video, p-video, pruna ai video, fast video generation, optimized video, wan t2v, wan i2v, economic video generation, cheap video generation, pruna text to video, pruna image to video

NeverSight/learn-skills.dev

Render videos from React/Remotion component code via inference.sh. Pass TSX code, get MP4. Supports all Remotion APIs: useCurrentFrame, useVideoConfig, spring, interpolate, AbsoluteFill, Sequence. Configurable resolution, FPS, duration, codec. Use for: programmatic video generation, animated graphics, motion design, data-driven videos, React animations to video. Triggers: remotion, render video from code, tsx to video, react video, programmatic video, remotion render, code to video, animated video, motion graphics code, react animation video

NeverSight/learn-skills.dev

Video ad creation with exact platform-specific specs for TikTok, Instagram, YouTube, Facebook, LinkedIn. Covers dimensions, duration limits, AIDA framework, and caption requirements. Use for: video ads, social media ads, paid media creative, video marketing, ad production. Triggers: video ad, social media ad, tiktok ad, instagram ad, youtube ad, facebook ad, linkedin ad, video creative, ad specs, paid media, video marketing, ad production, reels ad, stories ad, pre roll, bumper ad

NeverSight/learn-skills.dev

Best practices and techniques for writing effective AI video generation prompts. Covers: Veo, Seedance, Wan, Grok, Kling, Runway, Pika, Sora prompting strategies. Learn: shot types, camera movements, lighting, pacing, style keywords, negative prompts. Use for: improving video quality, getting consistent results, professional video prompts. Triggers: video prompt, how to prompt video, veo prompts, video generation tips, better ai video, video prompt engineering, video prompt guide, video prompt template, ai video tips, video prompt best practices, video prompt examples, cinematography prompts

NeverSight/learn-skills.dev

YouTube thumbnail design with specific dimensions, contrast rules, and mobile preview optimization. Covers safe zones, text placement, face expression psychology, and A/B testing. Use for: YouTube thumbnails, video cover images, click-through optimization. Triggers: youtube thumbnail, thumbnail design, video thumbnail, click through rate, ctr optimization, youtube cover, video cover image, thumbnail maker, thumbnail tips, youtube design, video preview image

NeverSight/learn-skills.dev

Landing page conversion optimization with layout rules, hero section design, and CTA psychology. Covers above-the-fold formula, social proof placement, mobile design, and F-pattern reading. Use for: startup landing pages, product pages, SaaS marketing, conversion optimization. Triggers: landing page, hero section, above the fold, conversion optimization, landing page design, cta button, hero image, landing page layout, saas landing page, product page design, conversion rate, landing page best practices

NeverSight/learn-skills.dev

Configure and use the hosted YouTube Data MCP end-to-end with minimal user input. Use when users want the agent to verify Node.js and `npx`, configure MCP server config (Windows/macOS, Cursor/Codex/OpenClaw/OpenCode), request API key at setup time, run post-install capability discovery (`tools/list` and `get_patch_notes`), and then strongly recommend helper skill and Python setup for full local document and spreadsheet workflows.

NeverSight/learn-skills.dev

Creates 120fps GPU-accelerated animations with Motion.dev (Framer Motion successor) for React, Next.js, Svelte, and Astro projects. Use when user requests animation, motion, scroll effects, parallax, hero animations, gestures, drag interactions, spring physics, whileHover effects, whileInView animations, animated UI, micro-interactions, page transitions, or layout animations. Generates production TypeScript/JSX code with accessibility (prefers-reduced-motion) and performance validation (≥60fps). Supports entrance animations, gesture interactions (hover/tap/drag), scroll-based reveals, and layout transitions using spring physics and natural timing. Do NOT use for CSS-only transitions (use native CSS), static sites without JavaScript, Vue animations (use motion-v variant instead), or SVG/Canvas complex animations (GSAP better suited).

NeverSight/learn-skills.dev

Static artifact craft skill for self-contained HTML/CSS/JS documents: docs, sheets, dashboards, explainers, slides, tools, and landing pages. Use when the user asks for a durable, openable, shareable web deliverable they'll keep or hand off — a report, a dashboard, a slide deck, a data table, a page. Local folder first, temporary public link via tunnel (localhost.run), optional durable publish to Surge, GitHub Pages, or Cloudflare. Not for quick look renders, inline snippets, or throwaway scratch. Not for SPA frameworks, backend APIs, database apps, or production product UI.

NeverSight/learn-skills.dev

FFmpeg commands for video/audio conversion, trimming, compression, and processing. Use when user mentions "ffmpeg", "convert video", "compress video", "extract audio", "trim video", "gif from video", "video codec", "transcode", "screen recording", "merge videos", "video to mp4", "reduce file size", or any media processing task.

NeverSight/learn-skills.dev

Vim keybindings, motions, text objects, and operators for efficient text editing. Use when user asks about "vim commands", "vim motions", "text objects", "vim keybindings", "vim cheat sheet", "learn vim", "vim in VS Code", or any Vim editing tasks.

NeverSight/learn-skills.dev

Grok Build ONLY. Turn a 2D character still into smooth animation sprites via image_gen/image_edit base → image_to_video (6s/10s run-in-place) → ffmpeg frames → magenta chroma-key → dense sampled sprites (strip/grid/GIF). Use when the user wants video-to-sprite, motion capture from generated video, smoother run/walk cycles from dense frames, or runs /video2dsprite. Do NOT use on Codex/Claude — only Grok Build has image_to_video. Prefer generate2dsprite for crisp pixel sheets without video.

Verwandte Skills