Communitygithub.com

NeverSight/learn-skills.dev

FFmpeg commands for video/audio conversion, trimming, compression, and processing. Use when user mentions "ffmpeg", "convert video", "compress video", "extract audio", "trim video", "gif from video", "video codec", "transcode", "screen recording", "merge videos", "video to mp4", "reduce file size", or any media processing task.

Was ist learn-skills.dev?

learn-skills.dev is a Claude Code agent skill that fFmpeg commands for video/audio conversion, trimming, compression, and processing. Use when user mentions "ffmpeg", "convert video", "compress video", "extract audio", "trim video", "gif from video", "video codec", "transcode", "screen recording", "merge videos", "video to mp4", "reduce file size", or any media processing task.

Funktioniert mit~Claude Code~Codex CLI~Cursor
npx skills add https://github.com/NeverSight/learn-skills.dev/tree/HEAD/data/skills-md/1mangesh1/dev-skills-collection/ffmpeg

In Ihrer bevorzugten KI fragen

Öffnet einen neuen Chat, in dem dieser Agent-Skill bereits geladen ist.

Dokumentation

FFmpeg

Probe and Inspect

# Show all stream info (codec, resolution, bitrate, duration)
ffprobe -v error -show_format -show_streams input.mp4

# One-line summary: duration, size, bitrate
ffprobe -v error -show_entries format=duration,size,bit_rate -of default=noprint_wrappers=1 input.mp4

# Show video resolution and frame rate
ffprobe -v error -select_streams v:0 -show_entries stream=width,height,r_frame_rate,codec_name -of csv=p=0 input.mp4

# List all supported codecs
ffmpeg -codecs

# List all supported formats
ffmpeg -formats

# List available encoders
ffmpeg -encoders

Convert Between Formats

# MP4 to MKV (copy streams, no re-encode -- fast)
ffmpeg -i input.mp4 -c copy output.mkv

# MKV to MP4 (re-encode if codecs are incompatible with MP4 container)
ffmpeg -i input.mkv -c:v libx264 -c:a aac output.mp4

# AVI to MP4
ffmpeg -i input.avi -c:v libx264 -c:a aac output.mp4

# MOV to MP4 (common for iPhone footage)
ffmpeg -i input.mov -c:v libx264 -c:a aac -movflags +faststart output.mp4

# MP4 to WebM (VP9 + Opus for web)
ffmpeg -i input.mp4 -c:v libvpx-vp9 -crf 30 -b:v 0 -c:a libopus output.webm

# Any format, let FFmpeg pick codecs for the target container
ffmpeg -i input.avi output.mp4

Compress Video

# CRF mode (constant quality). Lower CRF = better quality, bigger file.
# CRF 18 = visually lossless, 23 = default, 28 = smaller but visible loss.
ffmpeg -i input.mp4 -c:v libx264 -crf 23 -preset medium -c:a aac -b:a 128k output.mp4

# Faster encoding, larger file
ffmpeg -i input.mp4 -c:v libx264 -crf 23 -preset fast -c:a copy output.mp4

# Slower encoding, smaller file (use for archival)
ffmpeg -i input.mp4 -c:v libx264 -crf 23 -preset slow -c:a aac output.mp4

# H.265 for ~50% smaller files at same quality (slower to encode)
ffmpeg -i input.mp4 -c:v libx265 -crf 28 -preset medium -c:a aac output.mp4

# Target a specific file size (e.g., 25 MB for a 60s video)
# Calculate bitrate: (25 * 8192) / 60 = ~3413 kbps total. Subtract ~128 for audio.
ffmpeg -i input.mp4 -c:v libx264 -b:v 3285k -pass 1 -an -f null /dev/null && \
ffmpeg -i input.mp4 -c:v libx264 -b:v 3285k -pass 2 -c:a aac -b:a 128k output.mp4

Extract Audio

# Extract audio as MP3
ffmpeg -i input.mp4 -vn -c:a libmp3lame -q:a 2 output.mp3

# Extract audio as AAC (copy if already AAC)
ffmpeg -i input.mp4 -vn -c:a copy output.aac

# Extract audio as WAV (uncompressed)
ffmpeg -i input.mp4 -vn -c:a pcm_s16le output.wav

# Extract audio as FLAC (lossless)
ffmpeg -i input.mp4 -vn -c:a flac output.flac

Convert Audio Formats

# WAV to MP3 (VBR quality 2, roughly 190 kbps)
ffmpeg -i input.wav -c:a libmp3lame -q:a 2 output.mp3

# MP3 to AAC
ffmpeg -i input.mp3 -c:a aac -b:a 192k output.m4a

# FLAC to MP3
ffmpeg -i input.flac -c:a libmp3lame -q:a 0 output.mp3

# WAV to OGG Vorbis
ffmpeg -i input.wav -c:a libvorbis -q:a 5 output.ogg

# WAV to Opus (best quality-to-size ratio)
ffmpeg -i input.wav -c:a libopus -b:a 128k output.opus

# Any audio to WAV (for editing or compatibility)
ffmpeg -i input.mp3 -c:a pcm_s16le -ar 44100 output.wav

Trim and Cut

# Cut from 00:01:30 to 00:03:00 without re-encoding (fast, may have keyframe issues)
ffmpeg -ss 00:01:30 -to 00:03:00 -i input.mp4 -c copy output.mp4

# Cut with re-encoding (frame-accurate)
ffmpeg -ss 00:01:30 -to 00:03:00 -i input.mp4 -c:v libx264 -c:a aac output.mp4

# Cut first 30 seconds
ffmpeg -ss 0 -t 30 -i input.mp4 -c copy output.mp4

# Cut last 30 seconds (requires knowing duration, or use negative start)
ffmpeg -sseof -30 -i input.mp4 -c copy output.mp4

# Remove first 10 seconds
ffmpeg -ss 10 -i input.mp4 -c copy output.mp4

Create GIF

# Basic GIF (low quality, simple)
ffmpeg -i input.mp4 -vf "fps=10,scale=480:-1" output.gif

# High-quality GIF with palette generation (two-pass)
ffmpeg -i input.mp4 -vf "fps=15,scale=480:-1:flags=lanczos,palettegen" palette.png && \
ffmpeg -i input.mp4 -i palette.png -lavfi "fps=15,scale=480:-1:flags=lanczos [x]; [x][1:v] paletteuse" output.gif

# GIF from a specific segment (5 seconds starting at 00:00:30)
ffmpeg -ss 00:00:30 -t 5 -i input.mp4 -vf "fps=15,scale=480:-1:flags=lanczos,palettegen" palette.png && \
ffmpeg -ss 00:00:30 -t 5 -i input.mp4 -i palette.png -lavfi "fps=15,scale=480:-1:flags=lanczos [x]; [x][1:v] paletteuse" output.gif

Scale and Resize

# Scale to 1280x720
ffmpeg -i input.mp4 -vf "scale=1280:720" -c:a copy output.mp4

# Scale width to 1280, keep aspect ratio (-1 auto-calculates, -2 ensures even number)
ffmpeg -i input.mp4 -vf "scale=1280:-2" -c:a copy output.mp4

# Scale to 50% of original size
ffmpeg -i input.mp4 -vf "scale=iw/2:ih/2" -c:a copy output.mp4

# Scale to fit within 1920x1080, preserving aspect ratio (no upscale)
ffmpeg -i input.mp4 -vf "scale='min(1920,iw)':'min(1080,ih)':force_original_aspect_ratio=decrease" -c:a copy output.mp4

Subtitles

# Hardcode subtitles (burn into video, always visible)
ffmpeg -i input.mp4 -vf "subtitles=subs.srt" output.mp4

# Hardcode with custom style
ffmpeg -i input.mp4 -vf "subtitles=subs.srt:force_style='FontSize=24,PrimaryColour=&HFFFFFF'" output.mp4

# Soft subtitles (user can toggle on/off, MP4)
ffmpeg -i input.mp4 -i subs.srt -c copy -c:s mov_text output.mp4

# Soft subtitles (MKV, supports more formats including ASS)
ffmpeg -i input.mp4 -i subs.srt -c copy -c:s srt output.mkv

# Extract subtitles from video
ffmpeg -i input.mkv -map 0:s:0 output.srt

Merge and Concatenate

# Concatenate videos (same codec, resolution, framerate)
# First create a file list:
# file 'part1.mp4'
# file 'part2.mp4'
# file 'part3.mp4'
ffmpeg -f concat -safe 0 -i filelist.txt -c copy output.mp4

# Generate the file list from shell
for f in part*.mp4; do echo "file '$f'"; done > filelist.txt
ffmpeg -f concat -safe 0 -i filelist.txt -c copy output.mp4

# Concatenate with re-encoding (when formats differ)
ffmpeg -f concat -safe 0 -i filelist.txt -c:v libx264 -c:a aac output.mp4

# Combine video and audio from separate files
ffmpeg -i video.mp4 -i audio.mp3 -c:v copy -c:a aac -shortest output.mp4

# Replace audio track in a video
ffmpeg -i video.mp4 -i newaudio.mp3 -c:v copy -c:a aac -map 0:v:0 -map 1:a:0 output.mp4

Extract Frames

# Extract one frame at a specific timestamp
ffmpeg -ss 00:01:30 -i input.mp4 -frames:v 1 frame.png

# Extract one frame every second
ffmpeg -i input.mp4 -vf "fps=1" frames_%04d.png

# Extract one frame every 10 seconds
ffmpeg -i input.mp4 -vf "fps=1/10" frames_%04d.png

# Extract all frames (warning: generates many files)
ffmpeg -i input.mp4 frames_%06d.png

# Create a thumbnail sheet (4x4 grid)
ffmpeg -i input.mp4 -vf "select='not(mod(n\,100))',scale=320:180,tile=4x4" -frames:v 1 thumbnails.png

Screen Recording

# macOS -- record full screen
ffmpeg -f avfoundation -framerate 30 -i "1:0" -c:v libx264 -preset ultrafast -crf 18 output.mp4

# macOS -- list available devices
ffmpeg -f avfoundation -list_devices true -i ""

# macOS -- record screen with audio
ffmpeg -f avfoundation -framerate 30 -i "1:0" -c:v libx264 -preset ultrafast -c:a aac output.mp4

# Linux (X11) -- record full screen
ffmpeg -f x11grab -framerate 30 -video_size 1920x1080 -i :0.0 -c:v libx264 -preset ultrafast -crf 18 output.mp4

# Linux (X11) -- record a region (offset x=100,y=200, size 1280x720)
ffmpeg -f x11grab -framerate 30 -video_size 1280x720 -i :0.0+100,200 -c:v libx264 -preset ultrafast output.mp4

# Linux (PipeWire/Wayland) -- use pipewire screen capture
# Wayland does not support x11grab. Use pw-record or OBS with pipewire.

Speed Up and Slow Down

# Speed up video 2x (drop audio)
ffmpeg -i input.mp4 -vf "setpts=0.5*PTS" -an output.mp4

# Speed up video 2x and audio 2x
ffmpeg -i input.mp4 -vf "setpts=0.5*PTS" -af "atempo=2.0" output.mp4

# Slow down video 2x with audio
ffmpeg -i input.mp4 -vf "setpts=2.0*PTS" -af "atempo=0.5" output.mp4

# Speed up 4x (chain atempo filters; each atempo supports 0.5-2.0 range)
ffmpeg -i input.mp4 -vf "setpts=0.25*PTS" -af "atempo=2.0,atempo=2.0" output.mp4

Remove Audio

# Strip audio track, keep video as-is
ffmpeg -i input.mp4 -an -c:v copy output.mp4

Codec Reference

CodecUse CaseEncode FlagNotes
H.264General purpose, widest compatibility-c:v libx264Default choice. Works everywhere.
H.265Archival, streaming, 4K content-c:v libx265~50% smaller than H.264. Slower encode.
VP9Web delivery, YouTube-style hosting-c:v libvpx-vp9Royalty-free. Good browser support.
AV1Next-gen web/streaming (emerging)-c:v libaom-av1Best compression. Very slow to encode.
AACAudio in MP4 containers-c:a aacGood quality, universal support.
OpusVoice, music, streaming at low bitrates-c:a libopusBest quality-per-bit. WebM/OGG container.
MP3Legacy audio compatibility-c:a libmp3lameUse when AAC/Opus not accepted.

Quick decision guide:

  • Sharing on the web or need broad playback? H.264 + AAC in MP4.
  • Need smaller files and can wait longer to encode? H.265.
  • Targeting browsers specifically? VP9 + Opus in WebM.
  • Archiving and storage is a concern? H.265 or AV1.

Useful Flags

-y                  Overwrite output without asking
-n                  Never overwrite output
-hide_banner        Suppress FFmpeg build info
-v error            Only show errors (quieter output)
-movflags +faststart  Move MP4 metadata to beginning (better for streaming)
-map 0              Copy all streams from input
-shortest           Stop encoding when the shortest stream ends
-threads 0          Use all available CPU threads

Individual skills in this repo

This repo contains 20 individual skills — each has its own dedicated page.

NeverSight/learn-skills.dev

Use when generating or modifying Remotion video code, creating demo videos, or working with the demo-video/ directory

NeverSight/learn-skills.dev

Create AI avatar and talking head videos via inference.sh CLI. Recommended: P-Video-Avatar (fastest, cheapest, built-in TTS). Also: OmniHuman, Fabric, PixVerse. Audio: Inworld TTS-2 (100+ languages, emotion steering for characters), ElevenLabs, Kokoro. Capabilities: audio-driven avatars, text-to-avatar, lipsync videos, talking head generation, virtual presenters, UGC content. Use for: AI presenters, explainer videos, virtual influencers, dubbing, marketing videos, UGC ads, gaming avatars, NPC dialogue. Triggers: ai avatar, talking head, lipsync, avatar video, virtual presenter, ai spokesperson, audio driven video, heygen alternative, synthesia alternative, talking avatar, lip sync, video avatar, ai presenter, digital human, ugc, ugc video, ugc ad, avatar ugc

NeverSight/learn-skills.dev

Create AI marketing videos for ads, promos, product launches, and brand content. Models: Veo, Seedance, Wan, FLUX for visuals, Kokoro for voiceover. Types: product demos, testimonials, explainers, social ads, brand videos. Use for: Facebook ads, YouTube ads, product launches, brand awareness. Triggers: marketing video, ad video, promo video, commercial, brand video, product video, explainer video, ad creative, video ad, facebook ad video, youtube ad, instagram ad, tiktok ad, promotional video, launch video

NeverSight/learn-skills.dev

Generate AI videos with Google Veo, Seedance 2.0, HappyHorse, Wan, Grok and 40+ models via inference.sh CLI. Models: Veo 3.1, Veo 3, Seedance 2.0, HappyHorse 1.0, Wan 2.5, Grok Imagine Video, OmniHuman, Fabric, HunyuanVideo. Capabilities: text-to-video, image-to-video, reference-to-video, video editing, lipsync, avatar animation, video upscaling, foley sound. Use for: social media videos, marketing content, explainer videos, product demos, AI avatars. Triggers: video generation, ai video, text to video, image to video, veo, animate image, video from image, ai animation, video generator, generate video, t2v, i2v, ai video maker, create video with ai, runway alternative, pika alternative, sora alternative, kling alternative, seedance, happyhorse

NeverSight/learn-skills.dev

ElevenLabs automatic dubbing - translate and dub audio/video into 29 languages while preserving speaker voice via inference.sh CLI. Capabilities: auto speaker detection, voice-preserving translation, video dubbing, audio localization. Use for: content localization, video translation, multilingual content, international distribution. Triggers: dubbing, dub video, translate audio, video translation, audio translation, localize content, elevenlabs dubbing, eleven labs dub, multilingual dub, voice translation, auto dub, language dub, content localization

NeverSight/learn-skills.dev

Explainer video production guide: scripting, voiceover, visuals, and assembly. Covers script formulas, pacing rules, scene planning, and multi-tool pipelines. Use for: product demos, how-it-works videos, onboarding videos, social explainers. Triggers: explainer video, how to make explainer, product video, demo video, video production, video script, animated explainer, product demo video, tutorial video, onboarding video, walkthrough video, video pipeline

NeverSight/learn-skills.dev

Still-to-video conversion guide: model selection, motion prompting, and camera movement. Covers Wan 2.5 i2v, Seedance, Fabric, Grok Video with when to use each. Use for: animating images, creating video from stills, adding motion, product animations. Triggers: image to video, i2v, animate image, still to video, add motion to image, image animation, photo to video, animate still, wan i2v, image2video, bring image to life, animate photo, motion from image

NeverSight/learn-skills.dev

Generate talking head avatar videos with Pruna P-Video-Avatar via inference.sh CLI. Turn a portrait image into a realistic speaking video with built-in TTS. 18x faster and 6x cheaper than competitors. Models: P-Video-Avatar, P-Image (for portrait generation). Capabilities: text-to-avatar, audio-driven avatars, 30 voices, 10 languages, 720p/1080p, built-in TTS, dynamic backgrounds, full-body control. Use for: AI presenters, product demos, explainer videos, virtual influencers, marketing, education, multilingual content, UGC, gaming avatars. Triggers: avatar video, talking head, ai avatar, p-video-avatar, pruna avatar, video avatar, ai presenter, digital human, virtual presenter, lipsync, talking avatar, ai spokesperson, heygen alternative, synthesia alternative, veed alternative, fabric alternative, omnihuman alternative

NeverSight/learn-skills.dev

Generate videos with Pruna P-Video and WAN models via inference.sh CLI. Models: P-Video, WAN-T2V, WAN-I2V. Capabilities: text-to-video, image-to-video, audio support, 720p/1080p, fast inference. Pruna optimizes models for speed without quality loss. Triggers: pruna video, p-video, pruna ai video, fast video generation, optimized video, wan t2v, wan i2v, economic video generation, cheap video generation, pruna text to video, pruna image to video

NeverSight/learn-skills.dev

Render videos from React/Remotion component code via inference.sh. Pass TSX code, get MP4. Supports all Remotion APIs: useCurrentFrame, useVideoConfig, spring, interpolate, AbsoluteFill, Sequence. Configurable resolution, FPS, duration, codec. Use for: programmatic video generation, animated graphics, motion design, data-driven videos, React animations to video. Triggers: remotion, render video from code, tsx to video, react video, programmatic video, remotion render, code to video, animated video, motion graphics code, react animation video

NeverSight/learn-skills.dev

Video ad creation with exact platform-specific specs for TikTok, Instagram, YouTube, Facebook, LinkedIn. Covers dimensions, duration limits, AIDA framework, and caption requirements. Use for: video ads, social media ads, paid media creative, video marketing, ad production. Triggers: video ad, social media ad, tiktok ad, instagram ad, youtube ad, facebook ad, linkedin ad, video creative, ad specs, paid media, video marketing, ad production, reels ad, stories ad, pre roll, bumper ad

NeverSight/learn-skills.dev

Best practices and techniques for writing effective AI video generation prompts. Covers: Veo, Seedance, Wan, Grok, Kling, Runway, Pika, Sora prompting strategies. Learn: shot types, camera movements, lighting, pacing, style keywords, negative prompts. Use for: improving video quality, getting consistent results, professional video prompts. Triggers: video prompt, how to prompt video, veo prompts, video generation tips, better ai video, video prompt engineering, video prompt guide, video prompt template, ai video tips, video prompt best practices, video prompt examples, cinematography prompts

NeverSight/learn-skills.dev

YouTube thumbnail design with specific dimensions, contrast rules, and mobile preview optimization. Covers safe zones, text placement, face expression psychology, and A/B testing. Use for: YouTube thumbnails, video cover images, click-through optimization. Triggers: youtube thumbnail, thumbnail design, video thumbnail, click through rate, ctr optimization, youtube cover, video cover image, thumbnail maker, thumbnail tips, youtube design, video preview image

NeverSight/learn-skills.dev

Landing page conversion optimization with layout rules, hero section design, and CTA psychology. Covers above-the-fold formula, social proof placement, mobile design, and F-pattern reading. Use for: startup landing pages, product pages, SaaS marketing, conversion optimization. Triggers: landing page, hero section, above the fold, conversion optimization, landing page design, cta button, hero image, landing page layout, saas landing page, product page design, conversion rate, landing page best practices

NeverSight/learn-skills.dev

Configure and use the hosted YouTube Data MCP end-to-end with minimal user input. Use when users want the agent to verify Node.js and `npx`, configure MCP server config (Windows/macOS, Cursor/Codex/OpenClaw/OpenCode), request API key at setup time, run post-install capability discovery (`tools/list` and `get_patch_notes`), and then strongly recommend helper skill and Python setup for full local document and spreadsheet workflows.

NeverSight/learn-skills.dev

Creates 120fps GPU-accelerated animations with Motion.dev (Framer Motion successor) for React, Next.js, Svelte, and Astro projects. Use when user requests animation, motion, scroll effects, parallax, hero animations, gestures, drag interactions, spring physics, whileHover effects, whileInView animations, animated UI, micro-interactions, page transitions, or layout animations. Generates production TypeScript/JSX code with accessibility (prefers-reduced-motion) and performance validation (≥60fps). Supports entrance animations, gesture interactions (hover/tap/drag), scroll-based reveals, and layout transitions using spring physics and natural timing. Do NOT use for CSS-only transitions (use native CSS), static sites without JavaScript, Vue animations (use motion-v variant instead), or SVG/Canvas complex animations (GSAP better suited).

NeverSight/learn-skills.dev

Static artifact craft skill for self-contained HTML/CSS/JS documents: docs, sheets, dashboards, explainers, slides, tools, and landing pages. Use when the user asks for a durable, openable, shareable web deliverable they'll keep or hand off — a report, a dashboard, a slide deck, a data table, a page. Local folder first, temporary public link via tunnel (localhost.run), optional durable publish to Surge, GitHub Pages, or Cloudflare. Not for quick look renders, inline snippets, or throwaway scratch. Not for SPA frameworks, backend APIs, database apps, or production product UI.

NeverSight/learn-skills.dev

Vim keybindings, motions, text objects, and operators for efficient text editing. Use when user asks about "vim commands", "vim motions", "text objects", "vim keybindings", "vim cheat sheet", "learn vim", "vim in VS Code", or any Vim editing tasks.

NeverSight/learn-skills.dev

Download YouTube videos, extract/proofread/translate subtitles, and render them onto video. Use whenever the user asks to download a YouTube video with subtitles, translate video subtitles to Chinese, add subtitles/burn subtitles to a video, or do ASR transcription on video audio. Covers the full pipeline: video download → subtitle extraction → proofreading → translation → SRT generation → subtitle rendering. Also use for "下载视频加字幕", "视频翻译字幕", "把字幕烧录到视频中". For pure subtitle file creation without rendering (just SRT output), this skill handles that too — stop before Phase 5.

NeverSight/learn-skills.dev

Grok Build ONLY. Turn a 2D character still into smooth animation sprites via image_gen/image_edit base → image_to_video (6s/10s run-in-place) → ffmpeg frames → magenta chroma-key → dense sampled sprites (strip/grid/GIF). Use when the user wants video-to-sprite, motion capture from generated video, smoother run/walk cycles from dense frames, or runs /video2dsprite. Do NOT use on Codex/Claude — only Grok Build has image_to_video. Prefer generate2dsprite for crisp pixel sheets without video.

Verwandte Skills