Communitygithub.com

bg-szy/TOP-SKILLS

Process video files with audio extraction, format conversion (mp4, webm), and Whisper transcription. Use when user mentions video conversion, audio extraction, transcription, mp4, webm, ffmpeg, or whisper transcription.

What is TOP-SKILLS?

TOP-SKILLS is a Claude Code agent skill that process video files with audio extraction, format conversion (mp4, webm), and Whisper transcription. Use when user mentions video conversion, audio extraction, transcription, mp4, webm, ffmpeg, or whisper transcription.

Works with✓Claude Code✓Codex CLI~Cursor
npx skills add https://github.com/bg-szy/TOP-SKILLS/tree/HEAD/skills/marketplace/video-processor

Ask in your favorite AI

Open a new chat with this agent skill pre-loaded.

Documentation

Video Processor

Instructions

This skill provides video processing utilities including audio extraction, format conversion, and audio transcription using FFmpeg and OpenAI's Whisper model.

Prerequisites

Required tools (must be installed in your environment):

  • FFmpeg: Multimedia framework for video/audio processing

    # macOS
    brew install ffmpeg
    
    # Ubuntu/Debian
    apt-get install ffmpeg
    
    # Verify installation
    ffmpeg -version
    
  • OpenAI Whisper: Speech-to-text transcription model

    # Install via pip
    pip install -U openai-whisper
    
    # Verify installation
    whisper --help
    

Python packages (included in script via PEP 723):

  • click (CLI framework)
  • ffmpeg-python (Python wrapper for FFmpeg)

Workflow

Use the scripts/video_processor.py script for all video processing tasks. The script provides a simple CLI with the following commands:

1. Extract Audio from Video

Extract the audio track from a video file:

uv run .claude/skills/video-processor/scripts/video_processor.py extract-audio input.mp4 output.wav

Options:

  • --format: Output audio format (default: wav). Supports: wav, mp3, aac, flac
  • Output is suitable for transcription or standalone audio use

2. Convert Video to MP4

Convert any video file to MP4 format:

uv run .claude/skills/video-processor/scripts/video_processor.py to-mp4 input.avi output.mp4

Options:

  • --codec: Video codec (default: libx264). Common options: libx264, libx265, h264
  • --preset: Encoding speed/quality preset (default: medium). Options: ultrafast, fast, medium, slow, veryslow

3. Convert Video to WebM

Convert any video file to WebM format (web-optimized):

uv run .claude/skills/video-processor/scripts/video_processor.py to-webm input.mp4 output.webm

Options:

  • --codec: Video codec (default: libvpx-vp9). Options: libvpx, libvpx-vp9
  • WebM is optimized for web playback and streaming

4. Transcribe Audio with Whisper

Transcribe audio or video files to text using OpenAI's Whisper model:

# Transcribe video file (audio will be extracted automatically)
uv run .claude/skills/video-processor/scripts/video_processor.py transcribe input.mp4 transcript.txt

# Transcribe audio file directly
uv run .claude/skills/video-processor/scripts/video_processor.py transcribe audio.wav transcript.txt

Options:

  • --model: Whisper model size (default: base). Options:
    • tiny: Fastest, lowest accuracy (~1GB RAM)
    • base: Fast, good accuracy (~1GB RAM) [DEFAULT]
    • small: Balanced (~2GB RAM)
    • medium: High accuracy (~5GB RAM)
    • large: Best accuracy, slowest (~10GB RAM)
  • --language: Language code (default: auto-detect). Examples: en, es, fr, de, zh
  • --format: Output format (default: txt). Options: txt, srt, vtt, json

Transcription workflow:

  1. If input is video, FFmpeg extracts audio to temporary WAV file
  2. Whisper processes the audio file
  3. Transcription is saved in requested format
  4. Temporary files are cleaned up automatically

5. Combined Workflow Example

Process a video end-to-end:

# 1. Extract audio for analysis
uv run .claude/skills/video-processor/scripts/video_processor.py extract-audio lecture.mp4 lecture.wav

# 2. Transcribe to SRT subtitles
uv run .claude/skills/video-processor/scripts/video_processor.py transcribe lecture.mp4 lecture.srt --format srt --model small

# 3. Convert to web format
uv run .claude/skills/video-processor/scripts/video_processor.py to-webm lecture.mp4 lecture.webm

Key Technical Details

FFmpeg and Whisper Integration:

  • FFmpeg doesn't transcribe audio itself - it prepares audio for external transcription
  • The workflow is: Extract audio (FFmpeg) → Transcribe (Whisper) → Optional: Re-integrate with video
  • FFmpeg can pipe audio directly to Whisper for real-time processing (advanced use case)

Audio Format for Transcription:

  • Whisper works best with WAV or MP3 formats
  • Sample rate: 16kHz is optimal (script handles conversion automatically)
  • The script extracts audio with optimal settings for Whisper

Output Formats:

  • txt: Plain text transcript
  • srt: SubRip subtitle format (includes timestamps)
  • vtt: WebVTT subtitle format (web standard)
  • json: Detailed JSON with word-level timestamps

Error Handling

The script includes comprehensive error handling:

  • Validates input files exist
  • Checks FFmpeg and Whisper are installed
  • Provides clear error messages for missing dependencies
  • Handles temporary file cleanup on errors

Performance Tips

  • Use tiny or base models for quick drafts
  • Use small or medium for production transcriptions
  • Use large only when maximum accuracy is required
  • For long videos, consider extracting audio first, then transcribe in segments
  • WebM conversion with VP9 takes longer but produces smaller files

Examples

Example 1: Quick Video to MP4 Conversion

User request:

I have an AVI file from my old camera. Can you convert it to MP4?

You would:

  1. Use the to-mp4 command with default settings:
    uv run .claude/skills/video-processor/scripts/video_processor.py to-mp4 old_video.avi output.mp4
    
  2. Confirm the conversion completed successfully
  3. Inform the user about the output file location

Example 2: Extract Audio and Transcribe

User request:

I recorded a lecture video and need a transcript. Can you extract the audio and transcribe it?

You would:

  1. First extract the audio:
    uv run .claude/skills/video-processor/scripts/video_processor.py extract-audio lecture.mp4 lecture.wav
    
  2. Then transcribe using the base model (good balance of speed/accuracy):
    uv run .claude/skills/video-processor/scripts/video_processor.py transcribe lecture.mp4 transcript.txt --model base
    
  3. Share the transcript.txt file with the user

Example 3: Create Web-Optimized Video with Subtitles

User request:

I need to put this video on my website with subtitles. Can you help?

You would:

  1. Convert to WebM for web optimization:
    uv run .claude/skills/video-processor/scripts/video_processor.py to-webm presentation.mp4 presentation.webm
    
  2. Generate SRT subtitle file:
    uv run .claude/skills/video-processor/scripts/video_processor.py transcribe presentation.mp4 subtitles.srt --format srt --model small
    
  3. Inform user they now have:
    • presentation.webm (web-optimized video)
    • subtitles.srt (subtitle file for embedding)

Example 4: High-Quality Transcription with Language Specification

User request:

I have a Spanish interview video that needs an accurate transcript for publication.

You would:

  1. Use a larger model with language specified for best accuracy:
    uv run .claude/skills/video-processor/scripts/video_processor.py transcribe interview.mp4 transcript.txt --model medium --language es
    
  2. Optionally create SRT for review:
    uv run .claude/skills/video-processor/scripts/video_processor.py transcribe interview.mp4 transcript.srt --format srt --model medium --language es
    
  3. Review the transcript with the user and make any necessary corrections

Example 5: Batch Processing Multiple Videos

User request:

I have a folder of training videos that all need to be converted to WebM and transcribed.

You would:

  1. List all video files in the directory:
    ls training_videos/*.mp4
    
  2. For each video file, run the conversion and transcription:
    # For each video: video1.mp4, video2.mp4, etc.
    uv run .claude/skills/video-processor/scripts/video_processor.py to-webm training_videos/video1.mp4 output/video1.webm
    uv run .claude/skills/video-processor/scripts/video_processor.py transcribe training_videos/video1.mp4 output/video1.txt --model base
    
    # Repeat for each file
    
  3. Confirm all conversions and transcriptions completed
  4. Provide summary of output files

Summary

The video-processor skill provides a unified interface for common video processing tasks:

  • Audio extraction: Extract audio tracks in various formats
  • Format conversion: Convert to MP4 (universal) or WebM (web-optimized)
  • Transcription: Speech-to-text with multiple output formats
  • Flexible: CLI arguments for model selection, language, and output formats

All operations are handled through a single, well-documented script with sensible defaults and comprehensive error handling.

Individual skills in this repo

This repo contains 14 individual skills — each has its own dedicated page.

bg-szy/TOP-SKILLS

Beads Viewer - Graph-aware triage engine for Beads projects. Computes PageRank, betweenness, critical path, and cycles. Use --robot-* flags for AI agents.

bg-szy/TOP-SKILLS

CI red? Call us. Pipeline fire brigade deploys. Use when user mentions CI failures, build errors, test failures, or pipeline issues. Do NOT load for: local builds, standard implementation work, reviews, or setup.

bg-szy/TOP-SKILLS

CASS Memory System - procedural memory for AI coding agents. Three-layer cognitive architecture with confidence decay, anti-pattern learning, cross-agent knowledge transfer, trauma guard safety system. Bun/TypeScript CLI.

bg-szy/TOP-SKILLS

Create high-converting, visually distinctive landing pages. Use when building marketing pages, product launches, SaaS homepages, or any single-page conversion-focused website. Guides section-by-section composition with anti-AI-slop principles.

bg-szy/TOP-SKILLS

Using the Wonda CLI to generate images, videos, music, and audio from the terminal — plus LinkedIn, Reddit, and X/Twitter research and automation

bg-szy/TOP-SKILLS

Sets up and runs AFL++ for multi-core fuzzing of C/C++ projects built with afl-clang-fast or afl-gcc-fast. Covers instrumentation modes, parallel main and secondary campaigns, persistent mode, corpus minimization, and crash triage. Use when scaling fuzzing across cores, fuzzing a mature C/C++ codebase, reading the afl-fuzz status screen, or moving on after libFuzzer has plateaued.

bg-szy/TOP-SKILLS

>- Scans a codebase for security vulnerabilities using CodeQL's interprocedural data flow and taint tracking analysis. Triggers on "run codeql", "codeql scan", "build codeql database", "SAST scan", "taint analysis", "dataflow analysis", or "find vulnerabilities in this repo". Covers Python, JavaScript/TypeScript, Go, Java/Kotlin, C/C++, C#, Ruby, and Swift. Supports "run all" (security-and-quality + security-experimental) and "important only" (high-precision) scan modes, and creates data extension models for project-specific sources and sinks. For fast single-file pattern matching, or when no build is available for a compiled language, use the semgrep skill; to parse SARIF that already exists rather than produce it, use the sarif-parsing skill.

bg-szy/TOP-SKILLS

Use this skill whenever the user wants to create, read, edit, or manipulate Word documents (.docx files) or Word templates (.dotx files). Triggers include: any mention of 'Word doc', 'word document', '.docx', '.dotx', or requests to produce professional documents with formatting like tables of contents, headings, page numbers, or letterheads. Also use when extracting or reorganizing content from .docx or .dotx files, inserting or replacing images in documents, performing find-and-replace in Word files, working with tracked changes or comments, or converting content into a polished Word document. If the user asks for a 'report', 'memo', 'letter', 'template', or similar deliverable as a Word or .docx file, use this skill. Do NOT use for PDFs, spreadsheets, Google Docs, or general coding tasks unrelated to document generation.

bg-szy/TOP-SKILLS

Enforces authenticated gh CLI workflows over unauthenticated curl, WebFetch, and MCP fetch patterns. Use when working with GitHub URLs, API access, pull requests, or issues.

bg-szy/TOP-SKILLS

Builds custom fuzzers with LibAFL, the modular Rust fuzzing library. Covers composing observers, feedbacks, mutators, schedulers, and executors into a fuzzer for targets the standard tools do not fit. Use when writing a bespoke fuzzer or mutator, fuzzing a non-standard target or architecture, implementing a fuzzing research idea, or when libFuzzer and AFL++ lack the control you need.

bg-szy/TOP-SKILLS

Use this skill whenever the user wants to do anything with PDF files. This includes reading or extracting text/tables from PDFs, combining or merging multiple PDFs into one, splitting PDFs apart, rotating pages, adding watermarks, creating new PDFs, filling PDF forms, encrypting/decrypting PDFs, extracting images, and OCR on scanned PDFs to make them searchable. If the user mentions a .pdf file or asks to produce one, use this skill.

bg-szy/TOP-SKILLS

Use this skill any time a .pptx or .potx file is involved in any way — as input, output, or both. This includes: creating slide decks, pitch decks, or presentations; reading, parsing, or extracting text from any .pptx or .potx file (even if the extracted content will be used elsewhere, like in an email or summary); editing, modifying, or updating existing presentations; combining or splitting slide files; working with templates (.potx), layouts, speaker notes, or comments. Trigger whenever the user mentions \"deck,\" \"slides,\" \"presentation,\" or references a .pptx or .potx filename, regardless of what they plan to do with the content afterward. If a .pptx or .potx file needs to be opened, created, or touched, use this skill.

bg-szy/TOP-SKILLS

Sets up and runs Ruzzy, Trail of Bits' coverage-guided Ruby fuzzer and the only production-ready one for the language. Covers harness structure, fuzzing pure Ruby and the native C extensions in gems, and sanitizer builds. Use when fuzzing a Ruby library or gem, testing a Ruby C extension for memory safety, or asking how to fuzz Ruby at all.

bg-szy/TOP-SKILLS

Use this skill any time a spreadsheet file is the primary input or output. This means any task where the user wants to: open, read, edit, or fix an existing .xlsx, .xlsm, .xltx, .csv, or .tsv file (e.g., adding columns, computing formulas, formatting, charting, cleaning messy data); create a new spreadsheet from scratch or from other data sources; or convert between tabular file formats. Trigger especially when the user references a spreadsheet file by name or path — even casually (like \"the xlsx in my downloads\") — and wants something done to it or produced from it. Also trigger for cleaning or restructuring messy tabular data files (malformed rows, misplaced headers, junk data) into proper spreadsheets. The deliverable must be a spreadsheet file. Do NOT trigger when the primary deliverable is a Word document, HTML report, standalone Python script, database pipeline, or Google Sheets API integration, even if tabular data is involved.

Related Skills