Communitygithub.com

aiskillstore/marketplace

Process video files with audio extraction, format conversion (mp4, webm), and Whisper transcription. Use when user mentions video conversion, audio extraction, transcription, mp4, webm, ffmpeg, or whisper transcription.

O que é marketplace?

marketplace is a Claude Code agent skill that process video files with audio extraction, format conversion (mp4, webm), and Whisper transcription. Use when user mentions video conversion, audio extraction, transcription, mp4, webm, ffmpeg, or whisper transcription.

Funciona com✓Claude Code✓Codex CLI~Cursor
npx skills add https://github.com/aiskillstore/marketplace/tree/HEAD/skills/egadams/video-processor

Perguntar na sua IA favorita

Abre um novo chat com esta habilidade de agente já pré-carregada.

Documentação

Video Processor

Instructions

This skill provides video processing utilities including audio extraction, format conversion, and audio transcription using FFmpeg and OpenAI's Whisper model.

Prerequisites

Required tools (must be installed in your environment):

  • FFmpeg: Multimedia framework for video/audio processing

    # macOS
    brew install ffmpeg
    
    # Ubuntu/Debian
    apt-get install ffmpeg
    
    # Verify installation
    ffmpeg -version
    
  • OpenAI Whisper: Speech-to-text transcription model

    # Install via pip
    pip install -U openai-whisper
    
    # Verify installation
    whisper --help
    

Python packages (included in script via PEP 723):

  • click (CLI framework)
  • ffmpeg-python (Python wrapper for FFmpeg)

Workflow

Use the scripts/video_processor.py script for all video processing tasks. The script provides a simple CLI with the following commands:

1. Extract Audio from Video

Extract the audio track from a video file:

uv run .claude/skills/video-processor/scripts/video_processor.py extract-audio input.mp4 output.wav

Options:

  • --format: Output audio format (default: wav). Supports: wav, mp3, aac, flac
  • Output is suitable for transcription or standalone audio use

2. Convert Video to MP4

Convert any video file to MP4 format:

uv run .claude/skills/video-processor/scripts/video_processor.py to-mp4 input.avi output.mp4

Options:

  • --codec: Video codec (default: libx264). Common options: libx264, libx265, h264
  • --preset: Encoding speed/quality preset (default: medium). Options: ultrafast, fast, medium, slow, veryslow

3. Convert Video to WebM

Convert any video file to WebM format (web-optimized):

uv run .claude/skills/video-processor/scripts/video_processor.py to-webm input.mp4 output.webm

Options:

  • --codec: Video codec (default: libvpx-vp9). Options: libvpx, libvpx-vp9
  • WebM is optimized for web playback and streaming

4. Transcribe Audio with Whisper

Transcribe audio or video files to text using OpenAI's Whisper model:

# Transcribe video file (audio will be extracted automatically)
uv run .claude/skills/video-processor/scripts/video_processor.py transcribe input.mp4 transcript.txt

# Transcribe audio file directly
uv run .claude/skills/video-processor/scripts/video_processor.py transcribe audio.wav transcript.txt

Options:

  • --model: Whisper model size (default: base). Options:
    • tiny: Fastest, lowest accuracy (~1GB RAM)
    • base: Fast, good accuracy (~1GB RAM) [DEFAULT]
    • small: Balanced (~2GB RAM)
    • medium: High accuracy (~5GB RAM)
    • large: Best accuracy, slowest (~10GB RAM)
  • --language: Language code (default: auto-detect). Examples: en, es, fr, de, zh
  • --format: Output format (default: txt). Options: txt, srt, vtt, json

Transcription workflow:

  1. If input is video, FFmpeg extracts audio to temporary WAV file
  2. Whisper processes the audio file
  3. Transcription is saved in requested format
  4. Temporary files are cleaned up automatically

5. Combined Workflow Example

Process a video end-to-end:

# 1. Extract audio for analysis
uv run .claude/skills/video-processor/scripts/video_processor.py extract-audio lecture.mp4 lecture.wav

# 2. Transcribe to SRT subtitles
uv run .claude/skills/video-processor/scripts/video_processor.py transcribe lecture.mp4 lecture.srt --format srt --model small

# 3. Convert to web format
uv run .claude/skills/video-processor/scripts/video_processor.py to-webm lecture.mp4 lecture.webm

Key Technical Details

FFmpeg and Whisper Integration:

  • FFmpeg doesn't transcribe audio itself - it prepares audio for external transcription
  • The workflow is: Extract audio (FFmpeg) → Transcribe (Whisper) → Optional: Re-integrate with video
  • FFmpeg can pipe audio directly to Whisper for real-time processing (advanced use case)

Audio Format for Transcription:

  • Whisper works best with WAV or MP3 formats
  • Sample rate: 16kHz is optimal (script handles conversion automatically)
  • The script extracts audio with optimal settings for Whisper

Output Formats:

  • txt: Plain text transcript
  • srt: SubRip subtitle format (includes timestamps)
  • vtt: WebVTT subtitle format (web standard)
  • json: Detailed JSON with word-level timestamps

Error Handling

The script includes comprehensive error handling:

  • Validates input files exist
  • Checks FFmpeg and Whisper are installed
  • Provides clear error messages for missing dependencies
  • Handles temporary file cleanup on errors

Performance Tips

  • Use tiny or base models for quick drafts
  • Use small or medium for production transcriptions
  • Use large only when maximum accuracy is required
  • For long videos, consider extracting audio first, then transcribe in segments
  • WebM conversion with VP9 takes longer but produces smaller files

Examples

Example 1: Quick Video to MP4 Conversion

User request:

I have an AVI file from my old camera. Can you convert it to MP4?

You would:

  1. Use the to-mp4 command with default settings:
    uv run .claude/skills/video-processor/scripts/video_processor.py to-mp4 old_video.avi output.mp4
    
  2. Confirm the conversion completed successfully
  3. Inform the user about the output file location

Example 2: Extract Audio and Transcribe

User request:

I recorded a lecture video and need a transcript. Can you extract the audio and transcribe it?

You would:

  1. First extract the audio:
    uv run .claude/skills/video-processor/scripts/video_processor.py extract-audio lecture.mp4 lecture.wav
    
  2. Then transcribe using the base model (good balance of speed/accuracy):
    uv run .claude/skills/video-processor/scripts/video_processor.py transcribe lecture.mp4 transcript.txt --model base
    
  3. Share the transcript.txt file with the user

Example 3: Create Web-Optimized Video with Subtitles

User request:

I need to put this video on my website with subtitles. Can you help?

You would:

  1. Convert to WebM for web optimization:
    uv run .claude/skills/video-processor/scripts/video_processor.py to-webm presentation.mp4 presentation.webm
    
  2. Generate SRT subtitle file:
    uv run .claude/skills/video-processor/scripts/video_processor.py transcribe presentation.mp4 subtitles.srt --format srt --model small
    
  3. Inform user they now have:
    • presentation.webm (web-optimized video)
    • subtitles.srt (subtitle file for embedding)

Example 4: High-Quality Transcription with Language Specification

User request:

I have a Spanish interview video that needs an accurate transcript for publication.

You would:

  1. Use a larger model with language specified for best accuracy:
    uv run .claude/skills/video-processor/scripts/video_processor.py transcribe interview.mp4 transcript.txt --model medium --language es
    
  2. Optionally create SRT for review:
    uv run .claude/skills/video-processor/scripts/video_processor.py transcribe interview.mp4 transcript.srt --format srt --model medium --language es
    
  3. Review the transcript with the user and make any necessary corrections

Example 5: Batch Processing Multiple Videos

User request:

I have a folder of training videos that all need to be converted to WebM and transcribed.

You would:

  1. List all video files in the directory:
    ls training_videos/*.mp4
    
  2. For each video file, run the conversion and transcription:
    # For each video: video1.mp4, video2.mp4, etc.
    uv run .claude/skills/video-processor/scripts/video_processor.py to-webm training_videos/video1.mp4 output/video1.webm
    uv run .claude/skills/video-processor/scripts/video_processor.py transcribe training_videos/video1.mp4 output/video1.txt --model base
    
    # Repeat for each file
    
  3. Confirm all conversions and transcriptions completed
  4. Provide summary of output files

Summary

The video-processor skill provides a unified interface for common video processing tasks:

  • Audio extraction: Extract audio tracks in various formats
  • Format conversion: Convert to MP4 (universal) or WebM (web-optimized)
  • Transcription: Speech-to-text with multiple output formats
  • Flexible: CLI arguments for model selection, language, and output formats

All operations are handled through a single, well-documented script with sensible defaults and comprehensive error handling.

Individual skills in this repo

This repo contains 14 individual skills — each has its own dedicated page.

aiskillstore/marketplace

Create high-converting, visually distinctive landing pages. Use when building marketing pages, product launches, SaaS homepages, or any single-page conversion-focused website. Guides section-by-section composition with anti-AI-slop principles.

aiskillstore/marketplace

Create beautiful visual art in .png and .pdf documents using design philosophy. You should use this skill when the user asks to create a poster, piece of art, design, or other static piece. Create original visual designs, never copying existing artists' work to avoid copyright violations.

aiskillstore/marketplace

Using the Wonda CLI to generate images, videos, music, and audio from the terminal — plus LinkedIn, Reddit, and X/Twitter research and automation

aiskillstore/marketplace

Comprehensive document creation, editing, and analysis with support for tracked changes, comments, formatting preservation, and text extraction. When Claude needs to work with professional documents (.docx files) for: (1) Creating new documents, (2) Modifying or editing content, (3) Working with tracked changes, (4) Adding comments, or any other document tasks

aiskillstore/marketplace

Generate or edit raster images when the task benefits from AI-created bitmap visuals such as photos, illustrations, textures, sprites, mockups, or transparent-background cutouts. Use when Codex should create a brand-new image, transform an existing image, or derive visual variants from references, and the output should be a bitmap asset rather than repo-native code or vector. Do not use when the task is better handled by editing existing SVG/vector/code-native assets, extending an established icon or logo system, or building the visual directly in HTML/CSS/canvas.

aiskillstore/marketplace

A set of resources to help me write all kinds of internal communications, using the formats that my company likes to use. Claude should use this skill whenever asked to write some sort of internal communications (status reports, leadership updates, 3P updates, company newsletters, FAQs, incident reports, project updates, etc.).

aiskillstore/marketplace

Guide for creating high-quality MCP (Model Context Protocol) servers that enable LLMs to interact with external services through well-designed tools. Use when building MCP servers to integrate external APIs or services, whether in Python (FastMCP) or Node/TypeScript (MCP SDK).

aiskillstore/marketplace

Use for Codex models/pricing, scheduled tasks, skills, settings, setup, troubleshooting, customization, automations, and self-knowledge—including 'you,' 'your,' 'this app,' or 'this coding agent' when they refer to Codex—and for OpenAI APIs/products and ChatGPT Work. Also use for model choice/migration, prompting, SDKs, Responses, Realtime, agents, evals, and Chat/Work/Codex comparisons. Do not use for generic app/software tasks that merely mention Codex.

aiskillstore/marketplace

Comprehensive PDF manipulation toolkit for extracting text and tables, creating new PDFs, merging/splitting documents, and handling forms. When Claude needs to fill in a PDF form or programmatically process, generate, or analyze PDF documents at scale.

aiskillstore/marketplace

Presentation creation, editing, and analysis. When Claude needs to work with presentations (.pptx files) for: (1) Creating new presentations, (2) Modifying or editing content, (3) Working with layouts, (4) Adding comments or speaker notes, or any other presentation tasks

aiskillstore/marketplace

Perform a read-only, defect-first review of a specified code change and return every actionable finding. Use when another agent delegates review of uncommitted changes, a base-branch diff, a commit, or custom review instructions.

aiskillstore/marketplace

Create or update a Codex skill with appropriately scoped instructions and any needed supporting resources.

aiskillstore/marketplace

Toolkit for styling artifacts with a theme. These artifacts can be slides, docs, reportings, HTML landing pages, etc. There are 10 pre-set themes with colors/fonts that you can apply to any artifact that has been creating, or can generate a new theme on-the-fly.

aiskillstore/marketplace

Use this skill any time a spreadsheet file is the primary input or output. This means any task where the user wants to: open, read, edit, or fix an existing .xlsx, .xlsm, .csv, or .tsv file (e.g., adding columns, computing formulas, formatting, charting, cleaning messy data); create a new spreadsheet from scratch or from other data sources; or convert between tabular file formats. Trigger especially when the user references a spreadsheet file by name or path — even casually (like \"the xlsx in my downloads\") — and wants something done to it or produced from it. Also trigger for cleaning or restructuring messy tabular data files (malformed rows, misplaced headers, junk data) into proper spreadsheets. The deliverable must be a spreadsheet file. Do NOT trigger when the primary deliverable is a Word document, HTML report, standalone Python script, database pipeline, or Google Sheets API integration, even if tabular data is involved.

Habilidades Relacionadas