Communitygithub.com

JantonioFC/skillsbank

Extract key frames from MP4 videos at configurable intervals, run Tesseract OCR, and generate structured Markdown reports with video metadata and timestamped text transcripts.

Qu'est-ce que skillsbank ?

skillsbank is a Codex agent skill that extract key frames from MP4 videos at configurable intervals, run Tesseract OCR, and generate structured Markdown reports with video metadata and timestamped text transcripts.

Compatible avec~Claude Code✓Codex CLI~Cursor
npx skills add https://github.com/JantonioFC/skillsbank/tree/HEAD/skills/video-content-extractor

Demander à votre IA préférée

Ouvre une nouvelle conversation avec cette compétence d'agent déjà préchargée.

Documentation

Video Content Extractor

Overview

Automatically extracts key frames from MP4 video files at configurable time intervals, performs OCR text recognition on each frame, and generates a structured Markdown report. The report includes video metadata (duration, resolution, codecs) and frame-by-frame OCR transcripts with timestamp references.

This skill is designed for Codex CLI and requires FFmpeg and Tesseract OCR installed on the local machine.

When to Use This Skill

  • Use when you need to extract text content from video presentations, lectures, or screencasts.
  • Use when you want to create searchable transcripts from video files without embedded subtitles.
  • Use when you need to analyze video content programmatically and generate structured summaries.
  • Use when the user asks to "read what is on screen" or "extract the content from this video."

How It Works

Step 1: Analyze Video Metadata

The skill uses ffprobe to extract video metadata: duration, resolution, frame rate, codec information, and file size.

Step 2: Extract Key Frames

Using FFmpeg, the skill captures frames at the configured interval (default: every 30 seconds). Each frame is saved as a timestamped JPEG image.

Step 3: OCR Text Recognition

Each extracted frame is processed by Tesseract OCR. If the default PSM mode returns no meaningful text, it falls back to fully automatic page segmentation.

Step 4: Generate Markdown Report

All extracted data is assembled into a structured Markdown document.

Examples

Example 1: Basic Extraction

Agent prompt: Use the video-content-extractor skill to extract content from lecture.mp4

Output generates lecture.md and lecture_frames/ directory.

Example 2: Custom Interval

Parameters: video_path, output_dir, interval(seconds), lang Extract every 60 seconds with English-only OCR: python scripts/extract_video.py recording.mp4 ./output 60 eng

Example 3: Bilingual Content

Extract with default Chinese + English OCR: python scripts/extract_video.py lecture.mp4 . 15 chi_sim+eng

Best Practices

  • Use shorter intervals (10-15s) for fast-paced content with frequent text changes.
  • Use longer intervals (30-60s) for presentation slides or slow lectures to reduce duplicate frames.
  • For Chinese content, ensure Tesseract Chinese language pack is installed (chi_sim).

Limitations

  • Requires FFmpeg and Tesseract OCR to be installed and accessible via PATH.
  • Tesseract OCR accuracy depends on video quality, text size, and font clarity.
  • Does not extract audio or perform speech-to-text transcription.
  • Frame extraction is time-based (not scene-change-based), which may produce near-duplicate frames.
  • Large videos with short intervals can generate many frames - ensure sufficient disk space.

Security and Safety Notes

  • This skill only reads video files and writes extracted frames and Markdown reports.
  • It does NOT send any data over the network - all processing is local.
  • FFmpeg and Tesseract are invoked with fixed, pre-vetted arguments.
  • The skill does not modify or delete the original video file.

Common Pitfalls

  • Problem: Tesseract returns garbled text Solution: Ensure the correct language pack is installed. Run tesseract --list-langs to verify.

  • Problem: FFmpeg fails with "not found" Solution: Make sure FFmpeg is on PATH. Run ffmpeg -version to verify.

  • Problem: OCR is slow on large videos Solution: Increase the interval parameter to reduce frames processed.

Related Skills

  • @media-summarizer - For summarizing video content using visual and audio cues.
  • @document-ocr - For OCR on static images or scanned documents without video processing.

Individual skills in this repo

This repo contains 20 individual skills — each has its own dedicated page.

JantonioFC/skillsbank

AI-powered image editing with style transfer and object removal

JantonioFC/skillsbank

Internationalization and localization patterns. Detecting hardcoded strings, managing translations, locale files, RTL support.

JantonioFC/skillsbank

Analise e auditoria de editais de leilao judicial e extrajudicial. Riscos ocultos, clausulas perigosas, debitos, ocupante e classificacao da oportunidade.

JantonioFC/skillsbank

Embed Photopea in web apps using photopea.js. Covers embedding, file I/O, scripting, exporting, layers, text, filters, and the full Photoshop-compatible API.

JantonioFC/skillsbank

Best practices for Remotion - Video creation in React

JantonioFC/skillsbank

Generate walkthrough videos from Stitch projects using Remotion with smooth transitions, zooming, and text overlays

JantonioFC/skillsbank

Seek and analyze video content using Memories.ai Large Visual Memory Model for persistent video intelligence

JantonioFC/skillsbank

Writes complete, structured landing pages optimized for SEO ranking, AEO citation, and visitor conversion. Activate when the user wants to write or generate a landing page for a product, service, or offer.

JantonioFC/skillsbank

Video and audio perception, indexing, and editing. Ingest files/URLs/live streams, build visual/spoken indexes, search with timestamps, edit timelines, add overlays/subtitles, generate media, and create real-time alerts.

JantonioFC/skillsbank

Upload, stream, search, edit, transcribe, and generate AI video and audio using the VideoDB SDK.

JantonioFC/skillsbank

Skill for discovering and researching autonomous AI agents, tools, and ecosystems using the AgentFolio directory.

JantonioFC/skillsbank

Expert in building portfolios that actually land jobs and clients - not just showing work, but creating memorable experiences. Covers developer portfolios, designer portfolios, creative portfolios, and portfolios that convert visitors into opportunities.

JantonioFC/skillsbank

Tool lifecycle UI components for React/Next.js from ui.inference.sh. Display tool calls: pending, progress, approval required, results. Capabilities: tool status, progress indicators, approval flows, results display. Use for: showing agent tool calls, human-in-the-loop approvals, tool output. Triggers: tool ui, tool calls, tool status, tool approval, tool results, agent tools, mcp tools ui, function calling ui, tool lifecycle, tool pending

JantonioFC/skillsbank

Security audit, hardening, threat modeling (STRIDE/PASTA), Red/Blue Team, OWASP checks, code review, incident response, and infrastructure security for any project.

JantonioFC/skillsbank

When the user wants help with paid advertising campaigns on Google Ads, Meta (Facebook/Instagram), LinkedIn, Twitter/X, or other ad platforms. Also use when the user mentions 'PPC,' 'paid media,' 'ROAS,' 'CPA,' 'ad campaign,' 'retargeting,' 'audience targeting,' 'Google Ads,' 'Facebook ads,'...

JantonioFC/skillsbank

Create AI avatar and talking head videos with OmniHuman, Fabric, PixVerse via inference.sh CLI. Models: OmniHuman 1.5, OmniHuman 1.0, Fabric 1.0, PixVerse Lipsync. Capabilities: audio-driven avatars, lipsync videos, talking head generation, virtual presenters. Use for: AI presenters, expla...

JantonioFC/skillsbank

Create AI marketing videos for ads, promos, product launches, and brand content. Models: Veo, Seedance, Wan, FLUX for visuals, Kokoro for voiceover. Types: product demos, testimonials, explainers, social ads, brand videos. Use for: Facebook ads, YouTube ads, product launches, brand awarene...

JantonioFC/skillsbank

Generate AI videos with Google Veo, Seedance, Wan, Grok and 40+ models via inference.sh CLI. Models: Veo 3.1, Veo 3, Seedance 1.5 Pro, Wan 2.5, Grok Imagine Video, OmniHuman, Fabric, HunyuanVideo. Capabilities: text-to-video, image-to-video, lipsync, avatar animation, video upscaling, fole...

JantonioFC/skillsbank

Production-ready CI/CD configurations for Playwright — GitHub Actions, GitLab CI, CircleCI, Azure DevOps, Jenkins, Docker, parallel sharding, reporting, code coverage, and global setup/teardown.

JantonioFC/skillsbank

You are an expert copy editor specializing in marketing and conversion copy. Your goal is to systematically improve existing copy through focused editing passes while preserving the core message.

Skills associés