Communitygithub.com

ever-just/agentskills

Extract actionable intelligence from video content across YouTube, Vimeo, and company websites. Video reveals what text sources cannot: facility scale, team composition, product demonstrations, signage, event attendance, and operational reality.

agentskills 是什麼?

agentskills is a Codex agent skill that extract actionable intelligence from video content across YouTube, Vimeo, and company websites. Video reveals what text sources cannot: facility scale, team composition, product demonstrations, signage, event attendance, and operational reality.

相容平台~Claude Code✓Codex CLI~Cursor
npx skills add https://github.com/ever-just/agentskills/tree/HEAD/skills/video-intelligence

在你喜歡的 AI 中提問

開啟一個已預先載入此 Agent Skill 的新對話。

說明文件

Video Intelligence Extraction

Overview

Extract actionable intelligence from video content across YouTube, Vimeo, and company websites. Video reveals what text sources cannot: facility scale, team composition, product demonstrations, signage, event attendance, and operational reality.

This skill focuses on metadata + caption extraction (fast, free, no storage cost) and only downloads actual video files when visual analysis is required.


When to Use

  • Target company has a YouTube channel (even a small one — Shorts count)
  • Founder or leadership appears in podcast/interview videos
  • Product demos or walkthroughs exist on any platform
  • Event recordings show the company's booth, presentations, or attendance
  • You need to verify claims (facility size, team count, product range) with visual evidence
  • Company website has embedded video content

Prerequisites

ToolInstallPurpose
yt-dlppip install yt-dlp or brew install yt-dlpVideo metadata/caption download
Python 3.10+SystemVTT parsing, automation
Whisper (optional)pip install openai-whisperTranscription when no captions exist
ffmpeg (optional)brew install ffmpegAudio extraction for Whisper

Phase 1: Discovery

YouTube Channel Enumeration

# List ALL videos on a channel (returns video IDs + titles, no download)
yt-dlp --flat-playlist --print "%(id)s %(title)s" "https://www.youtube.com/@CHANNEL_HANDLE"

# Include upload dates for timeline analysis
yt-dlp --flat-playlist --print "%(upload_date)s %(id)s %(title)s" "https://www.youtube.com/@CHANNEL_HANDLE"

YouTube Search

# Search YouTube (top 10 results)
yt-dlp --flat-playlist --print "%(id)s %(title)s" "ytsearch10:COMPANY_NAME"

# Broader search (top 30)
yt-dlp --flat-playlist --print "%(id)s %(title)s" "ytsearch30:COMPANY_NAME founder interview"

Other Discovery Methods

SourceMethod
Google Video searchsite:youtube.com "Company Name"
Company websiteInspect source for <iframe src="youtube.com/embed/...">
Vimeoyt-dlp --flat-playlist "https://vimeo.com/USER"
LinkedInManual browse — videos cannot be enumerated programmatically
Event platformsCheck Sched/Emamo session pages for embedded video links

Discovery Checklist

  • Target company's own YouTube channel
  • Founder/CEO personal channel or guest appearances
  • Industry podcast channels where target was interviewed
  • Event/conference channels with target's presentations
  • Partner/supplier channels mentioning target
  • Local news/media channels covering target

Phase 2: Capture

Metadata + Captions (Default — No Video Download)

# Entire channel — metadata + auto-captions, no video files
yt-dlp --write-info-json --write-auto-sub --sub-lang en --skip-download \
  --sleep-interval 2 "https://www.youtube.com/@CHANNEL"

# Single video
yt-dlp --write-info-json --write-auto-sub --sub-lang en --skip-download \
  "https://www.youtube.com/watch?v=VIDEO_ID"

# Multiple specific videos from a list
yt-dlp --write-info-json --write-auto-sub --sub-lang en --skip-download \
  --sleep-interval 2 -a video_urls.txt

Key yt-dlp Flags

FlagPurpose
--write-info-jsonSaves full metadata (title, date, duration, description, tags, view count)
--write-auto-subDownloads auto-generated captions
--sub-lang enEnglish captions only
--skip-downloadNo video/audio file — metadata only
--sleep-interval 2Rate limiting (critical for channels with 30+ videos)
--flat-playlistList mode — enumerate without downloading anything
-f "bestvideo[height<=720]+bestaudio"Download video at reasonable quality
-x --audio-format mp3Extract audio only (for Whisper)

Download Video Files (When Visual Analysis Needed)

# 720p max (good balance of quality vs. file size)
yt-dlp -f "bestvideo[height<=720]+bestaudio" --merge-output-format mp4 VIDEO_URL

# Audio only for Whisper transcription
yt-dlp -x --audio-format mp3 VIDEO_URL

Output File Structure

project/
├── raw/video/
│   ├── Title of Video [VIDEO_ID].info.json    # Metadata
│   ├── Title of Video [VIDEO_ID].en.vtt       # Auto-captions (VTT format)
│   └── Title of Video [VIDEO_ID].mp4          # Video file (only if downloaded)
├── transcripts/
│   ├── VIDEO_ID_transcript.txt                # Parsed plain text
│   └── VIDEO_ID_transcript.txt
└── analysis/
    └── video_intelligence_summary.md

Phase 3: VTT Parsing

Convert VTT Captions to Plain Text

import re
import glob

def vtt_to_text(vtt_path):
    """Parse VTT subtitle file to clean plain text."""
    with open(vtt_path) as f:
        content = f.read()
    # Remove WEBVTT header and metadata
    content = re.sub(r'WEBVTT.*?\n', '', content, flags=re.DOTALL)
    # Remove timestamp lines
    content = re.sub(r'\d{2}:\d{2}:\d{2}\.\d+ --> .*', '', content)
    # Remove HTML-style formatting tags
    content = re.sub(r'<[^>]+>', '', content)
    # Deduplicate consecutive repeated lines and join
    lines = []
    for l in content.split('\n'):
        l = l.strip()
        if l and not l.isdigit() and (not lines or l != lines[-1]):
            lines.append(l)
    return ' '.join(lines)


def batch_parse_vtts(vtt_dir, output_dir):
    """Parse all VTT files in a directory to plain text."""
    import os
    os.makedirs(output_dir, exist_ok=True)
    for vtt_file in glob.glob(f"{vtt_dir}/*.vtt"):
        text = vtt_to_text(vtt_file)
        # Extract video ID from filename pattern "Title [VIDEO_ID].en.vtt"
        video_id = re.search(r'\[([^\]]+)\]', vtt_file)
        if video_id:
            out_name = f"{video_id.group(1)}_transcript.txt"
        else:
            out_name = os.path.basename(vtt_file).replace('.en.vtt', '_transcript.txt')
        with open(os.path.join(output_dir, out_name), 'w') as f:
            f.write(text)
        print(f"Parsed: {out_name} ({len(text.split())} words)")

Phase 4: Whisper Fallback (No Captions Available)

When auto-captions don't exist (common for Shorts, older videos, non-English content):

# Download audio only
yt-dlp -x --audio-format mp3 -o "%(id)s.%(ext)s" VIDEO_URL

# Transcribe with Whisper
whisper VIDEO_ID.mp3 --model base --language en --output_format txt

# For better proper noun accuracy (3-5x slower):
whisper VIDEO_ID.mp3 --model medium --language en --output_format txt

When Captions Typically Don't Exist

  • YouTube Shorts (< 60 seconds)
  • Videos uploaded before ~2015 (pre-auto-caption era)
  • Unlisted/private videos
  • Music-heavy content with minimal speech
  • Non-English content without specified language

Phase 5: Intelligence Extraction

Metadata Intelligence (from info.json)

Every .info.json file contains:

FieldIntelligence value
upload_dateTimeline of company activity and content strategy
durationShort = promo; Long = substantive interview
descriptionOften contains links, names, timestamps, partner mentions
tagsSEO strategy, self-identified categories
view_countWhich content resonates with their audience
like_countEngagement quality signal
channelWho published it (their channel vs. guest appearance)

Transcript Intelligence Checklist

  • Named people (employees, partners, clients)
  • Revenue/growth figures mentioned casually in interviews
  • Future plans or strategy discussed
  • Pain points or challenges admitted
  • Competitor mentions (positive or negative)
  • Hiring plans or team size references
  • Product roadmap or feature announcements
  • Customer names or case studies mentioned verbally

Visual Intelligence Checklist (Requires Video Download)

  • Facility/warehouse size and condition
  • Number of employees visible (team size proxy)
  • Products on shelves or in demos (inventory breadth)
  • Equipment and machinery (capital investment signals)
  • Signage, branding, and logo usage
  • Event booth size and design quality
  • Vehicle fleet (delivery capability)
  • Customer interactions caught on camera
  • Geographic/location identifiers in background

Edge Cases and Gotchas

  1. YouTube rate limiting — Use --sleep-interval 2 between downloads. DNS errors appear after 30+ rapid requests. For large channels, break into batches of 20.

  2. Channel download filenames — yt-dlp names files as "Title [VIDEO_ID].ext" when downloading from channels. Spaces and special characters in titles break shell scripts. Use -o "%(id)s.%(ext)s" for clean filenames.

  3. Duplicate uploads — Same interview uploaded under different channels or titles. Cross-reference by duration + upload_date to catch duplicates before spending time analyzing both.

  4. LinkedIn native videos — CANNOT be downloaded programmatically. LinkedIn CDN URLs are authenticated and expire within hours. Flag as manual task: "Save as Web Complete" in browser, or screenshot key frames.

  5. Shorts vs. full videos — Shorts URL format is youtube.com/shorts/ID but yt-dlp handles both formats transparently. Don't skip Shorts — they often show products, facilities, and team culture.

  6. Unlisted videos — Not in channel listings but accessible via direct URL. Search engines may have indexed them. Check Google cache and Wayback Machine for unlisted URLs that were once linked publicly.

  7. Auto-caption quality — No punctuation, sometimes garbled proper nouns (Celina → Selena, BROGAV → ProGraph). 95%+ word accuracy for standard English speech. Always verify company/person names against known-good sources.

  8. Upload date vs. content date — Videos may be uploaded weeks or months after recording. Check description and audio cues for actual recording date. Event videos especially lag.

  9. info.json is gold — Always capture with --write-info-json. The description field often contains timestamps, guest names, links to resources, and partner mentions that aren't in the video itself.

  10. VTT timestamps are valuable — The plain-text conversion loses temporal information. Keep original VTT files as source of truth. Timestamps let you find specific moments for visual verification.


Anti-Patterns

Don'tDo instead
Ignore Shorts because they're shortReview Shorts for visual intel (products, facility, team)
Assume all videos have captionsCheck; fall back to Whisper for uncaptioned content
Download video files by defaultStart with --skip-download; only get video for visual analysis
Try to automate LinkedIn video downloadFlag as manual task; CDN URLs are authenticated
Trust upload_date as content dateCross-reference with description, audio cues, event dates
Download entire channel at once without rate limitingUse --sleep-interval 2 and batch into groups of 20
Parse only transcripts, ignore info.jsonMetadata contains names, links, tags not in the audio
Treat video as inferior to text sourcesVideo uniquely confirms physical reality (facility, team, products)

Decision Tree

Content is on YouTube?
├── YES → yt-dlp --write-info-json --write-auto-sub --skip-download
│   ├── Has captions? → Parse VTT to text → Extract intelligence
│   └── No captions?
│       ├── Video > 60s with speech? → yt-dlp -x → Whisper → Analyze
│       └── Short/visual only? → Download video → Visual analysis
└── NO →
    ├── Vimeo → yt-dlp supports it (same workflow as YouTube)
    ├── LinkedIn video → Manual "Save as Web Complete" (flag as manual task)
    ├── Company website embed → Extract iframe URL → usually YouTube/Vimeo
    └── Other platform → Check yt-dlp supported sites list

Real-World Results (BROGAV Case Study)

MetricValue
Total videos discovered45 (9 full-length, 36 Shorts)
VTT caption files captured10
Transcripts parsed to text10
Substantive interviews analyzed2 (5,000+ words each)
Podcast via RSS/Whisper fallback1 (6,489 words)
Duplicate videos caught1 (same interview, two channels)
Unique intelligence extractedFounder origin story, revenue signals, growth plans, supplier relationships, facility details, team culture
Visual intelligence from ShortsProduct range, warehouse layout, branded cabinet line, seasonal inventory

Integration with Other Skills

SkillIntegration point
deep-researchVideo is Phase 3/4 source alongside web search
intelligence-dossierVideo transcripts feed People, Products, Financial Signals sections
client-discovery-osintCustomer testimonial videos name clients directly
supplier-verificationProduct demo videos show supplier logos and branded products
era-validated-linkedin-analysisVideo upload dates cross-reference LinkedIn activity timeline
contact-sheet-image-analysisVideo screenshots can be batch-analyzed via thumbnail grids

Individual skills in this repo

This repo contains 8 individual skills — each has its own dedicated page.

ever-just/agentskills

> Declarative animations for React. Springs, gestures, layout animations, and page transitions. GPU-accelerated, production-ready.

ever-just/agentskills

> Create animated logos, icons, and micro-interactions as lightweight Lottie JSON files. No After Effects needed — build animations entirely in code.

ever-just/agentskills

> A TypeScript library using generators to program animations with a real-time editor. Ideal for explainer videos, code walkthroughs, and educational content.

ever-just/agentskills

> Add animated subtitles and captions to Remotion videos from .SRT files. TikTok-style word-by-word highlights, karaoke effects, and more.

ever-just/agentskills

> Create videos programmatically using React components. The most mature framework for code-driven video production.

ever-just/agentskills

> Copy-paste video components for Remotion. Intros, text effects, transitions, and full scene templates. Like shadcn/ui but for video.

ever-just/agentskills

Optimize and embed an existing video as a fast, crisp, autoplaying hero/background/loop on a website. Use when adding or replacing a video on a marketing site and it is too heavy, will not autoplay (especially on mobile/Safari), shows a codec error, or looks soft — to transcode it to web-safe H.264 with the right pixel format, faststart, retina-aware sizing, and a poster frame, then embed and verify it actually plays. This is optimization + embedding of footage you already have, not creating/animating a video (for that use remotion/moviepy/motion-canvas). Use when the user says "add this video to the site", "the hero video is too big / won't play / isn't crisp", "replace the homepage video", or "make this loop on the landing page".

ever-just/agentskills

Automatically fetch YouTube video transcripts, generate structured summaries, and send full transcripts to messaging platforms. Detects YouTube URLs and provides metadata, key insights, and downloadable transcripts.

相關技能