Communitygithub.com

jsath/claude-skills

Transcribe audio or video on your own machine with Whisper (mlx-whisper, faster-whisper or whisper.cpp) and get word-level timestamps as words.json, a plain .txt and an .srt. No API key, nothing uploaded. This is the transcription step the editing skills (edit-style, longform-edit, reel-captions) depend on. Use when asked to transcribe a take, get word timings, make an SRT, re-transcribe a final render for a caption sync check, or 'run whisper locally'.

claude-skills とは?

claude-skills is a Claude Code agent skill that transcribe audio or video on your own machine with Whisper (mlx-whisper, faster-whisper or whisper.cpp) and get word-level timestamps as words.json, a plain .txt and an .srt. No API key, nothing uploaded. This is the transcription step the editing skills (edit-style, longform-edit, reel-captions) depend on. Use when asked to transcribe a take, get word timings, make an SRT, re-transcribe a final render for a caption sync check, or 'run whisper locally'.

対応✓Claude Code~Codex CLI~Cursor
npx skills add https://github.com/jsath/claude-skills/tree/HEAD/skills/use-local-whisper

お気に入りのAIに質問する

このエージェントスキルを事前に読み込んだ状態で新しいチャットを開きます。

ドキュメント

Use local Whisper

One script, three backends, one output format. Everything downstream (cut lists, caption timing, sync gates) reads <stem>.words.json:

[
{"w": "Your", "s": 0.02, "e": 0.27},
{"w": "AI", "s": 0.27, "e": 0.4}
]

w = word as heard, s/e = start/end in seconds.

Requirements (Claude Code, on a computer)

  • ffmpeg (any build): brew install ffmpeg / apt install ffmpeg
  • Python 3.9+
  • One backend:
backendinstalldefault modeluse it when
mlxpip install mlx-whispermlx-community/whisper-large-v3-turboApple Silicon. Best word TIMING of the three
faster-whisperpip install faster-whisperlarge-v3-turboLinux/Windows, or CUDA
whisper-cppbrew install whisper-cpp (binary is whisper-cli)none, pass --modelno Python ML stack wanted

whisper.cpp model download (466 MB, English):

mkdir -p ~/whisper-models
curl -L -o ~/whisper-models/ggml-small.en.bin \
  https://huggingface.co/ggerganov/whisper.cpp/resolve/main/ggml-small.en.bin

Install check (prints which backend auto would use):

ffmpeg -version | head -1
python3 -c "import mlx_whisper" 2>/dev/null && echo mlx || \
python3 -c "import faster_whisper" 2>/dev/null && echo faster-whisper || \
command -v whisper-cli || echo "NO BACKEND: install one from the table"

Not available on Claude.ai (no shell). There, ask the user to upload an SRT instead.

Run it

python3 scripts/transcribe.py take.mov                       # auto backend, English
python3 scripts/transcribe.py take.mov --backend whisper-cpp --model ~/whisper-models/ggml-small.en.bin
python3 scripts/transcribe.py final.mp4 --out qa/ --words-per-cue 2
python3 scripts/transcribe.py interview.m4a --lang auto      # detect language

It extracts 16 kHz mono with ffmpeg first, so any container works (mov, mp4, m4a, wav). Output lands next to the input unless --out is given: <stem>.words.json, <stem>.txt, <stem>.srt.

Which model for which job

jobusewhy
Caption cue TIMING, sync gateslarge-v3-turbo (mlx or faster-whisper)small models drift 0.1 to 0.4 s per word on the same audio
Transcript text, cut lists, retake detectionsmall.en or turbotext quality is close; small.en is fast on CPU
Non-Englishlarge-v3-turbo with --lang <code>.en models are English only

Rules that save a re-do

  1. Text from the script, timing from Whisper. Whisper mishears names ("Cloud" for Claude). For captions, align the script's words to words.json and keep the script spelling (edit-style tools/make_cues.py does this). Keep a fix map ({"cloud code": "Claude Code"}) for anything with no script.
  2. Write numbers as digits in scripts. Whisper emits 2,000, not "two thousand", so a spelled-out script fails to align.
  3. Re-transcribe the FINAL render for the published SRT and any sync check. Times from the raw take are wrong after every cut.
  4. Whisper stamps the first word near 0.00 even when speech starts later. Find the real onset from audio energy before trimming the head, then re-transcribe the first 2 s to prove the first word survived.
  5. Long files (over ~30 min): fine with all three backends; expect roughly real-time / 10 on Apple Silicon with turbo. Run under nice -n 15 if someone is recording on the same machine.

Failure handling

symptomfix
no local Whisper foundinstall one backend from the table
whisper-cpp needs --modeldownload a ggml model (above) or set WHISPER_MODEL
whisper-cli not found but it is installedset WHISPER_BIN=/path/to/whisper-cli (launchd/cron PATHs often miss /opt/homebrew/bin)
empty or 0-word outputcheck the file has audio: ffprobe -v error -select_streams a -show_entries stream=codec_name -of csv=p=0 FILE
wrong language / gibberishpass --lang <code> explicitly
words glued together ("2,000files")whisper.cpp quirk on some builds; use mlx or faster-whisper for timing work

Output to report

whisper-cpp: 17 words, 5.68s -> o/take.words.json / .txt / .srt

Then hand the words.json path to the next skill (edit-style, longform-edit, reel-captions).

Individual skills in this repo

This repo contains 3 individual skills — each has its own dedicated page.

jsath/claude-skills

Writes captions for a finished video from its transcript: pulls the line worth repeating, writes first-line options, writes per-platform versions (Instagram, TikTok, YouTube title and description, others on request), keeps the creator's voice using their best past captions, and checks the call to action matches what the video actually says. Use when a creator asks for captions, a caption hook, a YouTube title and description, or a CTA check. Drafts only; never posts or schedules.

jsath/claude-skills

Write CAPTIONS-PLATFORMS.md for one finished short video: a per-platform post caption for Instagram, TikTok, YouTube Shorts (title + description), Facebook, Threads, Bluesky and X, each within that platform's limits, with your comment-keyword CTA on the platforms where it belongs and never on X. Keyword, CTA, platforms, hashtags and limits come from a config file; a checker script fails the file on any miss. Use after an edit is finished, when asked to 'write the captions', 'caption this reel for every platform', or before scheduling.

jsath/claude-skills

Runs five agents that keep a YouTube channel's titles and descriptions matched to what people search for now: a Channel Reader that pulls the user's own video list and stats with their YouTube Data API key, YouTube Analytics OAuth or a pasted YouTube Studio export, a Niche Watcher that finds what is working right now for similar-size channels from public search results the API returns or the user provides, a Title Writer that drafts titles, descriptions and chapters for new videos, a Back Catalogue agent that picks old videos worth a refresh and drafts new versions, and a Change Tracker that records every change the user makes and compares before and after on matched windows. Use when a creator asks for YouTube SEO, better titles or descriptions, what is working in their niche, a back catalogue refresh, or says \"seo run\". Never invents search volumes or stats and never changes a video; drafts only.

関連スキル