Communitygithub.com

AlemTuzlak/skills

Use when the user wants to transcribe a video or audio file to text with word-level timestamps — spins up a self-contained local Whisper (whisper.cpp) Docker service bundled in this skill and returns transcript.txt, transcript.srt, and word-level transcript.words.json. Also used as a building block by /produce-video. Triggers on "transcribe this video", "get a transcript", "transcribe-video".

skills 是什么?

skills is a Claude Code agent skill that use when the user wants to transcribe a video or audio file to text with word-level timestamps — spins up a self-contained local Whisper (whisper.cpp) Docker service bundled in this skill and returns transcript.txt, transcript.srt, and word-level transcript.words.json. Also used as a building block by /produce-video. Triggers on "transcribe this video", "get a transcript", "transcribe-video".

兼容平台~Claude Code~Codex CLI~Cursor
npx skills add https://github.com/AlemTuzlak/skills/tree/HEAD/skills/transcribe-video

在你喜欢的 AI 中提问

打开一个已预加载此 Agent Skill 的新对话。

文档

Transcribe Video

Transcribes any video/audio file locally using a bundled whisper.cpp service. No external project or cloud API. Returns word-level timestamps (needed for synced overlays/captions).

When to use

  • "Transcribe this clip / video / audio", "get me a transcript with timestamps".
  • As a sub-step of /produce-video.

How it runs

Everything is driven by the bundled runner — never ask the user to manage Docker:

node scripts/transcribe.mjs <path-to-video-or-audio> [--out <dir>] [--port 9111] [--language en] [--no-word-ts] [--task transcribe|translate]
node scripts/transcribe.mjs --stop      # stop the warm container

The runner:

  1. Checks the Docker daemon is reachable (fails loud with guidance if not).
  2. Builds the transcribe-video-whisper image from assets/whisper-service/ if it is missing (first build is slow: it compiles whisper.cpp and bakes the ggml-base.en.bin model).
  3. Starts the transcribe-video-whisper container (host port 9111 → container 9001); reuses it if already running, docker starts it if stopped.
  4. Waits for GET /healthz to report ok.
  5. Extracts a 16 kHz mono WAV from the input with local ffmpeg (whisper.cpp only decodes WAV), then POSTs it to /transcribe with word_ts=true.
  6. Writes transcript.txt, transcript.srt, transcript.words.json to --out (default: the input file's directory), and prints a JSON result line to stdout.

Outputs

  • transcript.txt — plain text.
  • transcript.srt — subtitle text.
  • transcript.words.json — [{ "word", "start", "end" }], seconds, time-ordered. This is the sync source for overlays/captions.

stdout result line (for programmatic callers like /produce-video):

{ "ok": true, "outDir": "...", "wordCount": 38, "segmentCount": 2, "files": { "txt": "transcript.txt", "srt": "transcript.srt", "words": "transcript.words.json" } }

Requirements

  • Docker is REQUIRED — the Whisper service runs in a container, so Docker Desktop must be installed AND running. The runner checks this first and fails loud with install/start guidance if Docker is missing or its daemon is down; there is no non-Docker fallback.
  • ffmpeg is REQUIRED (used to extract a 16 kHz mono WAV for Whisper) — the runner checks for it up front and fails loud if it is not on PATH.
  • Node ≥ 22.

Notes

  • The container is left running (--restart unless-stopped) for warm reuse; stop with --stop.
  • If word timestamps are requested but none come back, the runner fails loud (overlay sync depends on them).
  • See references/ for the Docker lifecycle, the API shape, and the vendoring provenance.

Individual skills in this repo

This repo contains 14 individual skills — each has its own dedicated page.

AlemTuzlak/skills

Use when the user wants to write a blog post about a feature, product change, PR, git diff, or any technical topic - accepts marketing briefs, PRs, git refs, codebase paths, or freeform descriptions as input

AlemTuzlak/skills

Use when the user wants to generate a changelog, release notes, or document what changed between versions, tags, or PRs

AlemTuzlak/skills

Use when writing, editing, or organizing documentation, when planning what docs a feature needs, and whenever planning or implementing a new feature or change in a repo (docs ship with the code). Also use when tempted to write docs without showing the discovered readers to the user, without asking for tone, or without loading simple-english and i-have-adhd. Triggers on "write docs for X", "document this feature", "add a guide", "update the docs", "reorganize the docs", "plan feature X", "implement X", or /docs.

AlemTuzlak/skills

Use when a bug is in play: a test fails, CI is red, an API returns the wrong result, a stack trace appears, or the user says it is broken or fix this. Don't use for a new feature with no failure, for types-only work, or for docs.

AlemTuzlak/skills

Use when the user invokes /i-have-adhd, says they have ADHD, or asks for ADHD-friendly output. Also used as a required writing filter by the docs skill. Don't use for marketing copy or after the user says "stop adhd mode" or "normal mode".

AlemTuzlak/skills

Use when a settled change must be turned into an ordered stack of small blocks before anyone implements. Don't use for unsettled intent, typos, comments, formatting, docs-only work, or writing the implementation itself.

AlemTuzlak/skills

Use when the user wants to write a product update email, feature announcement newsletter, or digest email for users or subscribers

AlemTuzlak/skills

Use when the user wants to turn a raw talking-head / screen-share recording into a finished, edited, annotated video plus a full content package. Removes silences, flags mistakes for the user to cut, transcribes, adds transcript-synced overlays (code, on-screen code highlights, word highlights, lists, comparisons, diagrams, section labels, punch-in zooms), renders with original audio, then generates blog/socials/YouTube content. Triggers on "produce a video", "edit my video", "annotate my recording", "/produce-video".

AlemTuzlak/skills

Use when the user runs /prove-it or says prove it, prove the changes, show me in the browser, or asks to prove a UI or API change. Don't use only because the agent is about to say done, for types-only work, or for docs with no behavior to prove.

AlemTuzlak/skills

Use when the user wants to write, draft, or author an RFC (Request for Comments) / technical design doc for a feature, change, or architectural decision. Interactively interviews the user, grounds the proposal in the actual codebase, presents 2-3 concrete API/code-snippet approaches to choose from, then writes a review-ready RFC. Triggers on "write an RFC", "draft an RFC", "RFC for X", "design doc for X", or /rfc.

AlemTuzlak/skills

Use when the user wants to deeply learn a new topic from scratch. Runs a pre-interview (current knowledge, end-goal proficiency, depth, practice load, background, scope), researches online (articles, niche-influencer blogs, canonical docs, subtopic landscape), then produces a structured markdown course with mandatory visual diagrams, evidence-based learning-science features (retrieval practice, spaced callbacks, worked-example fading, concept ledger, jargon gate, analogy hygiene), and a self-contained interactive HTML mini-course. Triggers on /teach-me, "teach me about X", "I want to learn X", "deep dive on X", "create a course on X", "study X with me".

AlemTuzlak/skills

Use when the change intent is already settled and the agent must map what a behavior change touches before an implementation plan or any code. Use for new features, bug fixes, and refactors that move a boundary. Don't use for typos, comments, formatting, lockfile-only diffs, docs with no code, or while the user is still deciding what they want.

AlemTuzlak/skills

Use when the user wants to write a video script for a product demo, feature walkthrough, launch video, or social media video about a feature or product change

AlemTuzlak/skills

Use when the user wants YouTube metadata for a video — a click-worthy title, an SEO/above-the-fold description, tags, and timestamped chapters. Accepts a transcript (ideally timestamped SRT), a topic, a PR, or a freeform description. Used standalone and by /produce-video. Triggers on "youtube title", "youtube description", "youtube chapters", "youtube tags", "youtube metadata", "youtube-copy".

相关技能