Communitygithub.com

SkillMedev/skills

Produce an accurate, properly timed caption track (SRT or WebVTT) from a video's audio - transcribing or aligning to the voiceover script, timing cues to speech, and enforcing line-length and reading-speed rules so captions are readable and in sync. Use when someone says "generate captions", "make subtitles from the audio", "transcribe and caption this", "export an SRT or VTT", "the captions are out of sync", or "the subtitles flash by too fast to read". Do NOT use to burn-in, style, reframe, or position an existing caption track for a platform (9:16, brand styling, safe areas) - that is social-video-formatter; do NOT use to animate text word-by-word as a motion graphic - that is kinetic-typography; do NOT use to write the spoken script in the first place - that is narration-script.

skills 是什么?

skills is a Claude Code agent skill that produce an accurate, properly timed caption track (SRT or WebVTT) from a video's audio - transcribing or aligning to the voiceover script, timing cues to speech, and enforcing line-length and reading-speed rules so captions are readable and in sync. Use when someone says "generate captions", "make subtitles from the audio", "transcribe and caption this", "export an SRT or VTT", "the captions are out of sync", or "the subtitles flash by too fast to read". Do NOT use to burn-in, style, reframe, or position an existing caption track for a platform (9:16, brand styling, safe areas) - that is social-video-formatter; do NOT use to animate text word-by-word as a motion graphic - that is kinetic-typography; do NOT use to write the spoken script in the first place - that is narration-script.

兼容平台~Claude Code~Codex CLI~Cursor
npx skills add https://github.com/SkillMedev/skills/tree/HEAD/skills/captions-from-transcript

在你喜欢的 AI 中提问

打开一个已预加载此 Agent Skill 的新对话。

文档

Captions From Transcript

Captions are not optional polish - they are accessibility, they hold viewers watching muted, and for tutorials they reinforce exact UI terms. This skill produces a clean, accurately timed caption file. Styling and burn-in happen downstream.

Source the most accurate text first

Accuracy comes from where you start:

  • If a narration script exists, use it as ground truth. It is already correct on technical terms and UI labels. Align it to the audio rather than re-transcribing from scratch.
  • Otherwise transcribe the audio (Whisper, a platform's auto-caption, or any ASR), then correct it against the video. ASR reliably mangles product names, code, and acronyms - fix every one to match what is on screen exactly.

Never ship raw ASR output. The errors are always in the highest-value words.

Time the cues to speech

  • Each cue appears as its line is spoken and clears when it ends - align to speech, not to arbitrary intervals.
  • Minimum ~1 second on screen (even for a short cue), maximum ~7 seconds. Split anything longer.
  • No gaps mid-sentence; small gaps between sentences are fine and aid readability.

Line and reading rules

  • 1-2 lines per cue, never 3.
  • ~32-42 characters per line. Past that it crowds the frame and overruns safe areas.
  • Reading speed ≤ ~17 characters/second (≈160-180 wpm). If a cue exceeds it, the words flash by - split the cue or extend its duration.
  • Break lines at clause boundaries, never mid-phrase. Keep "to the Settings page" together; don't strand "the" alone on a line.

Choose the format

  • SRT - universal, index + HH:MM:SS,mmm --> HH:MM:SS,mmm + text. Use for upload to most platforms and editors.
  • WebVTT (.vtt) - WEBVTT header, HH:MM:SS.mmm timestamps, supports positioning/styling cues. Use for HTML5 <track> and the web player.

Produce valid, parseable output - correct timestamp punctuation (SRT uses a comma before milliseconds, VTT a period), blank line between cues, no trailing whitespace.

Clean vs verbatim

For tutorials, caption clean: drop filler ("um", "uh", false starts), keep meaning verbatim. Preserve technical terms, code, and UI labels exactly. Only go strict-verbatim if the user explicitly needs it (legal, research, exact-quote contexts).

QA the sync

Spot-check against the actual video at the start, a middle point, and the end - drift accumulates. Confirm cues land on their lines and clear before the next begins. Fix any caption that lingers over the wrong shot.

Hand off

The finished .srt / .vtt feeds social-video-formatter for burn-in and platform styling, or the web player's <track> element directly. Leave color, font, position, and animation to those steps - this skill outputs accurate, well-timed text and nothing more.

Don't

  • Don't style, color, position, or burn captions into the frame - that is social-video-formatter.
  • Don't animate words word-by-word - that is kinetic-typography.
  • Don't ship uncorrected ASR; the mistakes cluster exactly on the terms that matter.
  • Don't pack 3 lines or overrun the reading-speed budget to avoid splitting a cue - split it.

Individual skills in this repo

This repo contains 17 individual skills — each has its own dedicated page.

SkillMedev/skills

Designs REST API surfaces - resource naming, HTTP method and status-code semantics, error shapes, pagination, and filtering - and delivers an endpoint spec a consumer can build against without asking questions. Use when someone asks "how should I name this endpoint", "what status code should this return", "should this be PUT or PATCH", "how do I paginate this list", or is reviewing an API before it ships to external consumers. Do NOT use for planning breaking-change rollouts and deprecation windows - use api-versioning-strategist instead; for GraphQL type and resolver design - use graphql-schema instead; for generating client SDKs from an existing spec - use api-client-generator instead; for designing inbound webhook endpoints - use webhook-receiver-hardener instead.

SkillMedev/skills

Turns data and charts into a decision-driving narrative structured as headline finding, trend, implication, and recommended action - with finding-led chart titles, context for every number, annotation guidance, and honest flags on any conclusion the data cannot support. Use when someone says "turn these numbers into a story", "what's the takeaway from this data", "help me present these results to leadership", or has charts but no narrative. Do NOT use for compressing a long document into a one-pager - use executive-summary instead - or for running the analysis that produces the findings - use eda-playbook instead.

SkillMedev/skills

Use when a task needs live or historical money data - "convert USD to EUR", "current/past exchange rate", "FX rate on this date / over this range", or "current price of Bitcoin/Ethereum, market cap, 24h change". Frankfurter (ECB reference rates, no key) is the FX default; CoinGecko's free keyless tier covers crypto. Do NOT use for stock quotes or equities - no keyless stock API survives verification, say so instead of guessing; do NOT use for country economic indicators like GDP or inflation series - use government-open-data instead; if the request is a vague "I need live data", route through public-data-api-picker.

SkillMedev/skills

Builds a driver-based FP&A operating model linking business inputs to P&L, balance sheet, and cash flow outputs. Use when building an annual plan, preparing investor materials, running scenario analysis, or stress-testing the business.

SkillMedev/skills

Use when a task needs live geographic lookups - "geocode this address", "what's at these coordinates" (reverse geocoding), "lat/lon for this city", "which country/state is this ZIP or postal code in", or "country facts: capital, currency, population, flag". Nominatim (OpenStreetMap) is the geocoding default; Zippopotam for postal codes; APICountries for country facts. All keyless. Do NOT use for weather at a location - use weather-climate instead; do NOT use for country-level statistics over time (GDP, population trends) - use government-open-data instead; if the request is a vague "I need live data", route through public-data-api-picker.

SkillMedev/skills

Runs the full Getting Things Done loop - capture, clarify, organize, reflect, engage - building a trusted system of context lists, a projects list with defined next actions, and a weekly review habit. Use when someone says "I'm overwhelmed and things are slipping through the cracks", "set up GTD for me", "help me do a brain dump and organize it", or "my to-do list is a mess". Do NOT use for just running the weekly review ritual itself - use weekly-review instead - or for clearing an email backlog - use inbox-zero.

SkillMedev/skills

Processes any email backlog to zero using the 4Ds - Delete, Delegate, Defer, Do - with a mass-archive strategy for the obvious, a touch-each-email-once discipline, and a keep-it-clear system of batched processing windows, ruthless unsubscribing, filters, and a minimal folder setup. Use when someone says "I have 5,000 unread emails", "help me get to inbox zero", "email is eating my whole day", or treats their inbox as a to-do list. Do NOT use for drafting the reply emails themselves or prioritization rules for an ongoing support queue - use email-triage instead - or for protecting focus time around the email windows - use deep-work-planner instead.

SkillMedev/skills

Runs structured coaching sessions using values clarification and the GROW model, ending every session with one committed action, a deadline, and an if-then plan for the likely obstacle. Use when someone says "I feel stuck in my life", "help me figure out what I want", "hold me accountable to my goals", or "coach me through this decision". Do NOT use for building a stress toolkit - use stress-management instead - or a journaling practice - use journal-framework; for a standing goal-tracking system, use goals-accountability. Coaching, not therapy: signs of clinical distress route to a licensed professional.

SkillMedev/skills

Adapts English content for a target language and region - swapping idioms, setting the right formality register, replacing cultural references, converting dates, units, and currency - and delivers the localized text with a notes column explaining each adaptation and flagging what needs a native reviewer. Use when someone asks "localize this copy for the Japanese market", "adapt our marketing copy for the Mexican market", "make this translation sound native instead of literal", or a literal translation reads foreign, stiff, or risky. Do NOT use for auditing in-product UI text quality - use ux-writing-audit instead.

SkillMedev/skills

Turns a natural-language video brief into a complete, ready-to-preview Remotion composition - extracts duration, scenes, brand colors, aspect ratio, and real copy; plans the frame budget; and writes data-driven React/TypeScript using useCurrentFrame, interpolate, spring, AbsoluteFill, and Sequence, registered in Root.tsx. Use when someone says "build a 20-second product demo video", "animate a feature walkthrough", "write the Remotion code for this marketing clip", or wants a scene-by-scene composition they can scrub in Studio. Do NOT use for rendering, iterating, or batching the finished MP4 - use remotion-render instead - or for installing and scaffolding the project - use remotion-setup instead.

SkillMedev/skills

Scaffolds a new Remotion video project wired for Claude Code Agent Skills - Node check, create-video scaffold, skills install, folder conventions, Google Fonts, and a smoke-test render. Use whenever someone wants to start making videos with Remotion and Claude, even just "make product videos with Claude.

SkillMedev/skills

Classifies incident severity (SEV1-4) using impact, scope, and urgency signals and decides who to page. Use when an alert fires or a report comes in and a severity call must be made quickly.

SkillMedev/skills

Use the Skill Me catalog from inside any conversation - discover, install, and manage Claude skills through the Skill Me MCP, and load installed skills automatically each session.

SkillMedev/skills

Write one platform-native caption from a topic and brand voice, with a front-loaded hook, native length, and a single earned CTA. Use when the user gives a topic (and optionally brand voice or a described image/video) and asks for an Instagram, LinkedIn, X/Twitter, or TikTok caption. Do NOT use when the user wants a dedicated standalone LinkedIn post - use linkedin-post-writer instead; or a multi-post X thread - use tweet-thread-builder instead; or to adapt one existing piece of content into versions for several channels - use cross-platform-reformatter instead.

SkillMedev/skills

Reframe and export a finished video for social platforms - 9:16 ↔ 16:9 cropping, burned-in captions, a platform-native first frame, loop design, and per-platform specs (aspect, length, safe areas) for TikTok, Reels, Shorts, X, and LinkedIn. Use when someone says "make this vertical", "crop to 9:16", "format for TikTok/Reels/Shorts", "add captions/subtitles", "why is my video cut off on mobile", "resize for Instagram", or "export for social". Do NOT use when the question is about the storyboard or shot order before any video exists (use video-storyboard); the animation craft of how elements move (use motion-design-principles); animated word-by-word caption styling (use kinetic-typography); palette, contrast, or lighting (use motion-color-and-light); directing the product-demo content and screen choreography (use product-demo-director); or scoring, beat-syncing, or audio (use sound-and-music-sync) - this is the craft layer that decides WHAT to format for the feed, then hands a concrete spec to remotion-compo...

SkillMedev/skills

Writes and tunes PySpark jobs - join strategy and broadcast size limits, shuffle-partition sizing, skew diagnosis and salting, UDF avoidance, caching, and output file layout - with concrete size and skew thresholds. Use when someone asks "why is my Spark job slow", "should I broadcast this join", "one task takes forever while the rest finish", "my job OOMs during a join", or is writing a new PySpark ETL job. Do NOT use for Kafka topic, consumer-group, or streaming-pipeline design - use kafka-pipelines instead; do NOT use for single-machine dataframe work that fits in memory - use pandas-expert instead.

SkillMedev/skills

Builds clean, performant, accessible SwiftUI views with correct state ownership, scoped invalidation, and smooth list scrolling, and reviews existing SwiftUI code against a concrete frame-time and re-render budget. Use when someone asks "why does my SwiftUI list stutter", "should this be @State or @Observable", "my whole screen re-renders when one row changes", "how do I animate this transition", or wants a SwiftUI view built or refactored. Do NOT use for cross-platform React Native apps - use react-native-pro instead; do NOT use for Flutter widget trees - use flutter-widget-architect instead; do NOT use for Android Compose UIs - use jetpack-compose-builder instead.

相关技能