Communitygithub.com

sherhyf8311-ship-it/bilibili-subtitles

Retrieve available subtitles or create a clearly labeled transcript for Bilibili videos, including login-required captions through a user-scanned QR flow. Use when a user asks to fetch, transcribe, translate, or analyze Bilibili subtitles.

¿Qué es bilibili-subtitles?

bilibili-subtitles is a Codex agent skill that retrieve available subtitles or create a clearly labeled transcript for Bilibili videos, including login-required captions through a user-scanned QR flow. Use when a user asks to fetch, transcribe, translate, or analyze Bilibili subtitles.

Compatible con~Claude Code✓Codex CLI~Cursor
npx skills add sherhyf8311-ship-it/bilibili-subtitles

Preguntar en tu IA favorita

Abre un nuevo chat con esta habilidad de agente ya precargada.

Documentación

Bilibili Subtitles

Retrieve subtitles from a Bilibili video the user identifies, save them in a reusable format, and distinguish platform subtitles from machine transcription and translation.

Workflow

  1. Resolve the URL to its BV id or episode id, then query public video metadata and page count. For a bangumi URL, resolve the entire season's episode list and match the user's requested episode number to its ep_id, BV id, and cid; the shared URL may point to a different episode. Confirm the selected title and duration before retrieval. Record the title, uploader, duration, publication time, and part/cid values. Do not infer video contents from the title alone.
  2. Check the page and player subtitle metadata, then use yt-dlp --list-subs --skip-download to distinguish caption tracks from danmaku. Prefer downloading available manual or automatic caption tracks and preserve their language and timestamps.
  3. If yt-dlp explicitly says subtitles require login, immediately start the official Bilibili QR login through scripts/bilibili_qr_login.py. Once the local QR image is ready, display it and tell the user to scan and confirm it in their Bilibili app. Do not ask whether they want a QR first. Never scan or approve on their behalf. If login is not required but no caption track exists, do not create a QR just to search again; proceed to audio transcription when appropriate.
  4. The QR script runs yt-dlp with a temporary cookie jar and removes it in finally cleanup when subtitle retrieval finishes, times out, or is cancelled. Never print cookie contents or save them in the project or skill directory. Delete the QR image immediately after scan confirmation. Do not close the user's browser to unlock its cookie database.
  5. Interpret access signals precisely. A generic page/player notice such as purchase to watch full video does not by itself prove every format is inaccessible or that the user lacks ordinary access. Check the actual subtitle/audio request and any explicit account, purchase, DRM, or region error. Use only formats the user's current session can ordinarily access; never manipulate endpoints to defeat a restriction. QR login authenticates the user's account but does not grant entitlements it lacks.
  6. When captions are absent and the user wants a transcript, test whether yt-dlp can retrieve the audio-only format. If it succeeds without an access-control error, use an audio-only format and do not download the video stream. Prefer keeping the compressed audio temporary and converting it to mono 16 kHz for recognition when supported; avoid expanding a long episode into a large full-rate WAV unless the recognizer requires it. If audio retrieval returns a clear access denial or purchase restriction, stop and report it.
  7. Use an available speech-to-text model such as faster-whisper. Auto-detect language, preserve timestamps, and label the output as machine-generated. Do not call ASR output an official subtitle. Mark uncertain names, numbers, and technical terms for human review. Report the detected language, model, confidence when available, cue count, and last recognized timestamp. Compare that timestamp with media duration; voice-activity filtering can omit trailing silence, but unexplained missing speech means the transcript may be incomplete.
  8. If a model download from its configured source times out or stalls, check a reachable trusted mirror and resume partial downloads where the downloader supports it. Keep the model cache under a task-specific temporary directory. If assembling ranged downloads manually, verify every range's byte count and the final model's published checksum before loading it; otherwise do not use the assembled file.
  9. Save source captions as .vtt or .srt; save machine output separately from any translated version. Include the source URL, selected episode, retrieval date, language, method, coverage, and limitations in a short metadata note. After validating outputs, remove temporary audio, model files/cache, chunk files, helper scripts, QR image, and cookie jar; verify they are gone and report any cleanup failure. Do not remove unrelated pre-existing files.
  10. For analysis, cite transcript timestamps and separate what the speaker says from external claims that have not been independently verified.

yt-dlp examples

List subtitle tracks without downloading the video:

yt-dlp --list-subs "<video-url>"

Download available original/automatic captions:

yt-dlp --skip-download --write-subs --write-auto-subs --sub-langs "all" --sub-format "best" -o "%(title)s.%(ext)s" "<video-url>"

If the user's own logged-in browser session is required, use its cookies only with the user's existing entitlement. Do not ask for passwords or export/share cookie files. If authenticated access is unavailable, stop and explain what access is needed.

When browser-cookie extraction is blocked, use the QR flow above instead of asking the user to close the browser or export cookies. The QR and one-time login state are secrets: display the QR only in the local Codex session, never upload it to a QR website, delete the QR image immediately after scan confirmation, and remove both QR and cookie files after timeout or cancellation.

Install yt-dlp and qrcode[pil] into a temporary package directory if unavailable, then run the script with yt-dlp options after --, for example:

$runtime = Join-Path $env:TEMP "bilibili-subtitle-runtime"
py -3.11 -m pip install --target $runtime yt-dlp "qrcode[pil]"
$env:PYTHONPATH = $runtime
py -3.11 scripts/bilibili_qr_login.py --timeout 180 -- --skip-download --write-subs --write-auto-subs --sub-langs "all" --sub-format "best" -o "%(title)s.%(ext)s" "<video-url>"

When the script prints QR_IMAGE=<path>, immediately open that local image in the Codex session and tell the user to scan it. The script waits for Bilibili's scan confirmation and then runs the requested yt-dlp command. Never approve a login prompt on the user's behalf.

For ASR on media the user can access, use ffmpeg to extract mono 16 kHz audio, then run the selected local or hosted recognizer. Keep the original media only as long as needed, and remove temporary copies after confirming the requested outputs.

For a Bilibili audio-only fallback, yt-dlp can extract the best available audio stream without saving video:

$audioTemplate = Join-Path $env:TEMP "<bvid>-audio.%(ext)s"
yt-dlp -f ba -x --audio-format wav -o $audioTemplate "<video-url>"

For long episodes, keep the extracted compressed audio or convert it to a temporary mono 16 kHz file with ffmpeg before ASR to reduce disk use. If using faster-whisper with downloaded weights, point HF_HOME or its cache option at a task-specific directory under $env:TEMP, then remove that directory and the temporary audio after the SRT/TXT files are verified. A complete local CPU transcription can take substantially longer than metadata or audio retrieval; provide progress updates, wait for completion, and do not report success until the output file is complete and its timestamps have been checked against the media duration.

Reporting

State whether the result is an official subtitle, platform auto-caption, or machine transcript. Report missing tracks, access restrictions, failed language detection, and any incomplete coverage. Never fabricate missing words or present a translation as the original transcript.

Skills relacionados