YouTube Bilingual Hardburn
Turns a foreign-language video into a Chinese–English hardsub video. The deterministic steps are scripts; the translation step is done by the agent (LLM) and is where all the risk lives.
IRON LAW: ONE Chinese line per English line. NEVER split one English line into two Chinese lines (or merge two into one). After every translation batch, run verify_align.py. A single off-by-one shifts every later subtitle out of sync — and you will not notice until the burn is done.
When the source has punctuation vs. when it doesn't
- YouTube auto-captions (and most ASR) have almost no punctuation → the "split on
.!?" approach fails (collapses everything into one cue). This skill solves that by merging phrases up to a character cap. This is the default path. - If the source already has clean punctuated subtitles, the same pipeline still works.
Prerequisites
yt-dlp(download),ffmpeg(hardburn),python3(scripts). No Python packages needed.- A CJK-capable font installed. macOS:
PingFang SC(default). Linux: pass--font "Noto Sans CJK SC". Windows:--font "Microsoft YaHei". SCRIPTS=the absolute path to this skill'sscripts/directory. Set it once and reuse.
Workflow
Bilingual Hardburn Progress:
- [ ] Step 1: Get the video + captions ⚠️ REQUIRED
- [ ] Step 2: Extract + re-segment captions into readable lines
- [ ] Step 3: Translate every line (1:1) — the careful part ⚠️ REQUIRED
- [ ] Step 4: verify_align.py == OK ⛔ BLOCKING (cannot proceed until it passes)
- [ ] Step 5: Build bilingual .ass + confirm style on ONE test frame ⚠️ REQUIRED
- [ ] Step 6: Hardburn full video (long encode) + spot-check frames
Step 1: Get the video + captions ⚠️ REQUIRED
yt-dlp -o "%(upload_date)s_%(title)s.%(ext)s" \
-f "bestvideo[height<=1080][ext=mp4]+bestaudio[ext=m4a]/best[height<=1080]/best" \
--merge-output-format mp4 --write-subs --write-auto-subs --sub-langs "en" \
--no-mtime "<URL>"
- If
--write-auto-subsproduces no caption file (some videos have none), you must transcribe the audio first (e.g. whisper) into a.vtt/.srt, then continue. - Note the two output files:
VIDEO.mp4andVIDEO.en.vtt.
Step 2: Extract + re-segment ⚠️
python3 "$SCRIPTS/extract_phrases.py" "VIDEO.en.vtt" > phrases.json
python3 "$SCRIPTS/merge_lines.py" phrases.json lines.json # --max-chars 95 --gap 12
This produces lines.json ([{start,end,en}]) and lines_en.txt (a numbered idx<TAB>English file). Translate lines_en.txt. Tune later if needed:
- subtitles feel too long on screen → lower
--max-chars(e.g. 70) - you want breaks at natural pauses → lower
--gap(e.g. 2.5) — yields more, shorter lines
Step 3: Translate every line (1:1) ⚠️ REQUIRED — this is the careful part
Read lines_en.txt (it is numbered 0..N-1). Produce a plain-text ZH file: exactly one Chinese line for each English row, same order. Write newline-separated text (NOT JSON — Chinese quotes break JSON escaping).
Rules that prevent the Iron-Law violation:
- One row in → one line out. Translate the English row as its own unit. If a row breaks mid-sentence (it often will — that's normal for ASR), let the Chinese break at the same place. Do not "fix" it by flowing across rows.
- Work in batches of ~120–150 rows. After EACH batch write its file and immediately run
verify_align.pyon the cumulative result. - Keep proper nouns in their correct form (ChatGPT, OpenAI, Claude, Gemini, Grok, DeepSeek…). ASR mangles them ("chpd"→ChatGPT, "grock"→Grok, "ha cou"→haiku) — fix to the real name.
- Natural spoken Chinese, not stiff translation. Match the speaker's register.
Concatenate all batch files into one zh.txt (one ZH line per EN row, in order).
Step 4: verify_align.py ⛔ BLOCKING
python3 "$SCRIPTS/verify_align.py" lines.json zh.txt
- Prints
OK→ proceed. - Prints
MISMATCH+ a side-by-side table → find the row where EN and ZH meanings stop corresponding. That row is a split (one EN row became two ZH lines) or a merge. Fix it (merge the two ZH lines back into one, or split one into two), then re-run. Do not proceed until OK.
Step 5: Build .ass + confirm on ONE test frame ⚠️ REQUIRED
python3 "$SCRIPTS/build_bilingual_ass.py" lines.json zh.txt styled.ass
# Chinese on top (white 44px) + English below (grey 30px), bottom-centre.
# Linux/Windows: add --font "Noto Sans CJK SC" / --font "Microsoft YaHei"
Render ONE frame at a timestamp known to have a subtitle, and LOOK at it before the long encode:
ffmpeg -ss 60 -i "VIDEO.mp4" -copyts -vf "ass=styled.ass" -frames:v 1 -y test_frame.png
Confirm with the user: Chinese renders (no tofu boxes □□□), two-tier layout looks right, size/position acceptable. Adjust --zh-size/--en-size/--margin/--font and rebuild if needed. Do not start the full encode until the test frame looks correct.
Step 6: Hardburn full video + spot-check
ffmpeg -y -i "VIDEO.mp4" -vf "ass=styled.ass" \
-c:v libx264 -preset fast -crf 22 -c:a copy -sn "VIDEO_中英双语.mp4"
A 2-hour 1080p video takes ~15–40 min — run it in the background. When done, grab frames at the start, middle, and near-end to confirm subtitles stay in sync the whole way:
for t in 600 3900 7200; do ffmpeg -ss $t -i "VIDEO_中英双语.mp4" -frames:v 1 -y "chk_$t.png"; done
If a frame looks misaligned, re-confirm the same timestamp's cue in lines.json+zh.txt are the same source row before assuming a bug (low-res frames are easy to misread).
Anti-Patterns
- Splitting/merging lines during translation — the #1 cause of full-video desync. One EN row = one ZH line, always.
- Writing translations as JSON — Chinese curly quotes “ ” and 「」 silently break JSON parsing. Use newline-separated plain text.
- Trusting caption END times for ordering — rolling-caption ENDs overlap;
merge_lines.pysequences by START only. Don't reorder by END. - Skipping the test frame — burning 2 hours then discovering tofu boxes or a wrong font wastes the whole encode.
- Re-encoding audio — use
-c:a copy. Only the video track changes. - Using the old
build_bilingual.pyfrom the video-download skill on auto-captions — it splits on punctuation and collapses the whole transcript into one cue. This skill replaces it.
Pre-Delivery Checklist
-
verify_align.pyprintedOK(EN count == ZH count) before building the .ass - Test frame inspected: Chinese renders, no □ boxes, layout/size correct
- Spot-checked start/middle/end frames of the final mp4 — subtitles in sync
- Output named clearly (e.g.
<title>_中英双语.mp4); original video untouched - Proper nouns corrected (no "chpd"/"grock"/"ha cou" left in the Chinese)
Cross-platform (OpenAI Codex, etc.)
All scripts are pure python3 + ffmpeg/yt-dlp — no Claude-specific dependency. To run this workflow under Codex (or any shell-capable agent), load references/codex-usage.md.