Trilingual SRT builder (TH / EN / ZH)
A Short with subtitles in three languages can reach viewers who would never sit through a voice they don't understand. But bad subs are worse than none: out-of-sync lines, sentences cut mid-word, or Chinese that reads like a machine translation make a clip look cheap. This skill produces three SRT files that share exactly the same timings, read naturally, and fit a phone screen.
The bundled script scripts/make_srt.py (Python standard library only) writes the files and checks them. Use it whenever you can run code; if you can't, follow the same rules by hand and output the SRT text in code blocks.
Step 1 — Get the timed source text
Work out which of these the user has, best first:
- A transcript with timestamps (SRT/VTT from the editing app, YouTube auto-captions, a speech-to-text tool) → use its timings. Merge or split cues so each one is a single idea of 1–3.5 seconds.
- The final script + the clip length → split the script into cues (one spoken phrase each) and let the script estimate timings with
--estimate <seconds>. Tell the user the timings are estimates and must be nudged in their editor to match the voice — estimated subs are the most common cause of "subs feel off". - Only an audio/video file → if you have a speech-to-text tool in this session, transcribe with word timestamps first. If not, say so plainly and ask for a transcript export from their editing app (most editors and YouTube Studio can export captions). Never invent what was said.
Keep the source language exactly as spoken — subtitles that differ from the voice confuse viewers. Only fix obvious transcription errors (misheard words, brand names) and tell the user which ones you fixed.
Step 2 — Translate line by line
Translate each cue separately so the timings stay identical across languages. Rules:
- Spoken, short, natural. Subtitles are read in a second or two. Prefer "Turn it off at night" over "It is recommended that the device be switched off during night-time hours."
- Meaning over words. Thai filler (นะครับ, เลย, อ่ะ) disappears in English/Chinese. Idioms get an equivalent, not a literal copy.
- Keep numbers, units, model names and app names identical in all three files, and the same format (e.g. 80%, 20W, iOS 26). Check that every number in the Thai line also appears in the EN and ZH lines.
- Chinese = Simplified (zh-Hans) unless the user asks for Traditional (zh-Hant, for Taiwan/Hong Kong viewers). Use full-width punctuation (,。?!). Don't put spaces between Chinese characters.
- Menu labels and on-screen UI text: use the label as it actually appears in that language's version of the app/phone if you know it; if you're not sure of the official label, translate descriptively and flag it so the user can check on a device set to that language.
- Product/brand names, place names and people's names: keep the official romanisation; don't translate brand names into Chinese unless the brand has an official Chinese name you are confident about.
- Don't add information, jokes or emojis that aren't in the voice.
Step 3 — Write the JSON and run the script
{
"langs": ["th", "en", "zh-Hans"],
"segments": [
{"start": 0.0, "end": 2.6, "th": "ชาร์จมือถือข้ามคืน แบตเสื่อมจริงไหม?", "en": "Does charging overnight ruin your battery?", "zh-Hans": "整夜充电会伤电池吗?"},
{"start": 2.7, "end": 6.1, "th": "คำตอบสั้น ๆ คือ ไม่ได้พังทันที แต่ความร้อนคือตัวการ", "en": "Short answer: not instantly, but heat is the real enemy.", "zh-Hans": "简单说:不会马上坏,但真正的敌人是热量。"}
]
}
python3 scripts/make_srt.py segments.json --out-dir subs --name clip01
# no timings yet:
python3 scripts/make_srt.py segments.json --out-dir subs --name clip01 --estimate 38.5
Output: subs/clip01.th.srt, subs/clip01.en.srt, subs/clip01.zh-Hans.srt plus a QC report.
Line breaks: the script wraps long cues into two lines. Thai has no spaces between words, so it only breaks at the spaces Thai writers put between phrases — it will never cut a Thai word in half. If a Thai cue is flagged as too long, add a space between phrases where a break is natural, or force the break with | (space-pipe-space), e.g. "เปิดโหมดชาร์จ | แบบปรับให้เหมาะสม".
Step 4 — Fix every QC message
The script checks, per language:
| Check | Default limit | Why it matters |
|---|---|---|
| Characters per line | TH 24 · EN 38 · ZH 15 | Vertical screen; longer lines run under the like/comment buttons |
| Lines per cue | 2 | Three lines cover the picture |
| Reading speed (chars/sec) | TH 17 · EN 20 · ZH 9 | Faster than this and viewers can't finish reading |
| Time on screen | ≥ 0.7 s | Flashes are unreadable |
| Overlaps / end before start / empty text | none allowed | Platforms reject or mis-show the file |
Fix by shortening the translation (first choice), splitting a long cue into two, or merging a flash cue with its neighbour. Re-run until there are 0 errors; a few warnings are OK if you tell the user what they are. Limits can be changed with --max-line and --max-cps if the user's style differs.
Thai character counting ignores vowel and tone marks that sit above or below a letter, because they don't take width on screen.
Step 5 — Hand over
Give the user:
- The three
.srtfiles (or their contents in code blocks if you can't create files). - A short table of any lines you were unsure about (UI labels, slang, names) so they can check.
- How to use them:
- YouTube Studio: Subtitles → choose the video → Add language → Upload file → "With timing". Upload one file per language. Also set the video's original language first.
- Burned-in captions (text drawn on the video): most short-form viewers watch muted, so burning the main language into the video and uploading the others as SRT is a common choice. Many editors import SRT and style it.
- TikTok / Reels: support for uploading SRT files differs by app version and region — if there's no upload option, burn the captions into the video in the editor instead. Check the platform's current help page.
Example
User: "ทำซับ 3 ภาษาให้หน่อย คลิป 9 วิ" + script:
ชาร์จมือถือข้ามคืน แบตเสื่อมจริงไหม? / คำตอบสั้น ๆ คือ ไม่ได้พังทันที แต่ความร้อนคือตัวการ / เปิดโหมดชาร์จแบบปรับให้เหมาะสม ในตั้งค่าแบตเตอรี่
What you do: three cues → translate → run with --estimate 9 → the QC warns that Thai cue 3 has a 26-character line (เปิดโหมดชาร์จแบบปรับให้เหมาะสม, limit 24). Forcing a break wouldn't help here (the second line would become too long), so shorten the subtitle instead: "เปิดโหมดถนอมแบต ในตั้งค่าแบตเตอรี่". A subtitle may be a little shorter than the spoken line as long as the meaning is the same; if the user's style is word-for-word subs, ask before shortening. Re-run: 0 errors, 1 warning (timings estimated). Deliver the three files and say: "timings are estimated from text length — slide them to match your voice in your editor; please check the English/Chinese name of the battery setting on a phone set to that language".
clip01.th.srt result:
1
00:00:00,000 --> 00:00:02,653
ชาร์จมือถือข้ามคืน
แบตเสื่อมจริงไหม?
2
00:00:02,653 --> 00:00:06,347
คำตอบสั้น ๆ คือ ไม่ได้พังทันที
แต่ความร้อนคือตัวการ
3
00:00:06,347 --> 00:00:09,000
เปิดโหมดถนอมแบต
ในตั้งค่าแบตเตอรี่