hebrew-captions-asr
Hebrew transcription and burned-in captions: choose an ASR route by profile, get word timings, proofread by hand, build RTL-correct caption cards that animate in AND out, and prove them with caption_qa plus a coverage statement.
Rules that outrank the rest
- Force
language="he". HyperFrames' built-intranscribedefaults tosmall.en(poor Hebrew) andinitauto-transcribes with it: use--skip-transcribeand the local ivrit-ai route. - Never run an LLM over a whole transcript. Fix specific words by hand with a per-project spelling dictionary; decide ambiguous words by majority over several passes, then stop flip-flopping. Delivery bar: zero spelling errors (names, brands, quotes); a Hebrew reader accepts, a model critic only nominates.
- No cloud ASR/TTS call without a dated estimate and approval (
paid-generation-gate). The default is local and free; "local = $0" excludes setup, build time and correction minutes. - No
dir="rtl"on the HyperFrames root.lang="he"on<html>;direction: rtlonly on text elements. This is an engine/version-specific workaround (a 0.8.98 test did not reproduce the black render, E12): keep the rule, re-test per version, never teach "root RTL is invalid HTML". - Fonts from files (
hf/fonts/+@font-face); the HyperFrames browser does not find installed fonts by name and falls back silently. - Measured numbers are machine-bound. Quote speed as audio-seconds per wall-second with the timer scope named; tell students to re-measure on their own hardware.
- Law vs house preset (decision default Q5). Laws: exit animation as well as entrance, zero spelling errors, no root RTL, the look-alike test before a face is locked, back-transcription after a re-cut. House preset v1 (overridable): Rubik, 1-3 word cards, rail bottom <= y 1450, the dwell numbers below.
- Whisper tools break on non-ASCII file paths (
whisper-clion a Hebrew path): use an ASCII temp dir. Long ASR runs underrender_locklike any heavy job.
Inputs -> outputs
In: audio or video, DESIGN.md/brand type kit, the final cut if captions follow an edit. Out (project-relative): data/words.json ([{"text","start","end"}] seconds), transcript.srt, transcript.txt, data/captions.json (cards), the caption component in the composition, _work/qa/captions_report.json with a coverage statement.
ASR route by profile (measured on the owner's machine only)
| Profile | Route | Evidence and limit |
|---|---|---|
| any CPU (portable baseline) | faster-whisper + ivrit turbo CT2, int8, 8 physical-core threads | 731.4 s for 614.22 s of audio = 0.84 audio-s per wall-s (E08) |
| AMD/Windows with a Vulkan driver | whisper.cpp built with -DGGML_VULKAN=ON + the ivrit GGML model | 105.6 s for the same audio = 5.82 audio-s per wall-s; the same build on CPU 1,033.0 s (9.78x slower); RX 7600S needed GGML_VK_DISABLE_COOPMAT=1 (local-only: set it only if the run crashes, exit 127) |
| NVIDIA | faster-whisper CUDA 12 / cuDNN 9 or whisper.cpp CUDA | unmeasured |
| Apple Silicon | whisper.cpp Metal | unmeasured |
Corpus: 68 FLEURS he_il clips, 614.22 s of read speech, one pass, idle machine, whole-wrapper seconds. WER on that corpus (NFC, lowercase, punctuation -> space): whisper.cpp CPU 18.09 %, Vulkan 18.26 %, CT2 plain 17.74 %, CT2 + VAD 19.38 % (4-19 word edits of 1,161: close counts, not equivalence). Real dialogue, noisy phone audio and manual word-boundary accuracy are not measured. Full tables, build recipe and limits: references/asr-routes.md; dated facts: references/volatile-facts.md. |
Procedure
- Route. Probe the machine (
doctor), pick the route above, writeasr_route.json(model path/revision/hash, device, precision, threads, backend log line). Weights are pinned, never "latest". - VAD decision (optional per clip, decision default Q8). Without VAD Whisper returned non-empty Hebrew text for 3 s of digital silence and for 3 s of low noise (E02); the explicit-span VAD run in E08 was slower and less accurate than plain CT2 (225 vs 206 word edits). So: A/B on a 60 s sample of THIS audio, keep the lower WER, and always run the silence control when the audio has long non-speech.
- Transcribe with word timestamps; consume every segment inside the timed region. Save
words.json,.srt(--words-per-cue 3word-pop, 6 sentence mode),.txt. Sung or music-bed audio: a no-VAD "lyrics pass" with the confidence filters inreferences/asr-routes.md. - Proofread the whole text: names, brands, numbers (digits, not number words), the spelling dictionary, ambiguous words by majority. Flag fillers ("אה/אממ" are usually absent from Whisper output: find voiced gaps) and never auto-delete discourse words.
- After any re-cut ASR the FULL assembled VO and diff against the intended text (
join_diff); isolated join snippets said "clean" while the full file heard residues. Back-transcribe TTS output the same way. - Build cards (
references/caption-timing.md): 1-3 words, keywords in the keyword colour, bidi isolates, entrance + exit, 2-frame-early swaps, lead the voice 0.08-0.1 s, clamp starts at 0; font test perreferences/hebrew-typography.md; layout and bidi perreferences/rtl-and-bidi.md. - QA.
scripts/caption_lint.pyoncaptions.json,caption_qa --band <top>:1450on the render,hf_preflight, snapshots at every keyword frame (--describe false),frame_qa; write the coverage statement (frames decoded of expected, band, fps assumed, what was NOT checked).
Gates
States: pass | fail | blocked | n/a with a reason; a timeout, an empty card list or a missing file is blocked.
| Gate | Predicate | Evidence | If false | Owner | Recheck when |
|---|---|---|---|---|---|
| G1 ASR route | route chosen from a capability probe; model + weights pinned; language=he forced | asr_route.json (path, revision, device, precision, threads, timer scope) | fall back to CPU CT2 int8 | this skill | new machine, driver, model |
| G2 VAD | VAD on/off decided by an A/B on a 60 s real sample; the silence control run when the audio has long non-speech | both WERs (scripts/wer.py) + control output recorded | switch VAD off if WER is worse; add filters if hallucinations appear | this skill | new audio type |
| G3 transcript | every name/brand/number proofread; spelling dict applied; 0 spelling errors approved by a Hebrew reader | approved transcript + dict; wer.py against the corrected text on a sample | fix words by hand; never an LLM over the whole text | this skill + human | any text change |
| G4 assembled cut | ASR of the FULL assembled VO matches the intended words; no extra token; first/last 2 words of each sentence present | join_diff JSON | fix the join, re-cut, re-ASR | edit-talking-head G2 | every re-cut |
| G5 RTL / bidi | no root dir=rtl; lang="he"; Latin, digits, currency, URLs isolated; no stray bidi controls; the final frame viewed | scripts/caption_lint.py bidi block + hf_preflight + viewed frame | isolate the span; strip controls | this skill | any caption text change |
| G6 look-alikes / font | each keyword rendered at final size and at 360x640; ו/ז, ד/ר, ה/ח cannot be read as another word; font loaded from a file | specimen record in hf/QA.md (font file, hash, size, frame) + snapshot showing no fallback | swap the face for that word (Karantina "לבזבז" read as "לבובו") | this skill | any font, size or keyword change |
| G7 timing + animation | word >= 0.25 s, card >= 0.9 s, last word >= 0.25 s; exit animation present; swap overlap 1-2 frames; no negative start | caption_lint report + caption_qa + frame check | re-time; add the exit | this skill | any re-cut or retime |
| G8 safe zone + contrast + coverage | caption bottom <= y 1450 (house preset); brand-colour keyword >= 4.5:1 on the strip; caption_qa states its coverage | caption_qa report with coverage statement + contrast numbers | move, change colour, or add a backing | this skill | any layout change; new ratio |
House preset v1 numbers (each overridable; sources dated in the references)
Rubik: Black (900) keywords, Regular/600 small words; 54-66 px at 1080x1920 (E03 provisional default; Alef 700 and Noto Sans Hebrew 600 score equally; Karantina scored 1.8 vs 4.2 in a single-model static review, no human panel, no phone, no animation). Word-pop 1-3 words, 0.35-0.7 s per card; entrance mirrored by exit (blur-out-up, about 4 frames, rise 10 px, blur 6 px); caption rail y 900-1240 on A-roll, bottom <= 1450; key text y <= 1248. The 9:16 platform table is dated 2026-09 and unverified as law: run the overlay test on a phone.
References (load when)
references/asr-routes.md- choosing/building an ASR route, WER protocol, VAD, lyrics pass, cloud options (none run).references/hebrew-typography.md- choosing a font, the look-alike test, E03 table, licences.references/rtl-and-bidi.md- mixed Hebrew/English/number lines, HyperFrames and After Effects RTL traps.references/caption-timing.md- modes, exact timing recipe, safe zones, contrast script.references/volatile-facts.md- before quoting speeds, WER, pins, licences, safe-zone numbers.- Scripts:
scripts/caption_lint.py(card timing/bidi/rail lint),scripts/wer.py(WER/CER with a declared normalisation); both stdlib,--self-check; script paths are relative to this skill's folder. - Repo-level dated modules (owned elsewhere):
agent-content/references/asr-routes.md,agent-content/references/hebrew-rtl-captions.md,agent-content/references/platform-specs.md,agent-content/techniques/caption-collision.md; load when you need the dated tables or the box-collision method. - Siblings:
edit-talking-head,render-qa-deliver,video-analysis(analysis reports, not captions).
Evidence status
ASR from experiments E02/E08 (owner machine); fonts from E03; timing and exit rules from the owner's projects. No human timing ground truth, no cloud ASR run, no model-licence resolution for ONNX conversions. Specified; deterministic checks only; model eval not run (Q4).