Communitygithub.com

Malik-rgb43/editing-workflow

>- Transcribe Hebrew speech and build burned-in Hebrew captions: ASR route per hardware, word timestamps, proofreading, RTL and mixed English/number lines, caption font, entrance and exit animation, safe zone, caption QA. Triggers: תמלול, תמלל, כתוביות, כתוביות בעברית, מילים בולטות, קריוקי, המילה יצאה לא נכון, Hebrew captions, transcribe Hebrew, word timestamps, ו/ז look-alike. NOT for translation-only, non-Hebrew ASR, analysis reports (video-analysis), or caption-free motion launches.

What is editing-workflow?

editing-workflow is a Claude Code agent skill that >- Transcribe Hebrew speech and build burned-in Hebrew captions: ASR route per hardware, word timestamps, proofreading, RTL and mixed English/number lines, caption font, entrance and exit animation, safe zone, caption QA. Triggers: תמלול, תמלל, כתוביות, כתוביות בעברית, מילים בולטות, קריוקי, המילה יצאה לא נכון, Hebrew captions, transcribe Hebrew, word timestamps, ו/ז look-alike. NOT for translation-only, non-Hebrew ASR, analysis reports (video-analysis), or caption-free motion launches.

Works with~Claude Code~Codex CLI~Cursor
npx skills add https://github.com/Malik-rgb43/editing-workflow/tree/HEAD/.agents/skills/hebrew-captions-asr

Ask in your favorite AI

Open a new chat with this agent skill pre-loaded.

Documentation

hebrew-captions-asr

Hebrew transcription and burned-in captions: choose an ASR route by profile, get word timings, proofread by hand, build RTL-correct caption cards that animate in AND out, and prove them with caption_qa plus a coverage statement.

Rules that outrank the rest

  1. Force language="he". HyperFrames' built-in transcribe defaults to small.en (poor Hebrew) and init auto-transcribes with it: use --skip-transcribe and the local ivrit-ai route.
  2. Never run an LLM over a whole transcript. Fix specific words by hand with a per-project spelling dictionary; decide ambiguous words by majority over several passes, then stop flip-flopping. Delivery bar: zero spelling errors (names, brands, quotes); a Hebrew reader accepts, a model critic only nominates.
  3. No cloud ASR/TTS call without a dated estimate and approval (paid-generation-gate). The default is local and free; "local = $0" excludes setup, build time and correction minutes.
  4. No dir="rtl" on the HyperFrames root. lang="he" on <html>; direction: rtl only on text elements. This is an engine/version-specific workaround (a 0.8.98 test did not reproduce the black render, E12): keep the rule, re-test per version, never teach "root RTL is invalid HTML".
  5. Fonts from files (hf/fonts/ + @font-face); the HyperFrames browser does not find installed fonts by name and falls back silently.
  6. Measured numbers are machine-bound. Quote speed as audio-seconds per wall-second with the timer scope named; tell students to re-measure on their own hardware.
  7. Law vs house preset (decision default Q5). Laws: exit animation as well as entrance, zero spelling errors, no root RTL, the look-alike test before a face is locked, back-transcription after a re-cut. House preset v1 (overridable): Rubik, 1-3 word cards, rail bottom <= y 1450, the dwell numbers below.
  8. Whisper tools break on non-ASCII file paths (whisper-cli on a Hebrew path): use an ASCII temp dir. Long ASR runs under render_lock like any heavy job.

Inputs -> outputs

In: audio or video, DESIGN.md/brand type kit, the final cut if captions follow an edit. Out (project-relative): data/words.json ([{"text","start","end"}] seconds), transcript.srt, transcript.txt, data/captions.json (cards), the caption component in the composition, _work/qa/captions_report.json with a coverage statement.

ASR route by profile (measured on the owner's machine only)

ProfileRouteEvidence and limit
any CPU (portable baseline)faster-whisper + ivrit turbo CT2, int8, 8 physical-core threads731.4 s for 614.22 s of audio = 0.84 audio-s per wall-s (E08)
AMD/Windows with a Vulkan driverwhisper.cpp built with -DGGML_VULKAN=ON + the ivrit GGML model105.6 s for the same audio = 5.82 audio-s per wall-s; the same build on CPU 1,033.0 s (9.78x slower); RX 7600S needed GGML_VK_DISABLE_COOPMAT=1 (local-only: set it only if the run crashes, exit 127)
NVIDIAfaster-whisper CUDA 12 / cuDNN 9 or whisper.cpp CUDAunmeasured
Apple Siliconwhisper.cpp Metalunmeasured
Corpus: 68 FLEURS he_il clips, 614.22 s of read speech, one pass, idle machine, whole-wrapper seconds. WER on that corpus (NFC, lowercase, punctuation -> space): whisper.cpp CPU 18.09 %, Vulkan 18.26 %, CT2 plain 17.74 %, CT2 + VAD 19.38 % (4-19 word edits of 1,161: close counts, not equivalence). Real dialogue, noisy phone audio and manual word-boundary accuracy are not measured. Full tables, build recipe and limits: references/asr-routes.md; dated facts: references/volatile-facts.md.

Procedure

  1. Route. Probe the machine (doctor), pick the route above, write asr_route.json (model path/revision/hash, device, precision, threads, backend log line). Weights are pinned, never "latest".
  2. VAD decision (optional per clip, decision default Q8). Without VAD Whisper returned non-empty Hebrew text for 3 s of digital silence and for 3 s of low noise (E02); the explicit-span VAD run in E08 was slower and less accurate than plain CT2 (225 vs 206 word edits). So: A/B on a 60 s sample of THIS audio, keep the lower WER, and always run the silence control when the audio has long non-speech.
  3. Transcribe with word timestamps; consume every segment inside the timed region. Save words.json, .srt (--words-per-cue 3 word-pop, 6 sentence mode), .txt. Sung or music-bed audio: a no-VAD "lyrics pass" with the confidence filters in references/asr-routes.md.
  4. Proofread the whole text: names, brands, numbers (digits, not number words), the spelling dictionary, ambiguous words by majority. Flag fillers ("אה/אממ" are usually absent from Whisper output: find voiced gaps) and never auto-delete discourse words.
  5. After any re-cut ASR the FULL assembled VO and diff against the intended text (join_diff); isolated join snippets said "clean" while the full file heard residues. Back-transcribe TTS output the same way.
  6. Build cards (references/caption-timing.md): 1-3 words, keywords in the keyword colour, bidi isolates, entrance + exit, 2-frame-early swaps, lead the voice 0.08-0.1 s, clamp starts at 0; font test per references/hebrew-typography.md; layout and bidi per references/rtl-and-bidi.md.
  7. QA. scripts/caption_lint.py on captions.json, caption_qa --band <top>:1450 on the render, hf_preflight, snapshots at every keyword frame (--describe false), frame_qa; write the coverage statement (frames decoded of expected, band, fps assumed, what was NOT checked).

Gates

States: pass | fail | blocked | n/a with a reason; a timeout, an empty card list or a missing file is blocked.

GatePredicateEvidenceIf falseOwnerRecheck when
G1 ASR routeroute chosen from a capability probe; model + weights pinned; language=he forcedasr_route.json (path, revision, device, precision, threads, timer scope)fall back to CPU CT2 int8this skillnew machine, driver, model
G2 VADVAD on/off decided by an A/B on a 60 s real sample; the silence control run when the audio has long non-speechboth WERs (scripts/wer.py) + control output recordedswitch VAD off if WER is worse; add filters if hallucinations appearthis skillnew audio type
G3 transcriptevery name/brand/number proofread; spelling dict applied; 0 spelling errors approved by a Hebrew readerapproved transcript + dict; wer.py against the corrected text on a samplefix words by hand; never an LLM over the whole textthis skill + humanany text change
G4 assembled cutASR of the FULL assembled VO matches the intended words; no extra token; first/last 2 words of each sentence presentjoin_diff JSONfix the join, re-cut, re-ASRedit-talking-head G2every re-cut
G5 RTL / bidino root dir=rtl; lang="he"; Latin, digits, currency, URLs isolated; no stray bidi controls; the final frame viewedscripts/caption_lint.py bidi block + hf_preflight + viewed frameisolate the span; strip controlsthis skillany caption text change
G6 look-alikes / fonteach keyword rendered at final size and at 360x640; ו/ז, ד/ר, ה/ח cannot be read as another word; font loaded from a filespecimen record in hf/QA.md (font file, hash, size, frame) + snapshot showing no fallbackswap the face for that word (Karantina "לבזבז" read as "לבובו")this skillany font, size or keyword change
G7 timing + animationword >= 0.25 s, card >= 0.9 s, last word >= 0.25 s; exit animation present; swap overlap 1-2 frames; no negative startcaption_lint report + caption_qa + frame checkre-time; add the exitthis skillany re-cut or retime
G8 safe zone + contrast + coveragecaption bottom <= y 1450 (house preset); brand-colour keyword >= 4.5:1 on the strip; caption_qa states its coveragecaption_qa report with coverage statement + contrast numbersmove, change colour, or add a backingthis skillany layout change; new ratio

House preset v1 numbers (each overridable; sources dated in the references)

Rubik: Black (900) keywords, Regular/600 small words; 54-66 px at 1080x1920 (E03 provisional default; Alef 700 and Noto Sans Hebrew 600 score equally; Karantina scored 1.8 vs 4.2 in a single-model static review, no human panel, no phone, no animation). Word-pop 1-3 words, 0.35-0.7 s per card; entrance mirrored by exit (blur-out-up, about 4 frames, rise 10 px, blur 6 px); caption rail y 900-1240 on A-roll, bottom <= 1450; key text y <= 1248. The 9:16 platform table is dated 2026-09 and unverified as law: run the overlay test on a phone.

References (load when)

  • references/asr-routes.md - choosing/building an ASR route, WER protocol, VAD, lyrics pass, cloud options (none run).
  • references/hebrew-typography.md - choosing a font, the look-alike test, E03 table, licences.
  • references/rtl-and-bidi.md - mixed Hebrew/English/number lines, HyperFrames and After Effects RTL traps.
  • references/caption-timing.md - modes, exact timing recipe, safe zones, contrast script.
  • references/volatile-facts.md - before quoting speeds, WER, pins, licences, safe-zone numbers.
  • Scripts: scripts/caption_lint.py (card timing/bidi/rail lint), scripts/wer.py (WER/CER with a declared normalisation); both stdlib, --self-check; script paths are relative to this skill's folder.
  • Repo-level dated modules (owned elsewhere): agent-content/references/asr-routes.md, agent-content/references/hebrew-rtl-captions.md, agent-content/references/platform-specs.md, agent-content/techniques/caption-collision.md; load when you need the dated tables or the box-collision method.
  • Siblings: edit-talking-head, render-qa-deliver, video-analysis (analysis reports, not captions).

Evidence status

ASR from experiments E02/E08 (owner machine); fonts from E03; timing and exit rules from the owner's projects. No human timing ground truth, no cloud ASR run, no model-licence resolution for ONNX conversions. Specified; deterministic checks only; model eval not run (Q4).

Individual skills in this repo

This repo contains 5 individual skills — each has its own dedicated page.

Malik-rgb43/editing-workflow

>- Script, edit or review a paid-social ad, sponsored post or brand promo for a business, product, app or store (Meta, Instagram, TikTok, 15-60 s): exact offer and CTA, ranked hooks, hook variants, end card, blocking compliance table, licence checks. Triggers: ad, promo, sponsored, campaign; מודעה, פרסומת, קמפיין, ממומן, פרומו, מבצע, הנחה, הוק. NOT for testimonials (edit-testimonial), organic expert reels (edit-talking-head), AI-only films (edit-ai-generated), ratio exports (multi-video-variants).

Malik-rgb43/editing-workflow

>- Design, build and review a motion-graphics piece in HyperFrames/GSAP: product or app launch, feature announcement, kinetic typography, logo reveal, UI explainer, animated screenshots. Triggers: מושן גרפיקס, אנימציה, השקה, טיפוגרפיה קינטית, לוגו אנימציה, "like a Higgsfield / Linear launch". Not for a speaker-to-camera reel (edit-talking-head), AI-generated shots (edit-ai-generated), caption-only work (hebrew-captions-asr) or ad compliance (edit-ad-promo).

Malik-rgb43/editing-workflow

>- Deliver one approved video as several outputs (other aspect ratios, hook variants, platform versions, no-music/no-captions versions) or run several videos in parallel, with one shared mix, strict naming and a manifest that blocks stale derivatives. Triggers: export as 9:16 and 1:1, hook variants, variants batch; גרסאות, וריאציות הוק, כמה סרטונים יחד, 9:16 ו-16:9. NOT for a single deliverable (render-qa-deliver) or designing the master (the type skill).

Malik-rgb43/editing-workflow

>- Turn a video the user sends or links (file, or YouTube/TikTok/Instagram URL) or a folder of clips into measurements (cuts, pacing, per-frame data), keyframe sheets, a Hebrew transcript and sound analysis (SFX, BPM, key, beats, song ID). Triggers: watch, analyse, break down, transcribe this video/reel/ad; תנתח, תעבור על, תמלל, מה קורה בסרטון, קאטים, BPM. NOT for cutting or rendering (type skills), applying a style (reference-style-transfer), or captions for a final render (hebrew-captions-asr).

Malik-rgb43/editing-workflow

Turn a new video idea, brief or list of specifics into a checkable Concept Ledger and ask questions until every parameter is precise, before any concept, prompt or build. Use at the start of any new video, with or without footage, or when the user pastes a concept with specifics. Hebrew - סרטון חדש, בריף, יש לי רעיון, קונספט, תכין לי סרטון, בוא נתחיל פרויקט. NOT for notes on an existing draft (revision-round), analysing a reference (video-analysis) or paid generation (paid-generation-gate).

Related Skills