Auto Video Editor
Objective
Produce a real local edit while preserving originals. Codex performs semantic decisions; the workspace CLI performs deterministic media analysis, validation, FFmpeg rendering, and technical QA.
Workspace root: D:\workspace\codex-auto-video-lab.
Workflow
If the brief asks to learn from reference videos, perform 精剪, build a tier/list overlay, time product cutouts to speech, or use layered caption motion, load $tiktok-fine-cut-director before drafting the timeline. Its fine-cut specification becomes the semantic input to this deterministic render workflow.
-
Read the user's requirements and resolve every supplied local asset path. Make a safe, concise job id. Do not edit the source files in place.
-
Run the workspace doctor if this is the first edit in the task:
powershell -ExecutionPolicy Bypass -File .\scripts\doctor.ps1 -
Create the isolated job. Pass
--assetonce per source and include the user's exact brief:.\studio.ps1 create --job <job-id> --asset <path> --brief <requirements> -
Analyze the job. Add
--transcribewhen speech meaning, captions, filler removal, or quote selection affects the edit:.\studio.ps1 analyze --job <job-id> --transcribe -
Read
jobs\<job-id>\brief.txt,manifest.json,reports\analysis.json, and relevant transcript JSON. Use scene candidates and silence spans as evidence, not as automatic editorial truth. -
Generate the conservative draft:
.\studio.ps1 draft-plan --job <job-id> -
Edit
jobs\<job-id>\edit-plan.jsonto satisfy the brief. Read edit-plan.md before changing the plan. Keep anintent_summarythat explains selection, ordering, pacing, text, aspect ratio, and audio decisions. Never invent unavailable B-roll or factual claims. -
Validate, render, and run QA:
.\studio.ps1 validate --job <job-id>.\studio.ps1 render --job <job-id>.\studio.ps1 qa --job <job-id> -
Inspect all paths in
reports\qa.jsonundersample_framesandcritical_overlay_frameswith the local image viewer. If captions, crop, title safe areas, font glyphs, critical overlays, or black frames look wrong, revise the plan and rerender. -
Deliver the MP4, edit plan, analysis, and QA paths. State separately what was actually executed: source inspection, transcription, local rendering, full decode, sampled-frame visual inspection, or platform playback. Never call technical QA a semantic or publishing verification.
Decision Rules
- Make reasonable editorial assumptions when the brief leaves small details open; record them in
intent_summary. - Ask the user only when a missing choice would materially change the story, product claim, identity, or licensed asset usage.
- Prefer hard cuts for action, speech, proof, and ordinary montage. Use a non-cut transition only for a real time/place/chapter shift, soft bridge, or visibly matched directional movement; use text scale only on a supported hook, contrast, number, proof, payoff, or CTA. Use pace and shot selection before effects.
- Keep dialogue intelligible. Background music defaults near
-18 dBwith ducking enabled. - For vertical short-form output, default to 1080x1920 at 30 fps. Use
containinstead of destructive cropping when important content would leave frame. - Do not burn new captions over source footage that already contains readable captions unless the user asks for restyling.
- Treat scene detection, silence detection, and speech recognition as fallible evidence. Check transcript boundaries before semantic cutting.
Fast Baseline
For a request that explicitly needs no semantic deletion--only ordered concatenation, reframing, optional title, and captions--baseline-auto may run the whole mechanical pipeline. Label it as a baseline assembly, not custom editorial judgment:
.\studio.ps1 baseline-auto --job <job-id> [--transcribe] [--title <text>] [--burn-captions]
Current Boundary
Supported: local video/image ingestion, probing, scene candidates, silence spans, optional offline transcription, segment trims, 0.25x-4x speed, fill/contain reframing, exact source muting or source gain, controlled push/pull motion, evidence-reasoned dissolve/dip/directional-slide transitions, styled animated titles/captions/labels, one exact semantic color run inside a caption/title/label with optional bounded pulse or shake during that word's real speaking window, paired numbered-step labels, generated voiceover, background music with narration ducking, timed sound effects, commerce color finishing, H.264/AAC MP4, full-decode QA, black-frame scan, audio-level scan, and sampled frames.
Not supported directly in this FFmpeg renderer: asset knowledge bases, generative B-roll, layered Remotion motion graphics, object tracking, multicam sync, editable Premiere/DaVinci project export, or platform publishing. Route layered Remotion work through $openchatcut as specified by $tiktok-fine-cut-director.