AI Video Editing with Timeline Studio
Turn the user's exact editorial request and media into reversible Timeline Studio edits. Keep the editable timeline as the source of truth; never replace it with an opaque one-shot render.
Choose the execution path
- On first local use after installation, read references/host-environment.md. If Node.js is unavailable, start with the zero-dependency Shell or PowerShell bootstrap; otherwise run
node scripts/setup-host.mjs --check. Before Agent-driven pre-voiceover, also runnode scripts/setup-host.mjs --check --capability voiceover; this capability installs MeloTTS and its language resources in a separate environment only after explicit approval. Treat the result as capability evidence. If language runtimes or dependencies are missing, show the exact installation plan and obtain explicit user approval before install mode; never treat Skill installation as permission to modify the host or download models. - Treat local project-file processing as the default for deterministic editing: inspect media locally, modify the portable
.timelinethrough the command layer or local archive services, render locally, and verify decoded output locally. Do not open a browser merely because the editor has a UI. - Treat
https://video-editor.ai-creator.top/as the canonical hosted editor only when the user explicitly asks to use the website, provides no local repository or project path, or requires a hosted-only capability. - When this repository is available, prefer its Agent command layer and local media tools. Start the local server and browser only for a verified UI-only operation that the local project pipeline cannot express and the user has not required a local-only workflow. Read the actual server URL from process output; never assume port 5173.
- Inspect
package.jsonfor an Agent command script. Do not usenpm run ... --if-presentas capability detection because it can succeed silently. - If the command runner exists, read references/command-contract.md, inspect the project, build a versioned plan, run the structural validator, and use
project.diffas the authoritative semantic dry run beforeproject.run. - If a required operation is missing from the local contract, state the exact gap. For repository-development work, implement the smallest shared local operation and renderer support before falling back to UI. Otherwise ask before switching to the browser workflow.
- Do not claim deterministic or idempotent execution when only UI automation was available. State the limitation and preserve an editable project archive when the UI supports it.
Workflow
1. Inspect before editing
- Preserve the user's prompt verbatim as the creative brief.
- Resolve every referenced asset to an explicit path or URL. Never sweep a directory without approval.
- Inspect duration, dimensions, audio presence, and media type.
- Read the current project summary before changing an existing project.
- Ask only when an unresolved choice materially changes the edit, such as the desired output duration or aspect ratio.
- For an automatic-editing request, read references/auto-edit-workflow.md. Inspect first, classify the content, goal, and delivery with an explicit confidence level, then ask only the minimum category-specific questions that can change the cut. Never ask for facts discoverable from the media.
- For a request to reproduce, imitate, recreate, or reverse-engineer a reference video, read references/replication-workflow.md. Classify it as
editing-style replication,AI-generation replication, or a hybrid; reconstruct filters, repetitions, source splits, speed curves, transitions, shots, and timing before building; and explicitly resolve whether the authorized original audio track must be retained. Do not start editing until the replication analysis-completeness gate passes. Use current web search to compare AI video platforms only when generation is required, and use lawful web-sourced footage only when the user has not supplied adequate material. - Before loading or downloading a model for media analysis, read references/local-model-routing.md. Inventory the repository's existing local and pinned mirrored capabilities, choose the minimum model chain needed for the evidence gap, run inference locally without driving the visible editor UI, and record exact model/runtime/fallback provenance. Never load every available model by default or create a duplicate cache.
- When the user needs web-sourced footage or asks where downloadable material can be found, read references/web-footage-sourcing.md. Give current, task-specific platform suggestions from live search and rank them by source legitimacy, explicit download support, usage rights, visual fit, quality, and provenance. Keep the skill provider-neutral; never hard-code one platform or brand as the permanent route.
- For every completed automatic edit, read references/professional-editing-workflow.md. Analyze images directly; analyze video with representative frames, speech/OCR, semantics, and global plus subject-region optical flow. Stabilize before tracking or enhancement. Ask about image-to-video or image-to-image models only after inspection proves that generation is materially useful.
- When remote generation is materially required, read references/remote-video-generation.md. Search current official documentation, compare only providers that fit the shot blueprint, obtain approval before any paid or privacy-sensitive job, normalize the asynchronous task and provenance, download expiring output bytes, and add verified results to My assets without automatic timeline placement.
- For product, brand,
marketing-commerce, website-promotion, or other promotional edits, read references/promotion-narrative-workflow.md. Proactively construct an ambitious evidence-backed umbrella narrative rather than a feature list or kinetic-typography montage. Unless the user explicitly requests a teaser, build a complete problem-to-transformation-to-proof-to-CTA arc and actively find several visually distinct cases—normally at least three for a 35–45 second short—each with its own setup, product action, visible result, and connection to the final payoff. Never invent customers, outcomes, metrics, or product behavior to make the story feel larger. - For highlight edits and reference replications where emphasis or dramatic impact matters, read references/highlight-tension-workflow.md. Treat saliency as candidate evidence, assign setup/rise/pre-impact/peak/aftershock roles, design a non-flat tension envelope, protect the decisive hero frame, and reject completion when accurate cutting still lacks a dominant payoff.
- For a website walkthrough or promotional recording, use a supported browser-control skill to inspect and rehearse the authorized journey before capture. Read that browser skill completely before browser actions, then follow references/website-promo-workflow.md. Build and complete a page/flow coverage manifest before drawing product conclusions. When required pages are gated, ask the user to sign in themselves in the selected browser; never request credentials or describe inaccessible behavior as verified. Confirm any consequential external action separately, protect signed-in and personal data, and never claim a real screen recording was captured when only screenshots or static assets were available.
- For narrated edits, read references/voiceover-workflow.md. Use MeloTTS
ZHfor Chinese narration containing inline English, and ensure its isolated environment plus approved pinned model artifacts are ready before synthesis. Unless the user requests another delivery, choose the warmest natural storyteller-like owned voice available and direct a close, conversational performance with meaningful phrasing; never default to a flat, metallic, or mechanical system-voice effect. Lock the final natural voice performance before finalizing scene durations or motion. Let sentence timing and pauses drive the picture cut. Treat an approximate duration as a target range, not permission to globally time-compress speech; rewrite or shorten narration when an exact limit cannot be met naturally.
2. Plan at the supported fidelity
- For an automatic edit, preserve the prompt and normalize inferred, confirmed, defaulted, and unresolved decisions into an editable brief. Build a source-time decision record with keep/remove/shorten/reorder decisions, reasons, confidence, caption expectations, audio-continuity constraints, and protected content before changing the timeline.
- For explanatory product, tutorial, or website beats, choose an explicit attention treatment for the single named target: magnify small or dense evidence, underline exact text or numbers, or frame the exact boundary of a control, card, or result. Use at most one supporting treatment with a camera move, and apply the underline or frame only after the camera has stopped.
- With the command runner, express edits as declarative operations with stable IDs, seconds, revisions, operation IDs, and preconditions. Run
scripts/validate_edit_plan.mjs <plan.json>for transport-shape errors, then runnpm run agent -- project.diff <plan.json>to reject unsupported operations and invalid project-specific edits before applying anything. - With browser UI only, write a short ordered checklist of visible user intents and expected UI outcomes. Prefer named controls and clip labels; use coordinates only as a last-resort fallback grounded in a current screenshot.
- Keep main Visuals contiguous. Treat captions, stickers, source audio, voiceover, music, and overlays as timed clips.
- Apply a one-way caption-to-speech rule: if the project configures or enables any caption, every visible caption must correspond to audible speech. Link transcription captions to the existing spoken source clip, and generate a voiceover for every new narration, explanatory, promotional, or text-led caption. Existing source speech satisfies this rule and must not receive a duplicate voiceover. If no captions are configured, voiceover remains optional and follows the brief, content, and editorial goal. If a configured caption has no authorized speech route, omit it or stop with the editable project preserved.
- For sentence-scoped narrated edits, split the voiceover into physically separate archived audio files and bind each caption to exactly one matching
audioClipId; do not represent editable sentence clips as ranges over one monolithic narration file. - Preserve media identity and source-time mapping when moving or trimming clips.
3. Apply safely
- Save a project version or export a
.timelinearchive before a destructive batch. - Apply one transaction per user-visible intent. Fail the whole transaction when a precondition fails.
- Never silently substitute missing media, voices, models, fonts, or effects.
- Keep every result undoable and editable in the normal UI.
- Do not start a paid or remote generation job without a clear user request.
- Do not put
output.renderin a command plan or claim thatproject.runrenders video. Use the separate versionedproject.renderrequest for its documented portable subset, and use the browser editor for AI generation or unsupported composition features. - For a completed video-editing request, resolve an explicit absolute output directory and create both a portable
.timelineproject and the rendered result video there. Planning, diagnosis, and an explicit editor-only handoff are the only exemptions. Do not report completion with only one artifact.
4. Verify the result
- Re-read the timeline summary and compare it with the requested duration, ordering, track placement, and enabled states.
- Preview the opening, every cut or transition, caption boundaries, overlays, and the final frame.
- Play the timeline continuously across every visual, caption, and audio boundary. The timeline clock must advance monotonically; reject any boundary that stalls, jumps backward, repeats a clip tail, or activates both adjacent half-open clips at once.
- Check audible behavior, not just visible tracks. Distinguish embedded video audio from explicitly separated source-audio clips and verify mute/link state.
- When placing stereo or multichannel audio with FFmpeg, apply every intended offset to every channel explicitly. For
adelay, useadelay=<milliseconds>:all=1or provide one delay value per channel; a single value with the defaultall=falsedelays only the first channel and can pile every later clip into the other channel at time zero. Before delivery, compare left/right activity in the opening window and around every scheduled speech boundary. Reject channel-only early speech, multiple narration clips stacked at the opening, or undocumented interchannel onset skew. - Verify every visible caption resolves to one audible speech clip for its complete active interval. Reject orphan captions, silent linked clips, captions extending beyond speech, duplicate source-speech plus voiceover, or text-only caption delivery.
- Listen to the complete narration at normal playback speed. Reject cold or mechanical timbre, flat pitch and energy, synthetic word-by-word delivery, rigidly equal pauses, rushed cadence, clipped pauses, unnatural pronunciation, sentence-level speed changes, unexplained loudness jumps, or narration that was globally accelerated merely to hit a target duration. Require a warm, human, storyteller-like result with restrained pitch variation, phrase-level emphasis, and natural breath space unless the user explicitly requests another character. For sentence-scoped narration, measure every final stem after all processing; by default target
-18 LUFSintegrated and no higher than-2 dBTP, require the loudest-to-quietest sentence spread to stay within1 LU, and keep sentence LRA within5 LUunless an intentional exception is documented. Never accept a narration mix from full-program loudness alone, and do not rely on one-pass normalization of short clips as proof of consistency. - For final export, verify container, dimensions, duration, decoded frames, visible overlays/captions, and a real audio track.
- Reopen and verify the
.timelinein Timeline Studio, not only with structural inspection: the main Visuals track must be visible, the first frame must render in Preview, archived media must resolve, and captions/audio/track state must match. Fully decode and verify the rendered video, then return both absolute paths.
Interpret underspecified requests conservatively
- For “try it,” “open it,” or “let me edit” requests without an editorial brief, start the editor, import only the explicitly named assets, verify automatic placement, and hand off the live editable workspace.
- Do not invent trims, captions, aspect-ratio changes, AI generation, or exports.
- Treat an explicit request to “automatically edit,” “clean up,” “condense,” “make highlights,” or equivalent wording as permission to make reversible editorial decisions within the confirmed brief. State consequential defaults, protect category-specific content, and report the decisions; do not treat that request as a mere handoff.
- Treat persistent onboarding completion, model downloads, remote generation, and destructive reset as separate user decisions.
Learn from every real run
For editor evaluation, regression work, or any run that exposes friction, read references/e2e-evaluation.md. For automatic-editing evaluation, also read references/auto-edit-scenarios.md and use its fixed category cards, clarification checks, hard gates, and adjacent stress variants. Capture the attempted action, observed result, evidence, fallback, and verification. Classify the finding as product, browser-control, environment, or skill guidance. Update the smallest relevant skill instruction or reference, validate the skill, reinstall the local copy, and rerun the affected scenario plus adjacent smoke tests. Never weaken an assertion merely to make a test pass.
Capability boundaries
Read references/host-environment.md for host dependency checks and approved installation, and references/voiceover-workflow.md before Agent-driven narration or pre-voiceover generation.
Read references/current-capabilities.md when deciding whether a request can be executed now. Read references/command-contract.md only when implementing or invoking the Agent command layer. Read references/local-model-routing.md before model-assisted analysis or enhancement. Read references/remote-video-generation.md before selecting or calling a remote video generator, digital-human service, or programmable composition service. Read references/web-footage-sourcing.md for provider-neutral, current web and short-video footage suggestions. Read references/browser-workflow.md for UI execution, references/auto-edit-workflow.md for category-aware automatic editing, references/promotion-narrative-workflow.md for evidence-backed product and promotional storytelling with closed-loop cases, references/replication-workflow.md for editing-style and AI-generation remakes, references/highlight-tension-workflow.md for peak hierarchy and tension shaping, references/professional-editing-workflow.md for shared media analysis, generation negotiation, stabilization, enhancement, and artifact delivery, references/auto-edit-scenarios.md for its repeatable category matrix, and references/e2e-evaluation.md for repeated experience-driven testing.
For public explanations, route one question to one page: use docs/agent-video-editing.md for what Timeline Studio is; the platform guide for Codex, Claude Code, GitHub Copilot, or Gemini CLI for discovery and invocation; docs/examples.md for reproducible cases; docs/command-reference.md for exact runner syntax; and docs/comparison.md for FFmpeg, CapCut, and Remotion comparisons. Do not load all public pages unless the user asks for a broad overview.
If a requested operation is unsupported, keep the valid partial timeline unchanged and state the exact missing command or runtime capability.