What does paper2video do?
Make a video that a stranger with no context can follow, that stays true to the work, and that looks good enough to share: real data on screen early, one idea per beat, smooth motion that explains rather than decorates.
You are the lead: you plan, build the shared pieces, hand out well-scoped work to subagents when useful, integrate, and own correctness. Work autonomously end to end unless the user asks for checkpoints; show the user results (stills, drafts, renders), not questions, whenever a sensible default exists.
Supporting material (read when you reach that stage):
reference/rigor.md— claims ledger, data handling, independent fact-checkreference/voice.md— narration backends (ElevenLabs / free edge-tts / self-hosted GPT-SoVITS / the user's own recording / none), QA, pronunciation, open-source TTS know-howreference/visual-style.md— design system, motion rules, data honesty, review loopreference/platforms.md— YouTube + Bilibili deliverables: specs, covers, titles, descriptions, chaptersreference/vertical.md— the 9:16 phone cut: safe area, type sizes, one focus per beat, hook, building it from the same audioreference/handoff.md— using the user's own materials, and exporting for editing softwarereference/third-party.md— starting from only an arXiv link/title: finding source, code and data; extracting numbers from LaTeX/figures; small CPU illustrations; faithfulness when the paper is not the user'sscripts/— the pipeline (seescripts/README.md);template/— the Remotion starter (seetemplate/README.md)
0. Intake (keep it short)
You need: the work — an arXiv link or a paper title is enough (then follow reference/third-party.md to find source, code and data yourself); a repo, path, project page or readable remote machine also works. Whose work: the user's own (first person allowed) or someone else's (third person, faithful, credited — reference/third-party.md). Languages and platforms (default: an English cut for YouTube and a Chinese cut for Bilibili, both 16:9, plus covers in several sizes and titles/descriptions/chapters for both), and narration (see reference/voice.md; default: ElevenLabs if a key is available in env/.env, otherwise edge-tts, and say so). Ask only what you cannot default. Also ask once: first person ("our paper") or third person ("this paper"), and anything that must not appear (unpublished side projects, anonymous submissions, private data).
Treat every remote/shared resource as read-only unless told otherwise; list files and sizes before pulling anything large (ask above ~500 MB); never print or commit keys.
1. Set up the workshop
- Create a Remotion project in a new folder (e.g.
npx create-video@latest --blank, or copy the latest Remotion blank template); do not make the user install anything by hand — install Node deps yourself, and checkffmpeg,python3anduv(installuvvia its official script or pip if missing). - Install the official Remotion agent skills if they are not already available (
~/.agents/skills/remotion-*or listed in your skills):npx -y skills add remotion-dev/skills(see remotion.dev/docs/ai for the current command). Follow them for Remotion API details. - Copy
template/into the project (mergepackage.additions.jsonintopackage.json,npm install), copyscripts/to<project>/scripts/, createnotes/,data/(raw, git-ignored),public/data/,out/(current deliverables only, fixed names:reference/platforms.md→ Files),review/(disposable stills and contact sheets), and a projectCLAUDE.mdthat records: the goal, audiences, languages, voice choice, the "must not appear" list, source-of-truth document, and any user feedback as it arrives (dated). Keep it current — it is how later sessions (and subagents) stay consistent. npx tsc --noEmitand render one still of the template to prove the toolchain works.- Making several videos? Use a workspace: one folder holding a shared
package.json/node_modules(one install, ~400 MB instead of one per video; all@remotion/*at one version) and a shared.env, plus a workspaceCLAUDE.mdwith the standing rules for every video (Claude Code reads parentCLAUDE.mdfiles). Each video is its own folder (src/,public/,scenes.json,notes/,data/,out/,review/,audio_cache/,scripts/, its ownCLAUDE.md) withnode_modulesand.envlinked to the workspace root;scripts/new_video.sh <ShortName>(copy it to the workspace root) creates one as<NNN>_<ShortName>, numbering down from 999 so that sorting by name lists the videos first, newest on top. After changing the shared versions, type-check every video and compare a few rendered frames with its final file.
2. Understand the work before writing a word
Collect the material first (reference/third-party.md §1–2 if you only have a link or title). Read the paper end to end (and appendix), the README / project page, and skim the code for what the figures are made from. Find the core finding and why it is surprising, what makes the work fun to watch (the most surprising observation, the cleverest experiment, the example that makes people smile), the authors' intent (the question each experiment answers, why they set it up that way, what they tried), the 2–4 claims that carry the story, and the visuals that can show them with real data (figure-data exports, result JSONs, logged metrics, demo outputs). Prefer exported figure data or raw results over digitising plots. Write notes/understanding.md: the one-sentence story, the claims with their numbers and settings, the candidate visuals, and anything confusing or easy to overstate.
3. Story, script and claims
- Length and focus come first. Default: 3–4 minutes per cut, at most 4:00 unless the user asks for more — about 450–600 English words or 1,100–1,300 Chinese characters of narration. A video is not a reading of the paper: pick the one finding the whole film serves and the 2–3 pieces of evidence that carry it, and give everything else one sentence or leave it for the description. Faithful means nothing is distorted, not that every section appears.
- Make it fun, rigor first. Choose the most interesting experiments and examples as the evidence beats, and show each the way the authors set it up — the question, what they changed, what happened — so the viewer gets the idea and the authors' intent, not just a number. Humour and surprise come from the real material (a striking example, an unexpected result, a clever control), never from stretching a claim; if a fun line would overstate the paper, drop it.
- Time budget for a 4-minute film: hook 15–25 s (the paper's headline result is on screen, as real data, within the first 30 s) · primer 30–40 s · the evidence ≥ 40 % of the film · mechanism/theory ≤ 20 % and only what can be shown · takeaway ≤ 15 s. For a well-known older paper, an impact beat (≤ 20–30 s, before the takeaway): what it started or changed, told with sourced facts — named follow-up work with years, where the idea is used now; a citation count only with its source and date; no unsourced superlatives. No limitations scene by default: only when the paper itself treats its limitations as a central part (then ≤ 20 s); otherwise the scope lives in the settings you state with each result. Derivations, extra experiments and lists of implications get a scene only if they can be shown with real data in ≤ 20–30 s.
- Cut before you pay:
python3 scripts/vo.py build --lang <l> --dry-runprints the film length estimated from reading speed (within ~5 % of the voiced length); trim the script until it fits, then generate the voice. - Structure (adapt, don't force): hook with the strongest real visual (≤ ~15 s) → a full-screen title card (
template/src/components/TitleCard.tsx: the paper's title, authors and institutions, venue or arXiv id, its one-sentence claim, and "unofficial explainer" for someone else's paper; plus, as a dimmed background element, the first page of the paper's PDF — title, authors, affiliations, abstract; the first screen of its web page only when there is no PDF — easing in while the title and institutions are read:node scripts/paper_shot.mjs <arXiv id | PDF | URL>→public/shots/paper.png,TitleCard shot="shots/paper.png"; centre-right at 16:9, under the title at 9:16; not on the covers) while the narration names the paper and the authors' institutions (not the authors' names: TTS mangles them, and a name guessed into another script is often wrong — on screen and in the copy, write names exactly as the paper or arXiv spells them, never transliterated) — a scene of its own, on screen by ~20 s at the latest → a primer that lets a no-context viewer follow (what is the task/object, in plain words, animated) → the problem, shown with data → the idea → why it works (mechanism) → results, each with its setting → takeaway, ending on the paper's point. No "links are in the description" / 「论文和代码见简介」 line, spoken or on screen: viewers know where links live, and it ends the film on filler (the description and the title card carry the identifiers). Put good visuals early; don't save them for the end. - Write for the ear: short sentences, one number per sentence, concrete nouns, no "Not X but Y" tics, no fake Q&A, even tone across scene boundaries (an abrupt "Now the fun part!" jars). Give formulas and charts a beat of silence.
- Put the script in
scenes.json(schema:template/scenes.example.json): per scene and language, lines withtext(captions),tts(what is spoken: respellings, sparse audio tags),anchors(words that time visual beats), optionalpauseAfterMs; per scene an optionalchapter. Write each language natively rather than translating word for word; keep technical terms in English where the audience expects them. - Start
notes/claims.mdnow (reference/rigor.md): every number and factual sentence → its source. Writenotes/storyboard.md: per scene the visual, the anchor-driven beats, the data file, the source line.
4. Data → public/data/*.json
One agent (you or one subagent) owns turning raw material into clean JSON with a source field per file, and writes notes/data_report.md that cross-checks every headline number against the paper and lists discrepancies and pitfalls (metric variants, renamed terms, rounding). Scenes only read public/data.
5. Voice first, then picture
Build the narration before animating (reference/voice.md): scripts/vo.py build --lang <l> produces per-scene audio and public/data/vo.<l>.json with word timings and anchor frames. The picture is then keyed to anchors, so both languages stay in sync automatically. Check the report: ASR round-trip errors, odd prosody, missing anchors; fix wording/respellings, regenerate only affected scenes (outputs are cached by content hash).
6. Build the scenes
- You build the theme, shared components and anything reused across scenes first (see
reference/visual-style.md); then scenes can be parallelised across subagents with explicit file ownership (one scene set per agent, nobody edits shared files — they report requested changes back to you). Give each agent:CLAUDE.md, the storyboard entry, thescenes.jsonlines and anchor names, the data files, the style rules, the chapter number and title for each of its scenes (parallel agents otherwise number their chapter tags independently), and the requirement to render and inspect stills in every language before reporting (scripts/review.py). Afterwards re-runscripts/zh_font_subset.pyonce (new Chinese characters fall back to a system font until then) and, after every Chinese voice build,scripts/zh_breaks.py(the captions of both cuts break Chinese only between words). - The title card's paper page: render it once the paper is fixed (
node scripts/paper_shot.mjs <arXiv id>downloads the PDF todata/and renders page 1; a local PDF or a web article's URL also works; never a hand-made screen grab), pass it asshot, keybeats.shotto the title words and, at 9:16,beats.shotDimto where the narration moves on to the institutions; check both cuts' stills (text stays readable over the page, captions clear). - Every scene: headline, beats driven by anchors (with fallbacks), a source line on data shots, SCHEMATIC on illustrations, both languages' on-screen strings, content clear of the caption band.
- Review: render stills at anchor-based frames for the whole film in every language and tile them into contact sheets (
scripts/review.py <Composition> <cut> [--scenes ...], one bundle for all frames), and look — overlaps, clipping, empty frames, illegible text, anything off-message. Then review the pacing, reading the sheets as the story a viewer gets: flag any stretch over ~20 s without a new real visual, any list of three or more items read out, any formula or theory that is not tied to data on screen, and any scene that does not move the one finding forward — compress or cut those. Iterate.
7. Independent fact-check
Spawn a fresh agent that did not build anything to check the script and every on-screen string/number against the paper (prompt in reference/rigor.md). One light pass is enough by default: the script and copy right after the script is final (before the paid voice, when wording fixes are cheapest) — numbers, settings, overclaims. Check the screens yourself against notes/claims.md while reviewing the sheets; add a second independent pass (screens, plotted data, covers) only for a high-stakes film or when the user asks. Fix MUST-FIX items (including voice-over wording, then regenerate those scenes and re-time the music: music.py --refit for a generated bed, the same --file … --keep-ending N again for a library track) and most SHOULD-FIX items.
8. Sound, render, verify
Music (optional) ducked under speech: by default a royalty-free track (scripts/music.py --file track.mp3 --keep-ending 20 — CC0 / public domain / CC BY with the credit line in every description; keep a shared library of tracks and their credits in a workspace); composing one with --generate costs credits and is opt-in; a few procedural SFX on visual beats (scripts/sfx_synth.py, cues in src/timeline/sfx.ts; subtle, ≤2 "hits" per film). Render and master with scripts/finalize.sh (−14 LUFS / −1.5 dBTP), export captions with scripts/srt.py. Verify every file with scripts/verify.py <mp4> --cut <lang>: streams, duration against the timeline, loudness, a full-mix ASR pass against the script (with the key terms), and full-frame motion runs (zoom/drift/jitter) to look at. Report honestly what was checked.
9. Deliverables and handoff
Per reference/platforms.md: horizontal cuts per language, a vertical cut for phones (native 9:16 scenes from the template's src/vertical/, same audio, reference/vertical.md), covers (16:9 and vertical; scripts/covers.sh renders the set into out/covers/), SRTs, and out/social_copy.md with titles/descriptions/chapters per platform (plus any other platform the user asks for — research its current conventions then). Everything goes into out/ under the fixed names in reference/platforms.md (Files); a new version overwrites the old one — no _v2 / _old copies (move a version you must keep to the Trash or review/). If the user wants to finish in an editor or bring their own footage, follow reference/handoff.md.
Efficiency without losing quality
Quality first; cut only waste. Most waste is rework and repetition:
- Decide before you pay. Before generating any voice, settle structure and length: hook → title card by ~20 s, 2–4 minutes, one story; run
vo.py build --dry-runfor the length estimate; apply the voice's known text rules (respellings, no English in a single-language open-source voice, no fragile short clauses); anchors are words oftext. A script change after the voice exists costs credits, a re-pick and a re-render. - Agents only where they save wall time: independent heavy work (figure-data extraction across many figures, many scenes in parallel). A ~10-scene film is usually faster built by the lead, which already holds the context; every subagent re-reads the material. Give each one a self-contained brief. One light fact-check, not a checker per stage.
- Read once, then look things up. Read the paper yourself, in full, once — as clean text (the LaTeX sections or the extracted article text, not raw HTML with scripts and data URIs). Distil the paper into
notes/understanding.mdandnotes/claims.md; later use grep / smallsed -nranges, inspect a JSON's keys instead of dumping it, and don't re-read files you already summarised. - Review what changed. Contact sheets only for the scenes you touched (
review.py --scenes,vreview.py --scenes, scale 0.5); one full-film sheet before the final render. Render full cuts only after the sheets look right; verify with--asr nonewhen the audio did not change. - Long jobs in the background (renders, voice builds, ASR): keep working meanwhile and wait for the completion notice; no sleep/poll loops.
- Keep tool output short. Everything printed stays in the context and is re-read on every later call (by mid-film that is a few hundred thousand tokens per call): tail logs, grep for PASS/FAIL/WARN, print counts and diffs of a few lines, write long reports to files and read only the part you need.
- Reuse: components from earlier films (title card, figure panels, colour bars), voice settings that already worked, music from the library.
Running on a headless machine
Everything works without a display or GPU: Remotion renders with headless Chrome (downloaded on first use; on Linux it needs the usual shared libraries — npx remotion browser ensure reports what is missing), the Python/ffmpeg tools are CLI-only, faster-whisper runs on CPU. Preview by rendering stills/MP4s, or run npx remotion studio and forward the port (ssh -L 3000:localhost:3000 host). Install Node (nvm), uv and a static ffmpeg in the user's home directory if there is no root.
Iterating with the user
Expect feedback on tone, pacing, clarity and specific frames (they quote timestamps — map them to scene + anchor with scripts/timeline_info.py). Record durable preferences in the project CLAUDE.md. Re-render only what changed; keep chapters/timestamps in the copy in sync after timing changes.
Hard-won rules (short list — details in the references)
- Fun from the real material, never at the cost of rigor: the paper's own surprises, clever experiments and examples, shown the way the authors meant them.
- Hook → title → body: within 10–20 s the viewer must know what they are watching (which paper, by whom, what it claims). Keep the hook short, then a full-screen title card that the narration reads out; do this in every cut, vertical included.
- Shorter is better: the goal is the paper's idea, its main experiments and its conclusion — usually 2–4 minutes (≤ 4 unless the user asks for more); cut anything that doesn't serve that. One finding and its best evidence; a viewer who is bored at minute three never sees the rest.
- Truth over hype: no number without a source; don't state claims more strongly than the paper; say the setting behind each number (model, task, metric). Skip the limitations section unless the paper makes it central.
- Real data on real-data charts; illustrations labelled SCHEMATIC; no count-up number animations; bars from zero unless the axis says log.
- No full-frame shake/punch-in/zoom effects — emphasise the specific card; full-frame scale changes also caused visible text jitter in renders.
- The production tool is a credit line, not the pitch (unless the user explicitly wants that angle).
- Never print, log or commit keys; check paid-API quota before every batch and stop if it would be exceeded.
- Keep the disk clean: renders and bundles are big. Use
scripts/stills.mjs(it deletes its bundle and closes its browser, also on Ctrl-C); never call Remotion'sbundle()without deleting the result; review stills go toreview/, not scratch or system temp folders; check free space before a full render. Close what you start when you are done with it (Remotion Studio, background renders, servers, background shells) — a leftover headless browser keeps eating memory. If the user keeps work on an external/NAS volume, point temp files and tool caches there too: setPAPER_VIDEO_TMPDIR(the template'sremotion.config.ts,stills.mjsandcommon.pycopy it intoTMPDIR; some launchers, Claude Code included, resetTMPDIRitself) andUV_CACHE_DIR/HF_HOME/npm_config_cachein.claude/settings.json→env.