Communitygithub.com

josueh04/product-video-skills

Gather pixel references for UI that the code cannot show, or to check a rebuild against the real thing, using frames and timed OCR text from screen recordings, captures recovered from past Claude Code session transcripts, a local instance of the app built like production, and web research for third-party apps, plus side-by-side parity images and contact sheets. Use it whenever someone hands over a screen recording (.mov or .mp4) of the product, when a screen has no usable source code (a stale checkout, another company's UI such as a sign-in page, calendar or CRM, runtime output from a backend not in the repos), when asked "what does it really look like", "match the recording", "how long does that animation take in the app", "compare our render to the real app", or when screenshots from an earlier session might already exist. Never uses the reviewer's personal browser.

product-video-skills 是什麼?

product-video-skills is a Claude Code agent skill that gather pixel references for UI that the code cannot show, or to check a rebuild against the real thing, using frames and timed OCR text from screen recordings, captures recovered from past Claude Code session transcripts, a local instance of the app built like production, and web research for third-party apps, plus side-by-side parity images and contact sheets. Use it whenever someone hands over a screen recording (.mov or .mp4) of the product, when a screen has no usable source code (a stale checkout, another company's UI such as a sign-in page, calendar or CRM, runtime output from a backend not in the repos), when asked "what does it really look like", "match the recording", "how long does that animation take in the app", "compare our render to the real app", or when screenshots from an earlier session might already exist. Never uses the reviewer's personal browser.

相容平台✓Claude Code~Codex CLI✓Cursor
npx skills add https://github.com/josueh04/product-video-skills/tree/HEAD/skills/ui-reference-capture

在你喜歡的 AI 中提問

開啟一個已預先載入此 Agent Skill 的新對話。

說明文件

UI reference capture

The code gives structure, styles, strings and icons. It does not give what happens at runtime (what a backend answers, in which order things appear, how long the real app waits), and it says nothing about another company's UI. This skill fills those gaps with references, in a fixed order of preference, and it never improvises in someone's personal browser.

That last rule has a history. In the production these skills come from, an agent was allowed into the reviewer's own logged-in browser to capture a live app. The live flows hung, the agent kept doing "cleanup" steps after the reviewer said stop, it left stray items in a test workspace, and the session ended badly after about 68 minutes with nothing usable built. The same captures were later recovered from the session transcript in minutes. Offline references first; a browser only when the user asks for one, isolated, and stopped the moment they say stop.

Where things go

  • Inputs (recordings, exports, screenshots the user hands over): copy them into products/<slug>/references/<name>/ as soon as they arrive. Downloads folders and the OS's temporary capture folders get cleaned, and one recording in the source production lived in a folder the system would have deleted. macOS recording names contain a narrow no-break space (U+202F) before AM/PM: find them with a glob, not by typing the name.
  • Working material (thousands of frames, recovered captures, raw OCR): a scratch folder outside every repo, e.g. $TMPDIR/pvs-<slug>-<name>/.
  • Results worth keeping (the OCR timeline, key frames with real data cropped out, the analysis): references/<name>/, cited from SOURCES.md.

Recordings and captures of a real workspace hold real customer names, emails, phone numbers, internal ids and other people's conversations. They never go into the workbench repo, and the video never shows them: section 1 step 6 lists what must change.

Choose the technique

You haveTechniqueRead
A screen recording of the product1. Analyze the recordingreferences/recording-analysis.md
Past sessions that looked at the app2. Recover captures from transcriptsthis file, section 2
An app that can run locally3. Local instance built like productionreferences/local-instance.md
Another company's product4. Web research, public sources onlyreferences/third-party-ui-prompt.md
Nothing (a phone call, an end card, generated text)Design it from the kit and mark itproduct-truth (mock)

Try them in that order. Before asking anyone to record, check section 2: captures may already exist. If none of the four works, say so; do not fall back to the reviewer's browser.

1. Analyze a screen recording

PVS_HOME="$(cd "$(cd "${CLAUDE_SKILL_DIR}" && pwd -P)/../.." && pwd)"
S="$PVS_HOME/skills/ui-reference-capture/scripts"
W="$TMPDIR/pvs-acme-board"                                    # scratch, outside the repo
bash "$S/frames.sh" references/board/rec.mov "$W" --fps 10 --logical-width 1440
bash "$S/ocr.sh" "$W/logical" "$W/ocr.jsonl"                  # macOS only
"$PVS_HOME/bin/pvs-py" "$S/ocr_events.py" "$W/ocr.jsonl" references/board/ocr-timeline.txt --fps 10 \
  --region "panel=0.6,0,1,[email protected]"
"$PVS_HOME/bin/pvs-py" "$S/grid.py" sheet "$W/overview.png" "$W/logical" --every 20 --cols 4
  1. frames.sh writes small thumbs (change detection) and frames at the app's logical width (a retina recording is 2x: logical width is half the pixel width), so one frame pixel is one CSS px of the rebuilt app.
  2. ocr.sh compiles ocr.swift (Apple Vision) once into a cache folder and writes one JSON line per frame. On Linux it exits with "unsupported": read the contact sheets instead.
  3. ocr_events.py turns per-frame OCR into stable spans, start-end [region] text. That timeline gives the exact strings the real app shows and when, which beats guessing from a video by eye.
  4. grid.py sheet makes contact sheets: an overview every 2 s, then dense 10 fps sheets per phase. grid.py overlay draws a logical-px grid on a frame for measuring.
  5. Save key frames as references/<name>/frames/<t>-<state>.jpg (cropped of real data).
  6. Write references/<name>/analysis.md with the structure in references/recording-analysis.md: the real timeline with exact strings, measured motion (how long a panel takes to open, how long the cursor rests), a "must change for the video" table (real names, emails, phones, workspaces, internal ids, long waits, OS chrome such as the Dock), and product gaps to raise before anyone builds on them (a step the story needs that the recording shows failing).

Recordings hide gaps when waits are compressed; note every wait you shorten.

2. Recover captures from past sessions

Every image a tool returned in a Claude Code session (browser screenshots, window captures, read images) and every image pasted into the chat is stored in that session's JSONL.

"$PVS_HOME/bin/pvs-py" "$S/recover_transcript_images.py" ~/.claude/projects/<project>/<session>.jsonl \
  [--from-line N] [--tool screenshot]

With no output folder it writes to a fresh temp folder and refuses any folder inside a git work tree, because these captures show real accounts. It writes index.tsv (file, time, tool, input) so you can find the screens you need. Review them in contact sheets (`grid.py sheet

3. A local instance built like production

The most faithful reference when the app can run on the user's machine: build it from the production commit with the production build arguments, seed it with fictional data, isolate it from every shared system, and capture it with a headless browser. Read references/local-instance.md before starting; the user must agree to running their app, and it never uses their accounts or shared databases.

4. Third-party UI

For another company's screens, launch a research subagent with references/third-party-ui-prompt.md: public sources only (official docs, help centers, press kits, marketing screenshots, public embed demos owned by the vendor), never a login, never a submit or a booking, about 30 minutes. Every value comes back marked verified, measured or inferred. Prefer the product's own component when it has one (a product that renders a preview of a chat app has a preview component; use it instead of rebuilding the chat app). Show a full third-party screen only when it proves a real integration; elsewhere use logos and badges inside the product.

Parity: prove the rebuild matches

"$PVS_HOME/bin/pvs-py" "$S/pair.py" snapshots/nocam-12.4s.png references/board/frames/12.4-board.png \
  640 40 1100 300 "$W/pair-board-header.png" --ref-scale 2 --logical-width 1440 --render-width 1920

Take our snapshot with the camera at scale 1 (a copy of the composition with #camera{transform:none!important}), crop the same logical box from both, and stack them. The script prints the mean absolute difference per channel; iterate until it stops dropping and the stacked image shows no offset. Measure in app px and convert explicitly: confusing scales was a recurring source of wrong coordinates.

Rules and why

  • Never the reviewer's personal browser or profile, and nothing taken from it (fonts, icons, storage). When anyone says stop, stop every browser action at once, cleanup included.
  • Never sign in, never read tokens or browser storage. If a capture needs a login, the person types their own password in their own session, or the capture does not happen.
  • Recovered and recorded material stays private. Temp folder first, cropped copies only.
  • Mark every inferred value. A reference that is half guessed must say which half.

Individual skills in this repo

This repo contains 14 individual skills — each has its own dedicated page.

josueh04/product-video-skills

Extract, once per product, everything every video of it reuses and write it to kit/ and product.yaml (design tokens, font subsets as woff2, icon subsets as SVG from the product's own icon packages, logos in light, dark and app-tile variants from the repo, a fictional cast proposed once for veto and then frozen, the canonical-names map, pronunciations, banned terms and the read-only tool list). Use it when a product is set up or its kit is missing or incomplete, when a video needs an icon, font or logo that is not in kit/ yet, when someone asks for demo names, fake customers, phone numbers or emails, when a brand word is mispronounced or an old product name shows up, and when checking that demo data is fictional.

josueh04/product-video-skills

Interview the user about a product, then create products/<slug>/ with its own git history, a filled product.yaml and linked skills, fetch its sources and build its kit. Run only when the user types /product-new.

josueh04/product-video-skills

Back every sentence of a video's narration (audio/lines.tsv) and every screen it shows with a citation into the pinned source code (role/path:line@sha) or a docs URL, mark what is visible in the UI versus backend-only, flag restricted or unreleased features, and cut or rewrite anything unbacked; writes the video's TRUTH.md and checks it with truth_check.py. Use it whenever a script or narration is drafted or edited, before voice is generated, before a build, when someone asks "can we say this?", "is this true?", "does the product really do X?", when a reviewer asks for a feature or a claim the product may not support, and when a source document (pitch deck, PRD, marketing page) makes claims the video wants to repeat.

josueh04/product-video-skills

Check a rendered product video before anyone else sees it: worker-pattern flicker, black frames, loudness and true peak, clipping, clicks at clip edges, overlapping narration, speech to text against the script, banned terms and legacy names, camera zoom, contact sheets, frame strips at transitions and parity against the approved version; then write qa/REPORT.json, the only thing deliver.py accepts. Use it after every HyperFrames render, whenever someone asks "is the render clean", "QA this", "check the video", "check the audio", "why does it flicker", "there is a click", "compare v3 with v2", "did the approved part change", or before showing, sending, uploading or delivering any MP4, even when the request does not say QA. Also use it to triage a defect a reviewer reported in a render.

josueh04/product-video-skills

Write a product video's narration and turn it into voice clips with word timings, pronunciation fixes, sound effects and even loudness. The script becomes a table of moments and then audio/lines.tsv (one clip per sentence, with role and speed columns); tts.py voices it with ElevenLabs or the free macOS say voice, maps brand respellings back to the on-screen spelling, normalizes every clip and writes audio/timings.json for the composer; make_sfx.py builds typing tracks from real keystrokes and places recorded click and pop sounds. Use it whenever a video needs a script, narration, voice-over, lines.tsv, timings.json, TTS, a new take, a voice or casting choice, a pronunciation fix ("it says the name wrong"), a changed sentence, a tone note ("too hype", "sounds cut off"), audio levels, a click at the end of a clip, typing or click sounds, or when the build stage asks for the voice. Also use it for silent loops, which still need a moment table and SFX.

josueh04/product-video-skills

The animation rules that keep HyperFrames' parallel render workers from dropping, flashing or flickering elements, plus a static lint (lint_motion.py) that finds the violations in a video's template before it costs a render. Use it whenever you write or edit GSAP tweens, timelines, cursors, camera moves, typing, scrolls or pop-ups in a HyperFrames composition or a video's src/template.tpl, whenever a render shows flicker, stutter, an element that vanishes on some frames, a title that flashes, or a "WORKER PATTERN" line from scan_render.py or qa.py, and whenever the preview looks right but the MP4 does not. Also use it to review someone else's timeline code before rendering.

josueh04/product-video-skills

Pin a product's source code (read-only exports in sources/ plus sources.lock), confirm that the pinned commit is what runs in production, and map every screen of a video brief to its route, components, i18n strings and state, written to the video's SOURCES.md. Use it whenever a video needs to know where a screen lives in the code, when sources/ is missing or stale, before ui-spec-from-code or product-truth start on a video, after the product's frontend changed ("what changed", "which videos are affected", "refresh the sources", "is this checkout current", "which commit is in prod"), and when there is no code and you need an inventory of the no-code references (recordings, recovered captures, docs) a screen can be rebuilt from.

josueh04/product-video-skills

Compose a narrated product demo in HyperFrames: the stage (the product UI rebuilt at its real viewport and scaled to 1080p, camera, rack focus with veil, chapter titles, cursor and clicks, typing, streaming text, pop-ups, toasts, scrolls, end screen and lockup) and the build.py that anchors every beat to a word of the narration. Use it whenever you write or edit a video's video/build.py, src/template.tpl or src/app.css, place a beat on a word, add a chapter, a click, a pop-up or a push-in, frame a screen, build the end screen or lockup, snapshot setup beats, or render a draft or delivery MP4 of a product video in this workbench. Also use it when someone says "the cursor is off", "too zoomed in", "too fast", "it feels chaotic", "the title flashes", "sync the UI to the voice", or asks for a walkthrough, demo or pitch video of a UI.

josueh04/product-video-skills

Turn a product's real frontend code into 1:1 rebuild specs for a video, one read-only subagent per screen, each returning static HTML, CSS with every variable resolved to its literal value and cited (role/path:line@sha), every state, transitions with exact durations and easings, icons from the code's own icon sets, and the exact i18n strings; plus resolve_tokens.py to write the product's design tokens to kit/tokens.css. Use it whenever a screen of the product has to appear in a video, when writing or fixing video/src/app.css or the template markup, when someone asks for exact sizes, colors, fonts, paddings, animations or icons of a screen, when a rebuilt screen "looks off" next to the real app, and when the design tokens or theme of a product need extracting. Framework adapters cover Angular with PrimeNG (proven), React, Vue, Tailwind and plain HTML (unproven).

josueh04/product-video-skills

Say where this session stands in the Product Video Skills workbench (setup state, which product and video the current folder belongs to, the stage of every video) and the exact next command to type. Also answers "how do I..." questions about the workbench from its docs.

josueh04/product-video-skills

Coordinate the build of one or more signed videos of the current product with subagents (source recon, product truth, UI specs, voice, one builder per video), re-run QA itself, then deliver. A light coordinator that never builds itself. Run only when the user types /video-build.

josueh04/product-video-skills

Start a new video of the current product - create videos/<video>/, write BRIEF.md, the feature coverage matrix (COVERAGE.md) and the claims sheet (CLAIMS.md), propose chapters, then stop for the reviewer's sign-off. Never builds. Run only when the user types /video-new.

josueh04/product-video-skills

Turn a batch of reviewer feedback on the product's videos into one table per video, fix every video that got notes in parallel (one subagent each) while keeping approved parts, re-run QA and parity, bump versions and deliver. Run only when the user types /video-review.

josueh04/product-video-skills

Check this machine and install the pinned video toolchain of the Product Video Skills workbench (HyperFrames CLI, its rendering Chrome and its agent skills from the same release, the Python environment, the speech model for QA), then run the self-check. Safe to run again. `/video-setup check` only reports.

相關技能