Communitygithub.com

DeepDiveTechLab/six

Build SIX — Scrolling Immersive Xploration: an immersive scroll-scrubbed "fly through the world" landing page for any industry or brand using MuAPI. As the visitor scrolls, a pre-rendered camera flies from outside each scene into its interior — or in one continuous forward arc from the top scene into the ones below it — then flows on to the next scene with NO cuts: one continuous connected flight (exploded-view arc, isometric diorama world, or any art direction you pick). The skill interviews the user for the topic, the story beats/sections, and brand kit, then generates cohesive scenes + seamless camera clips with MuAPI and wires a portable, framework-agnostic scroll-scrub engine. Use when the user wants a "3D world" / "browse-through-the-industry" hero, a scroll cinematic, a diorama landing, or to turn a business into a scrollable world.

Was ist six?

six is a Claude Code agent skill that build SIX — Scrolling Immersive Xploration: an immersive scroll-scrubbed "fly through the world" landing page for any industry or brand using MuAPI. As the visitor scrolls, a pre-rendered camera flies from outside each scene into its interior — or in one continuous forward arc from the top scene into the ones below it — then flows on to the next scene with NO cuts: one continuous connected flight (exploded-view arc, isometric diorama world, or any art direction you pick). The skill interviews the user for the topic, the story beats/sections, and brand kit, then generates cohesive scenes + seamless camera clips with MuAPI and wires a portable, framework-agnostic scroll-scrub engine. Use when the user wants a "3D world" / "browse-through-the-industry" hero, a scroll cinematic, a diorama landing, or to turn a business into a scrollable world.

Funktioniert mit✓Claude Code~Codex CLI~Cursor
npx skills add DeepDiveTechLab/six

In Ihrer bevorzugten KI fragen

Öffnet einen neuen Chat, in dem dieser Agent-Skill bereits geladen ist.

Dokumentation

SIX — Scrolling Immersive Xploration

Produces a landing page where scroll drives a camera: it dives from outside a scene into its interior, then flies out and into the next scene, continuously, with no visible cuts. The visuals are AI-generated (MuAPI); the page just scrubs pre-rendered video by scroll position. This is the same technique behind Apple's scroll-through product pages — the camera genuinely moves, scroll only drives time.

What you generate: N scene stills → N "dive-in" camera clips → N-1 "connector" clips that join consecutive scenes seamlessly → a portable scrub engine that plays the whole chain as one flight.

The one rule that makes or breaks it: seams must be frame-identical. Read The seamless chain before generating any connector. Getting this wrong is the single most common failure and produces a visible "pop" between scenes.

Do not assume a frontend framework. The scrub engine in resources/scrub-engine.js is self-contained vanilla JS (it builds its own DOM + injects its own CSS into a container you give it, and exposes itself as mountSIX), so it drops into plain HTML, Next.js, Vue, a Python-served page, anything. The value of this skill is the MuAPI pipeline, the prompts, and the seam method — not the framework.

File map: SKILL.md (this file) → scripts/pipeline.md (copy-paste bash for the whole generation run) + scripts/knockout.py (background removal) → references/prompts.md (intake checklist + every prompt template) + references/muapi-model-roster.md (verified model/tier/endpoint mapping — read this if anything below needs auditing) → resources/scrub-engine.js + resources/index-template.html (drop-in engine + minimal page).


Step 0 — Bootstrap

  1. MuAPI CLI. If muapi is not on $PATH, install it per your environment's usual package flow, then run muapi auth configure (writes MUAPI_API_KEY/its config file — interactive, you cannot run the auth flow itself if it requires a browser). Confirm it works with muapi auth status. There is no documented CLI credits-balance command; if a batch stalls with a credits-flavoured error, check the muapi.ai dashboard directly.
  2. ffmpeg / ffprobe on $PATH (frame extraction + encoding).
  3. An image tool for background knockout if you want floating scenes: PIL (python3 -c "import PIL"), or cwebp/sips. Optional — see Step 3.
  4. jq on $PATH — used throughout the pipeline to parse --output-json responses and extract uploaded-asset URLs.
  5. Caveats: macOS ships bash 3.2 (no declare -A); don't use associative arrays in scripts. MuAPI generations take minutes — always run them detached (background) and poll (--async + muapi predict wait <request_id>), never a foreground blocking call. MuAPI's video/image endpoints take hosted URLs, not local file paths — this is the opposite of Higgsfield's convention, and the single most likely migration bug: every local reference image (a scene still, an extracted boundary frame) must go through muapi upload file <path> --output-json --jq '.url' first, and the returned URL is what gets passed to --start-image/--end-image/--image-urls. Exact flag names and resolution/quality options aren't uniformly documented across every model — verify with muapi video from-image --help / muapi image generate --help before batching, and see references/muapi-model-roster.md for what's already been cross-checked.

Step 1 — Interview the user

The subject is the user's to state — ask it as an open question in plain prose, never a fabricated multiple-choice. A made-up list of industries biases them and reads as you deciding their business for them; let them answer in their own words (their real business, a client's, or any idea). Reserve structured multiple-choice (AskUserQuestion in Claude Code; a plain either/or question elsewhere) for the genuinely enumerable, lower-stakes choices below — art direction and brand-kit approach — and even there, signal they can go their own way ("Other"). Ask only what you can't sensibly default. Cover:

  1. Subject (ask openly, not multiple-choice) — "What should this world be about? Your business, a client's, or any idea — a word or a sentence is fine." Capture the industry/product + a one-line pitch (e.g. "a bubble tea company, from leaf to last sip"), and a brand name if they have one; otherwise you'll propose one below.
  2. Brand kit — offer three paths, pick one:
    • Import from a URL: fetch the page yourself (web-fetch it) and read off the visible palette, name, and tone — no dedicated MuAPI brand-kit importer is documented, so this step is you doing the extraction, not a CLI call. Say so plainly if the user expects a one-command import.
    • The user hands you palette + name + tone directly.
    • You propose a palette + name and let them approve. Capture 4–6 named hex values, a display name, and a tone word or two.
  3. Art direction — default is "soft matte low-poly clay diorama, isometric, tilt-shift miniature, warm light." Offer alternatives (flat papercraft, glossy toy, claymation, neon night, photoreal architectural). Whatever is chosen becomes the shared style preamble reused verbatim in every scene prompt (this is what makes the world cohesive).
  4. The journey (sections) — the ordered scenes the camera flies through. Propose a set derived from the subject's own value chain and let the user edit. 5–7 works well. Boba example: farms → pearl kitchen → flagship shop → delivery → community plaza → the hero product. Each section needs: a short subject description (what's IN the diorama), an eyebrow, a headline, one line of body, and 0–3 tag pills. The last section is usually the hero product + the CTA.
  5. Mobile version (beta) — ALWAYS ask this; never silently generate both. Ask as a two-option choice (AskUserQuestion in Claude Code; a plain question elsewhere): "Want a mobile-optimized version too? Mobile support is in beta — the scroll-scrub mechanic is desktop-native; on phones you get lighter encodes and engine hardening, but portrait crops the 16:9 frame and low-end devices may still stutter." Options: "Desktop only" / "Desktop + mobile (beta)". The beta disclaimer must be stated to the user, not just implied. What the answer gates:
    • Yes → produce the -m.mp4 mobile encodes (Step 6) and wire clipMobile/connectorsMobile (Step 7); run the full mobile QA (Step 8). If any scene's focal subject sits off-centre, offer the 9:16 hero-variant escape hatch (extra MuAPI credits — say so).
    • No → skip the mobile encodes and wiring entirely. The engine's phone hardening (seek-coalescing, iOS priming, safe-area CSS) is always on regardless — that's not a "mobile version," it's just the page not breaking when a phone visits — so a desktop-only build still degrades gracefully.

Video model is not an interview question — default seedance-2-first-last-frame (tier global) silently. If the user names a preference, honor it only if it can frame-lock a seam — i.e. it accepts a start image and, for connectors, an end image too (references/muapi-model-roster.md has the verified roster). This skill only ships seamless output, so a model that can't frame-lock is declined with a one-line why, not substituted in — use a roster model instead.

Keep the scroll mechanic fixed (continuous fly-through) — that's the point of the skill. See references/prompts.md for the intake checklist and copy structure.


Step 2 — Generate the scene stills

One image per section, all sharing the same style preamble for cohesion. Default model gpt-image-2-text-to-image (crisp, great at isometric illustration; returns a solid/white background which is perfect for floating diorama "islands"). Use nano-banana-pro if the brief is character/cartoon-heavy, or google-imagen4-ultra for the photoreal architectural style (see references/muapi-model-roster.md for the full image-model table).

Prompt shape (full templates in references/prompts.md):

<STYLE PREAMBLE, identical every time>. On a plain solid <bg> background with a soft
contact shadow. <PALETTE hexes>. No text, no letters, no logos, centered, 3:2.
Subject: <what is in THIS diorama>.
  • Run all N concurrently, detached. Command per scene:
    muapi image generate --model gpt-image-2-text-to-image --prompt "$(cat scene_i.txt)" \
      --aspect-ratio 3:2 --async --output-json scene_i.job.json
    id=$(jq -r '.request_id' scene_i.job.json)
    muapi predict wait "$id" --json > scene_i.json
    
  • Result URL is under .result_url (or .output_url/.url — the exact key isn't uniformly documented; inspect one response and adjust) in the poll output. curl it down.
  • A generation may fail transiently — re-roll that one individually; don't restart the batch.
  • Review the stills before continuing. They must read as one cohesive world (same angle, palette, light). If one is off-style, regenerate it with muapi image generate again (a fresh call, same preamble) — don't reach for muapi image edit with a prior still as reference to "lock" style; edit mode fuses/clones that reference's actual room rather than just matching its style.

See scripts/pipeline.md for the exact batch script.


Step 3 — (Optional) Float the scenes

If you want the dioramas to float over an atmospheric background instead of sitting in a solid box, knock out the flat background to transparency with scripts/knockout.py (border-connected flood fill — preserves interior colour that matches the bg, e.g. cream walls). Then encode to webp. If you'd rather keep it simple, just make the page background the same colour as the scene background and skip this.

These stills double as video posters and lazy-load fallbacks, so keep them.


Step 4 — Camera architecture (pick one — this makes or breaks the feel)

How the camera moves between scenes is the single biggest quality lever. Two shapes; pick by aesthetic.

Video model — pick ONE for the whole chain

This skill only ships seamless output, so the only usable models are ones that can frame-lock a seam: every chained clip must accept --start-image, and connectors also need --end-image. That capability — not preference — is the selection rule. Full evidence trail in references/muapi-model-roster.md; condensed here:

Modeltierstart/end imageNotes
seedance-2-first-last-frame (default)global✓ / ✓--mode first-last — one image = a leg/dive (no end), two images = a connector. Same model covers both roles.
seedance-2-first-last-frame-fastglobal✓ / ✓Same model, fast-queue variant — the cheap previz tier. Run the whole chain here first to validate seams before spending on the full render.
seedance-2-vip-first-last-framevip✓ / ✓NSFW / content-filter fallback — VIP's filter band differs and documented to tolerate realistic human faces in references, unlike global/chinese. Use for one stubborn clip, not the default chain.

The chinese tier is not in the roster: first-last (two-image conditioning) isn't available on it, only single-image i2v — so it can't hold a seam. Skip anything whose media inputs are single-image-only for the same reason (it can only condition a generation, not continue a shot).

Rules:

  • One model+tier for all chained clips. Each renderer/tier has its own motion/color/grain character; mixing mid-chain keeps position continuity (frames still hand off) but the render-character shift reads as a subtle pop. The one sanctioned exception is the NSFW fallback for a single stubborn clip (Gotchas) — a slight character shift on one 5s connector beats a missing connector.
  • Default to seedance-2-first-last-frame on global; honor a user's stated preference only if the model qualifies (frame-locking). If it doesn't, say so and use a supported model — never ship a non-seamless build to satisfy a model request.
  • The pipeline scripts take the model/tier as $VMODEL/$VTIER with per-tier flags already cased out (scripts/pipeline.md).

A) Continuous forward take — RECOMMENDED for grounded / realistic / walkthrough

One camera that only ever glides forward, first scene through last, as a single take. Generate the legs sequentially: leg 0 from scene-0's still (glide forward into it); then each leg's --start-image = the previous leg's ACTUAL last frame (extract with ffmpeg, upload it), prompt "continue gliding smoothly FORWARD into [scene i], never pulling back" (or an expressive mid-leg move under the motion-handoff contract — see Camera grammar below), and no --end-image — an end-image of a wide establishing shot forces the camera to pull back, which is the #1 cause of stutter. Extract each leg's last frame to feed the next. Result: every seam is frame-identical and the camera never reverses. There are no connectors (skip Step 5) — the legs ARE the journey. Wire each leg as a section clip with connectors: [] and a small crossfade (~0.08). Even without an --end-image the legs still arrive at distinct rooms (the prompt steers the content). Cost: strictly sequential (can't parallelize) and slower; interiors are more likely to trip the content filter, so build in re-rolls (3 attempts/leg).

B) Dive-in + aerial connector — only for diorama / miniature / god's-eye worlds

A "dive into each scene" clip + a connector that pulls up and out and flies over to the next scene (Step 5). The pull-out reverses camera direction at every seam (forward dive → backward pull-out). In a miniature/diorama world that reads as an intentional "zoom out to the map, fly to the next island"; in a grounded first-person walkthrough it reads as a jarring rewind/stutter. Use B only for the map-like aesthetic. When in doubt, use A.

Camera grammar — the move should fit the concept (A is NOT "forward only")

"Forward only" is the seam rule, not the leg rule. The physics of the chain:

  • Position continuity at a seam comes from the frame handoff (next leg starts from the previous leg's actual last frame).
  • Velocity continuity at a seam means the camera must never reverse across a seam — that's the rewind stutter.
  • Inside a single leg the camera is free. One leg is one continuous render — there is no seam to break mid-leg, so orbits, crane-ups, lateral tracking, even a push-in that eases back out are all safe within the clip. Reversals are only fatal across seams.

So give each leg an expressive move chosen from the scene's own logic, under a motion handoff contract: every leg ends by settling into a slow, steady forward drift toward the next destination (final ~1 s), and every leg begins by continuing that same drift. Keep both clauses in the prompts verbatim (templates in references/prompts.md).

Pick the grammar from the concept:

Concept / toneMid-leg move
Product / luxury retailslow half-orbit around the hero object, then continue past it
Real estate / hospitalitysteadicam glide through doorways; gentle crane-up in atria
Industrial / process / logisticslow lateral track alongside the line, foreground parallax
Travel / outdoors / campusdrone-style rise-and-reveal, then a descending swoop
Food / craft / detail-drivenpush in close to the craft moment, ease back, carry on
Playful miniature (arch. B)dives + aerial hops — the connector IS the grammar

Honest costs: expressive mid-leg moves raise re-roll odds — the model can end a fancy move in a state that isn't a clean forward drift. Mitigations: keep the final-second settle clause verbatim; eyeball each leg's last frame before chaining the next (it should look like a frame from a gentle forward glide — if not, re-roll before wasting the next leg); budget ~1 extra re-roll per expressive leg. A plain forward glide stays the zero-risk default — use it for legs where the scene itself is the show.

Two related pacing knobs live in the engine (Step 7): per-section scroll (more scroll distance = longer dwell in that scene) and linger (the camera settles mid-scene exactly while the copy peaks, then picks up speed toward the seam). Prefer expressive motion in the clip and restraint in the scrub mapping — they compound.

And remember scroll is a scrubber: visitors can scroll up, so every move also plays in reverse. That's free and expected — no extra work — but it's another reason seam velocity must be consistent in both directions (a seam that reads fine forward reads as a stutter backward too if velocity flips).

For B, one camera flight per scene: starts high/outside, descends into the interior, structure opens. Model: the chain model you picked above (default seedance-2-first-last-frame), --start-image = the scene still (uploaded first).

  • Use the solid-background still (not the knocked-out transparent one) as the start image, so the video has a full frame.
  • Prompt: "Single continuous cinematic camera move, no cuts. Begin high and far looking at the whole from outside … descend and fly inside toward … the roof/walls gently open to reveal the interior. , smooth graceful slow motion. No text." plus a one-line Audio directive (Seedance 2.0 generates audio natively — an unstated direction produces random ambient sound; template in references/prompts.md).
  • Params: --tier global --aspect-ratio 16:9 --duration 8 --generate-audio false. No --resolution flag is documented for video — MuAPI's Seedance 2.0 renders in the 480p–720p band and there's no way to request 1080p; encode at whatever ffprobe reports (Step 6), never upscale.
  • Run concurrently, detached, then download each result. Re-roll individual failures. Keep the raw sources — you need their frames next.

Step 5 — Connectors (architecture B only)

Skip this whole step for architecture A — the forward take has no connectors; its legs already chain seamlessly. This step applies to B (diorama/miniature), and note the reversal caveat from Step 4.

The connector clips are what make the world feel connected instead of cut. A connector flies from the end of scene i out and into the start of scene i+1. Both of its endpoints must be the ACTUAL RENDERED FRAMES of the neighbouring clips — never the original diorama still.

Why: every MuAPI generation renders slightly differently. If a connector ends on a fresh render of "the kitchen diorama," but the next dive clip starts on its own different render of that same diorama, the two won't match and you get a pop at the seam. The fix is to hand off the exact pixels:

For each connector between dive_i and dive_{i+1}:
  start-image = the LAST frame extracted from dive_i's rendered video
  end-image   = the FIRST frame extracted from dive_{i+1}'s rendered video

Now every seam is frame-identical on both sides: dive_i.end == connector.start and connector.end == dive_{i+1}.start.

Extract the boundary frames from the rendered dives (not the stills):

ffmpeg -sseof -0.15 -i dive_i.mp4   -frames:v 1 -q:v 2 dive_i_last.png    # interior of i
ffmpeg -ss 0      -i dive_{i+1}.mp4 -frames:v 1 -q:v 2 dive_next_first.png # establishing of i+1

Both extracted PNGs must be uploaded before the call — muapi upload file <path> --output-json --jq '.url' — since MuAPI's video endpoints take hosted URLs, not local paths.

Generate the connector (--duration 5 is plenty). Connectors need --end-image, so the tier must support first-last — that means global or vip, never chinese:

s=$(muapi upload file dive_i_last.png --output-json --jq '.url')
e=$(muapi upload file dive_next_first.png --output-json --jq '.url')
muapi video from-image --model "$VMODEL" --tier "$VTIER" \
  --prompt "$(cat connector_i.txt)" \
  --start-image "$s" --end-image "$e" \
  --aspect-ratio 16:9 --duration 5 --generate-audio false \
  --async --output-json connector_i.job.json

Connector prompt: "Single continuous camera move, no cuts. Pull up and back out of , rise into the sky, glide across the connected miniature world, and arrive above <scene i+1>, beginning to descend toward it. Seamless flowing aerial transition.

Verwandte Skills