Communitygithub.com

NeverSight/learn-skills.dev

Grok Build ONLY. Turn a 2D character still into smooth animation sprites via image_gen/image_edit base → image_to_video (6s/10s run-in-place) → ffmpeg frames → magenta chroma-key → dense sampled sprites (strip/grid/GIF). Use when the user wants video-to-sprite, motion capture from generated video, smoother run/walk cycles from dense frames, or runs /video2dsprite. Do NOT use on Codex/Claude — only Grok Build has image_to_video. Prefer generate2dsprite for crisp pixel sheets without video.

What is learn-skills.dev?

learn-skills.dev is a Claude Code agent skill that grok Build ONLY. Turn a 2D character still into smooth animation sprites via image_gen/image_edit base → image_to_video (6s/10s run-in-place) → ffmpeg frames → magenta chroma-key → dense sampled sprites (strip/grid/GIF). Use when the user wants video-to-sprite, motion capture from generated video, smoother run/walk cycles from dense frames, or runs /video2dsprite. Do NOT use on Codex/Claude — only Grok Build has image_to_video. Prefer generate2dsprite for crisp pixel sheets without video.

Works with✓Claude Code✓Codex CLI~Cursor
npx skills add https://github.com/NeverSight/learn-skills.dev/tree/HEAD/data/skills-md/0x0funky/agent-sprite-forge/video2dsprite

Ask in your favorite AI

Open a new chat with this agent skill pre-loaded.

Documentation

Video2dsprite (Grok Build only)

Convert a base 2D character image into dense animation sprites using Grok Build's native video tools.

base still → image_to_video (in-place motion) → extract frames → chroma key → sample/normalize → strip / grid / GIF

Platform gate (read first)

RuntimeSupported?
Grok Build (xAI)Yes — requires image_gen / image_edit + image_to_video (or reference_to_video)
Codex / Claude / other agentsNo — they lack Grok video tools. Tell the user this skill is Grok Build only and offer $generate2dsprite instead

If image_to_video is missing from the tool list, stop and explain. Do not fake motion with code-drawn frames.

This skill is an optional denser-motion path. It does not replace $generate2dsprite:

Use $generate2dsprite when…Use $video2dsprite when…
Crisp pixel sheets, fixed grids, identity-critical heroesUser wants denser intermediate poses / smoother feeling loops
Attack/cast body sheets, prop packs, engine atlasesExperimenting with video-sourced run/walk/idle motion
Production default for most game spritesUser explicitly asks for video → frames → sprites

Video softens pixels, drifts identity, and leaves chroma fringes. Always QC; for production heroes, prefer $generate2dsprite unless the user wants the video look.

Parameters

Infer from the user request:

  • subject: character / creature description, or path to existing still
  • action: run | walk | idle | attack | custom motion phrase
  • view: usually side (side-scroller). topdown is harder — warn and keep camera locked
  • duration: 6 (default) or 10 seconds
  • frame_counts: which denser sets to export, default 8,16,24,48
  • cell_size: output sprite cell, default 128
  • anchor: feet (default for side locomotion) | center
  • bg: solid #FF00FF (required for chroma)
  • name: output slug
  • out_dir: working folder (default ./sprites/video2dsprite/<name>/ or project-relative)

Agent rules

  1. Grok-only. Refuse on non-Grok runtimes with a short explanation + $generate2dsprite alternative.
  2. Still → video, never text-to-video alone. Stage frame 1 as a clean still (image_gen or image_edit from a reference). Then call image_to_video.
  3. In-place motion. Prompt for run/walk in place facing a fixed direction. No camera pan, no background scroll, no scene change. Subject stays roughly centered.
  4. Solid magenta background on the base and preserved in the video prompt (#FF00FF / pure magenta). Required for flood-fill chroma.
  5. Do not invent art with PIL/Canvas. Base art comes from image_gen / image_edit or a user/local still. Scripts only postprocess.
  6. Do not put experimental outputs into the game unless the user asks to integrate.
  7. Prefer one locomotion cycle for game use. Dense sample across a full 6s multi-cycle clip is fine for previews; for engine sheets, optionally re-sample a single cycle (12–16 frames) after visual QC.
  8. Report absolute paths of video, cleaned frames, strips, and preview GIFs when done.

Workflow

1. Plan

Pick the smallest useful run:

  • Side-view run/walk loop → this skill
  • Multi-action hero kit → still use $generate2dsprite per action; only use video for locomotion if requested
  • FX / projectile / prop packs → $generate2dsprite, not video

Create:

<out_dir>/
  base/
  video/
  frames-raw/
  frames-clean/
  sprite/          # default 8-frame set + denser x16/x24/x48
  prompt-used.txt
  pipeline-meta.json
  README.txt

2. Build the base still

Options:

  • A. Existing sprite: open with image tools / read image, composite onto solid #FF00FF if needed
  • B. New character: image_gen with solid magenta background, full body, side view, centered
  • C. Match reference: image_edit from user reference onto magenta, preserve identity

Base requirements:

  • Full body visible, generous magenta margin
  • Side view for run/walk (profile or 3/4 side), feet near bottom third
  • Same art style as the rest of the project when a reference exists
  • No text, UI, watermark, or second character

Save as <out_dir>/base/<name>-base.png.

Write the exact image prompt into prompt-used.txt.

3. Animate with image_to_video

Call Grok image_to_video:

  • image: path to the base still
  • duration: 6 (default) or 10
  • resolution_name: 480p unless user asks 720p
  • prompt: one short present-tense shot (see references/prompt-rules.md)

Mandatory motion constraints in the prompt:

  • Subject runs/walks in place (treadmill style)
  • Camera locked — no pan, zoom, or orbit
  • Background stays flat solid magenta
  • Identity, costume, palette stable for the whole shot
  • Single continuous action only

Copy the returned video to <out_dir>/video/<name>-<duration>s.mp4.

If video tools are unavailable, stop (platform gate).

4. Extract + chroma + sample (local script)

Run the processor (ffmpeg + Pillow + numpy):

python skills/video2dsprite/scripts/video2dsprite.py process \
  --video <out_dir>/video/<name>-6s.mp4 \
  --out-dir <out_dir> \
  --name <name> \
  --frame-counts 8,16,24,48 \
  --cell-size 128 \
  --body-height 100 \
  --foot-y 118 \
  --fps 0

Notes:

  • --fps 0 = extract every decoded frame (use source fps)
  • Magenta flood-fill from corners + despill
  • Even sampling for each count in --frame-counts
  • Feet-normalized cells, horizontal strip, grid, loop GIF per count

Optional: only re-sample denser sets from existing cleaned frames:

python skills/video2dsprite/scripts/video2dsprite.py sample \
  --clean-dir <out_dir>/frames-clean \
  --out-dir <out_dir> \
  --frame-counts 16,24,48 \
  --cell-size 128

5. QC

Visually check:

  • Preview GIF loops without huge pops
  • Magenta gone (no solid pink blocks); fringe acceptable or re-key
  • Feet stay on a stable baseline (no hop from bad crop)
  • Identity roughly stable (face/clothes not morphing every frame)
  • Action is in-place (not sliding out of frame)
  • For game use: pick one count (often 16 or 24) or cut one true cycle

If identity drifts hard or pixels are too soft, fall back to $generate2dsprite for production sheets and keep the video set as motion reference only.

6. Deliver

Report paths only (unless user asked to wire into a game):

  • Video: video/*.mp4
  • Dense sprites: sprite/x16|x24|x48/
  • Strips / grids / GIFs: sprite/run-strip-N.png, run-grid-N.png, run-preview-N.gif
  • Meta: pipeline-meta.json

Do not modify game code unless requested.

Defaults

  • Duration: 6s
  • Action: side run in place, facing right
  • Export counts: 8, 16, 24, 48
  • Cell: 128², body height ~100, feet at y≈118
  • Background: #FF00FF
  • Prefer image_to_video over reference_to_video (compose multi-ref with image_edit first if needed)

Tradeoffs (tell the user once)

Pros: denser intermediates → often feels smoother than 4–8 discrete gen poses.
Cons: softer pixels, identity drift, chroma fringe, multi-cycle 6s clips are not a single perfect loop, heavier assets.
Rule of thumb: 8→16→24 usually gains smoothness; 48 is often diminishing returns; 145 raw frames are for sampling, not all for runtime.

Resources

Relationship to other skills

  • $generate2dsprite — primary sheet pipeline (Codex + Grok when image gen exists)
  • $generate2dmap — maps; not used here
  • $video2dsprite — Grok Build exclusive motion densification path

Individual skills in this repo

This repo contains 20 individual skills — each has its own dedicated page.

NeverSight/learn-skills.dev

Create AI avatar and talking head videos via inference.sh CLI. Recommended: P-Video-Avatar (fastest, cheapest, built-in TTS). Also: OmniHuman, Fabric, PixVerse. Audio: Inworld TTS-2 (100+ languages, emotion steering for characters), ElevenLabs, Kokoro. Capabilities: audio-driven avatars, text-to-avatar, lipsync videos, talking head generation, virtual presenters, UGC content. Use for: AI presenters, explainer videos, virtual influencers, dubbing, marketing videos, UGC ads, gaming avatars, NPC dialogue. Triggers: ai avatar, talking head, lipsync, avatar video, virtual presenter, ai spokesperson, audio driven video, heygen alternative, synthesia alternative, talking avatar, lip sync, video avatar, ai presenter, digital human, ugc, ugc video, ugc ad, avatar ugc

NeverSight/learn-skills.dev

Create AI marketing videos for ads, promos, product launches, and brand content. Models: Veo, Seedance, Wan, FLUX for visuals, Kokoro for voiceover. Types: product demos, testimonials, explainers, social ads, brand videos. Use for: Facebook ads, YouTube ads, product launches, brand awareness. Triggers: marketing video, ad video, promo video, commercial, brand video, product video, explainer video, ad creative, video ad, facebook ad video, youtube ad, instagram ad, tiktok ad, promotional video, launch video

NeverSight/learn-skills.dev

Generate AI videos with Google Veo, Seedance 2.0, HappyHorse, Wan, Grok and 40+ models via inference.sh CLI. Models: Veo 3.1, Veo 3, Seedance 2.0, HappyHorse 1.0, Wan 2.5, Grok Imagine Video, OmniHuman, Fabric, HunyuanVideo. Capabilities: text-to-video, image-to-video, reference-to-video, video editing, lipsync, avatar animation, video upscaling, foley sound. Use for: social media videos, marketing content, explainer videos, product demos, AI avatars. Triggers: video generation, ai video, text to video, image to video, veo, animate image, video from image, ai animation, video generator, generate video, t2v, i2v, ai video maker, create video with ai, runway alternative, pika alternative, sora alternative, kling alternative, seedance, happyhorse

NeverSight/learn-skills.dev

ElevenLabs automatic dubbing - translate and dub audio/video into 29 languages while preserving speaker voice via inference.sh CLI. Capabilities: auto speaker detection, voice-preserving translation, video dubbing, audio localization. Use for: content localization, video translation, multilingual content, international distribution. Triggers: dubbing, dub video, translate audio, video translation, audio translation, localize content, elevenlabs dubbing, eleven labs dub, multilingual dub, voice translation, auto dub, language dub, content localization

NeverSight/learn-skills.dev

Explainer video production guide: scripting, voiceover, visuals, and assembly. Covers script formulas, pacing rules, scene planning, and multi-tool pipelines. Use for: product demos, how-it-works videos, onboarding videos, social explainers. Triggers: explainer video, how to make explainer, product video, demo video, video production, video script, animated explainer, product demo video, tutorial video, onboarding video, walkthrough video, video pipeline

NeverSight/learn-skills.dev

Still-to-video conversion guide: model selection, motion prompting, and camera movement. Covers Wan 2.5 i2v, Seedance, Fabric, Grok Video with when to use each. Use for: animating images, creating video from stills, adding motion, product animations. Triggers: image to video, i2v, animate image, still to video, add motion to image, image animation, photo to video, animate still, wan i2v, image2video, bring image to life, animate photo, motion from image

NeverSight/learn-skills.dev

Generate talking head avatar videos with Pruna P-Video-Avatar via inference.sh CLI. Turn a portrait image into a realistic speaking video with built-in TTS. 18x faster and 6x cheaper than competitors. Models: P-Video-Avatar, P-Image (for portrait generation). Capabilities: text-to-avatar, audio-driven avatars, 30 voices, 10 languages, 720p/1080p, built-in TTS, dynamic backgrounds, full-body control. Use for: AI presenters, product demos, explainer videos, virtual influencers, marketing, education, multilingual content, UGC, gaming avatars. Triggers: avatar video, talking head, ai avatar, p-video-avatar, pruna avatar, video avatar, ai presenter, digital human, virtual presenter, lipsync, talking avatar, ai spokesperson, heygen alternative, synthesia alternative, veed alternative, fabric alternative, omnihuman alternative

NeverSight/learn-skills.dev

Generate videos with Pruna P-Video and WAN models via inference.sh CLI. Models: P-Video, WAN-T2V, WAN-I2V. Capabilities: text-to-video, image-to-video, audio support, 720p/1080p, fast inference. Pruna optimizes models for speed without quality loss. Triggers: pruna video, p-video, pruna ai video, fast video generation, optimized video, wan t2v, wan i2v, economic video generation, cheap video generation, pruna text to video, pruna image to video

NeverSight/learn-skills.dev

Render videos from React/Remotion component code via inference.sh. Pass TSX code, get MP4. Supports all Remotion APIs: useCurrentFrame, useVideoConfig, spring, interpolate, AbsoluteFill, Sequence. Configurable resolution, FPS, duration, codec. Use for: programmatic video generation, animated graphics, motion design, data-driven videos, React animations to video. Triggers: remotion, render video from code, tsx to video, react video, programmatic video, remotion render, code to video, animated video, motion graphics code, react animation video

NeverSight/learn-skills.dev

Video ad creation with exact platform-specific specs for TikTok, Instagram, YouTube, Facebook, LinkedIn. Covers dimensions, duration limits, AIDA framework, and caption requirements. Use for: video ads, social media ads, paid media creative, video marketing, ad production. Triggers: video ad, social media ad, tiktok ad, instagram ad, youtube ad, facebook ad, linkedin ad, video creative, ad specs, paid media, video marketing, ad production, reels ad, stories ad, pre roll, bumper ad

NeverSight/learn-skills.dev

Best practices and techniques for writing effective AI video generation prompts. Covers: Veo, Seedance, Wan, Grok, Kling, Runway, Pika, Sora prompting strategies. Learn: shot types, camera movements, lighting, pacing, style keywords, negative prompts. Use for: improving video quality, getting consistent results, professional video prompts. Triggers: video prompt, how to prompt video, veo prompts, video generation tips, better ai video, video prompt engineering, video prompt guide, video prompt template, ai video tips, video prompt best practices, video prompt examples, cinematography prompts

NeverSight/learn-skills.dev

YouTube thumbnail design with specific dimensions, contrast rules, and mobile preview optimization. Covers safe zones, text placement, face expression psychology, and A/B testing. Use for: YouTube thumbnails, video cover images, click-through optimization. Triggers: youtube thumbnail, thumbnail design, video thumbnail, click through rate, ctr optimization, youtube cover, video cover image, thumbnail maker, thumbnail tips, youtube design, video preview image

NeverSight/learn-skills.dev

Landing page conversion optimization with layout rules, hero section design, and CTA psychology. Covers above-the-fold formula, social proof placement, mobile design, and F-pattern reading. Use for: startup landing pages, product pages, SaaS marketing, conversion optimization. Triggers: landing page, hero section, above the fold, conversion optimization, landing page design, cta button, hero image, landing page layout, saas landing page, product page design, conversion rate, landing page best practices

NeverSight/learn-skills.dev

Configure and use the hosted YouTube Data MCP end-to-end with minimal user input. Use when users want the agent to verify Node.js and `npx`, configure MCP server config (Windows/macOS, Cursor/Codex/OpenClaw/OpenCode), request API key at setup time, run post-install capability discovery (`tools/list` and `get_patch_notes`), and then strongly recommend helper skill and Python setup for full local document and spreadsheet workflows.

NeverSight/learn-skills.dev

Creates 120fps GPU-accelerated animations with Motion.dev (Framer Motion successor) for React, Next.js, Svelte, and Astro projects. Use when user requests animation, motion, scroll effects, parallax, hero animations, gestures, drag interactions, spring physics, whileHover effects, whileInView animations, animated UI, micro-interactions, page transitions, or layout animations. Generates production TypeScript/JSX code with accessibility (prefers-reduced-motion) and performance validation (≥60fps). Supports entrance animations, gesture interactions (hover/tap/drag), scroll-based reveals, and layout transitions using spring physics and natural timing. Do NOT use for CSS-only transitions (use native CSS), static sites without JavaScript, Vue animations (use motion-v variant instead), or SVG/Canvas complex animations (GSAP better suited).

NeverSight/learn-skills.dev

Static artifact craft skill for self-contained HTML/CSS/JS documents: docs, sheets, dashboards, explainers, slides, tools, and landing pages. Use when the user asks for a durable, openable, shareable web deliverable they'll keep or hand off — a report, a dashboard, a slide deck, a data table, a page. Local folder first, temporary public link via tunnel (localhost.run), optional durable publish to Surge, GitHub Pages, or Cloudflare. Not for quick look renders, inline snippets, or throwaway scratch. Not for SPA frameworks, backend APIs, database apps, or production product UI.

NeverSight/learn-skills.dev

FFmpeg commands for video/audio conversion, trimming, compression, and processing. Use when user mentions "ffmpeg", "convert video", "compress video", "extract audio", "trim video", "gif from video", "video codec", "transcode", "screen recording", "merge videos", "video to mp4", "reduce file size", or any media processing task.

NeverSight/learn-skills.dev

Vim keybindings, motions, text objects, and operators for efficient text editing. Use when user asks about "vim commands", "vim motions", "text objects", "vim keybindings", "vim cheat sheet", "learn vim", "vim in VS Code", or any Vim editing tasks.

NeverSight/learn-skills.dev

Download YouTube videos, extract/proofread/translate subtitles, and render them onto video. Use whenever the user asks to download a YouTube video with subtitles, translate video subtitles to Chinese, add subtitles/burn subtitles to a video, or do ASR transcription on video audio. Covers the full pipeline: video download → subtitle extraction → proofreading → translation → SRT generation → subtitle rendering. Also use for "下载视频加字幕", "视频翻译字幕", "把字幕烧录到视频中". For pure subtitle file creation without rendering (just SRT output), this skill handles that too — stop before Phase 5.

NeverSight/learn-skills.dev

Use when generating or modifying Remotion video code, creating demo videos, or working with the demo-video/ directory

Related Skills