Communitygithub.com

VoxFlowStudio/skills

Use when the user wants AI-generated short-form video — knowledge cards (picstory / 小红书 / TikTok / Reels), narrated explainers, presentations, AI clips, or slides — covering picstory, present, slides, explain, and image generation. For article-to-card reels (Slice — 13 themes including paper-slide), use voxflow:slice. For shareable HTML/CSS card images or narrated card MP4 videos (`voxflow card render`) use voxflow:card.

What is skills?

skills is a Claude Code agent skill that use when the user wants AI-generated short-form video — knowledge cards (picstory / 小红书 / TikTok / Reels), narrated explainers, presentations, AI clips, or slides — covering picstory, present, slides, explain, and image generation. For article-to-card reels (Slice — 13 themes including paper-slide), use voxflow:slice. For shareable HTML/CSS card images or narrated card MP4 videos (`voxflow card render`) use voxflow:card.

Works with✓Claude Code~Codex CLI~Cursor✓Gemini CLI
npx skills add https://github.com/VoxFlowStudio/skills/tree/HEAD/skills/video

Ask in your favorite AI

Open a new chat with this agent skill pre-loaded.

Documentation

VoxFlow Video Skill

Generate short-form videos with AI: LLM writes the script, AI draws cards or scenes, TTS narrates, FFmpeg / Remotion renders the final MP4.

For article-to-card reels (Slice — 13 themes: paper-slide / editorial-mag / bold-poster / notion-card / brutalist / glass-dark / editorial-stencil / broadsheet / blueprint / daisy-pastel / showa-catalog / photo-feature / atmospheric), switch to voxflow:slice. For shareable HTML/CSS card image sets or narrated card-to-MP4 export, switch to voxflow:card.

Five entry points — pick by what the user wants:

CommandOutputUse when
picstoryVertical/landscape MP4 with hand-drawn cards or cinematic scenes"知识卡片视频", 小红书 / Twitter edu, sketchnote tutorials
present1080×1920 narrated card video, 5 visual schemesPitch decks, explainer reels, branded short-form
explainMP4 explainer with title / bullets / summary scenes"What is X?" tutorials, course intros
slidesSelf-contained HTML deck with embedded TTS audioProduct launches, talks, share-as-link
imageSingle PNGOne-off illustrations / thumbnails (Hunyuan TextToImage)

Prerequisites

  • npm install -g voxflow and voxflow login
  • ffmpeg installed (brew install ffmpeg / sudo apt install ffmpeg) — required for MP4 render
  • For present / explain (Remotion-backed): the local plugin install includes remotion-cards/. If present says "Remotion not ready", run npm install inside the bundled remotion-cards/ directory. For explain, you can skip local Remotion with --cloud.

🎴 picstory — knowledge-card video

LLM writes a structured script, AI draws one card per scene, TTS narrates, FFmpeg assembles. Best for 小红书 / Twitter edu / TikTok.

Quick start

# Default: Chinese, sketchnote style, portrait, 5 scenes
voxflow picstory --topic "AI Agent 入门指南"

# 2-scene quick test (no full video render)
voxflow picstory --topic "AI 入门" --scenes 2 --image-only

# English landscape video
voxflow picstory --topic "How React Hooks Work" --language en --ratio landscape --style photo

Output: picstory-<timestamp>.mp4 + .json (script).

Card-type styles (structured heading + key points)

--styleLookBest for
sketchnote (default)Colorful hand-drawn bullet journalKnowledge sharing, tutorials
neon_noirCyberpunk dark with neon glowTech, startup
minimal_3dSoft 3D clay on pastel gradients小红书 lifestyle
chalkboardWhite chalk on dark greenScience, academic

Scene-type styles (free-form illustration)

--styleLookBest for
photoCinematic / photo-realStorytelling, travel
manga_panelJapanese manga ink lineworkDrama, step-by-step guides
vintage_newspaper1940s broadsheetHistory, factual stories

Ratios

--ratioPixelsPlatform
portrait (default)1080×1920小红书, TikTok, Reels, 抖音
landscape1920×1080YouTube, B站
square1080×1080Instagram, Twitter

Script-model presets (--script-model)

PresetProviderStrength
omittedserver configBalanced default (gpt-4o-mini)
gemini-flashOpenRouterMultilingual, good Chinese
deepseekDeepSeekCheapest, excellent Chinese
hunyuan腾讯混元Chinese-native
moonshotMoonshotChinese long context

Server enforces an allowlist — only the preset names above work; arbitrary model IDs are rejected.

Image quality (--quality)

Image generation is the only meaningful cost (~$0.005-0.08 per image). LLM script (~2K tokens) is negligible.

--qualityProviderStrength5-image cost
fast (default)OpenRouter Gemini FlashCheapest, balanced~$0.025
hdOpenRouter Gemini ProHigher detailmid
ultraOpenRouter gpt-5.4-image-2Best overall, ~16× cost~$0.40
fast-aibermAiberm Gemini FlashCheap Aiberm tierlow
hd-aibermAiberm Gemini ProStrongest Chinese text rendering — best for 小红书 cards with Chinese headersmid

Use fast for iteration; hd-aiberm when cards must contain accurate Chinese characters; ultra for hero exports.

Full options

FlagDefaultDescription
--topic <text>required (or --text)Story topic
--text <content>—Paste full article instead of a topic
--style <name>sketchnoteSee styles above
--ratio <name>portraitportrait | landscape | square
--language <code>zhzh | en | ja | ko | ...
--scenes <n>52–10. Use 2 for quick tests.
--script-model <preset>server defaultSee presets above
--quality <tier>fastSee table above
--voice <id>defaultTTS voice from voxflow voices
--speed <n>1.0TTS speed 0.5–2.0
--bgm <file>—Background music (mp3/wav) mixed under narration
--bgm-volume <n>0.1BGM volume 0–1
--fade <n>0.4Per-scene fade in/out seconds (0 to disable)
--image-onlyfalseSave images + audio without final video render
--output-dir <dir>—Directory for all outputs
--output <path>autoFinal MP4 path

Quota cost

OperationQuota
LLM script100
TTS / scene50
Image / scene500
2-scene test~1,200
5-scene full~2,850

Free tier (10K/month) ≈ 3 full picstory videos.

Pipeline

Topic / Text
  ├─[1] LLM script → { title, scenes: [{ heading, keyPoints, narration }] }
  ├─[2] TTS per scene → PCM audio
  ├─[3] Image per scene (parallel) → JPEG via /api/image/generate
  └─[4] FFmpeg: image + WAV → Ken Burns zoompan MP4 → concat → +BGM → final.mp4

Troubleshooting

ProblemFix
ffmpeg not foundbrew install ffmpeg (mac) / sudo apt install ffmpeg
Image generation failedRetry; if stuck on hd, drop to --quality fast
Render too slowKen Burns is CPU-heavy. Use --scenes 2 --image-only to validate content first.
model_not_allowedUse a --script-model preset name, not a raw model ID
Chinese text mangled in cardsSwitch to --quality hd-aiberm

📑 present — narrated card video (Remotion local)

Text or URL → LLM cards → TTS → Remotion render. 1080×1920, 5 visual schemes.

voxflow present --text "Claude Code 是一个 AI 编程工具" --style aurora
voxflow present --url https://example.com/article --style noir
voxflow present --text "2025 AI 芯片格局" --web-search --style neon
voxflow present --cards pre-generated.json --no-audio
FlagDefaultNotes
--text / --url / --cardsone requiredinput source
--styleauroranoir | neon | editorial | aurora | brutalist
--voice <id>v-female-R2s4N9qJTTS voice
--speed <n>1.00.5–2.0
--no-audiofalseSilent video only
--web-searchfalseAugment LLM with up-to-date web facts
--output <path>./present-<ts>.mp4.mp4 or .wav

If present reports "Remotion not ready", run npm install inside the bundled remotion-cards/ directory to set up the local renderer.


🧠 explain — AI explainer video

Title / bullets / summary scene flow. Best for "What is X?" tutorials.

voxflow explain --topic "What is React?"
voxflow explain --topic demo --output demo.mp4         # built-in demo (no API call)
voxflow explain --topic "区块链入门" --style chalkboard --voice v-male-Bk7vD3xP
voxflow explain --topic "Machine Learning" --audio-only
voxflow explain --topic "AI Agent 入门" --cloud         # render on server
FlagDefaultNotes
--topic <text>requiredUse demo for built-in demo script
--stylemodernmodern | playful | corporate | chalkboard
--languageenen | zh | ja | ko | ...
--voice <id>v-female-R2s4N9qJ
--scenes <n>53–12
--audio-onlyfalseSkip render, output WAV only
--cloudfalseUse cloud Remotion instead of local
--output <path>./explain-<ts>.mp4

📊 slides — HTML presentation with TTS

Generates a self-contained HTML deck with embedded base64 audio per slide. Open in any browser, no server needed.

voxflow slides "AI in Healthcare"
voxflow slides "Q4 Revenue Report" --template report --theme paper
voxflow slides "React Tutorial" --template tutorial --model balanced
voxflow slides "Startup Pitch" --template pitch --theme ocean --no-audio
FlagDefaultNotes
--text <text> (or positional)requiredTopic
--templatefreeproduct | report | tutorial | pitch | free
--thememidnightmidnight | paper | ember | forest | ocean
--modelswiftswift | balanced | pro | creative
--voice <id>v-female-R2s4N9qJ
--no-audiofalseSkip TTS, slides only
--output <path>./slides-<ts>.html

10 layouts: hero, title-bullets, two-column, three-cards, image-left, image-right, quote, timeline, stats, section. Templates auto-pick a layout sequence.


🖼 image — single Hunyuan illustration

Synchronous text → image (PNG). Useful for thumbnails or one-off art.

voxflow image "a sleeping cat in a sunlit window" --resolution 1024:1024 -o cat.png

Resolutions: 768:768, 768:1024, 1024:768, 1024:1024, 720:1280, 1280:720, 768:1280, 1280:768, 1080:1920, 1920:1080.

Prompt max 1000 chars. Output: local PNG + COS URL.


Pick-the-right-tool checklist

"小红书风格知识卡片"           → picstory
"AI 短视频 + caption"          → picstory --style sketchnote
"explainer / What is X?"        → explain
"branded short with my text"    → present
"already have cards/script"     → present --cards
"shareable HTML deck w/ audio"  → slides
"single illustration"           → image

Rules

  1. Search voices with voxflow voices before passing --voice. Never guess IDs.
  2. Check quota before video calls (voxflow status): picstory ≈ 3K, present/explain ≈ 500–2K.
  3. Test cheap first: picstory --scenes 2 --image-only validates the script before paying for full render.
  4. After render finishes, auto-play: open output.mp4 (macOS).

Individual skills in this repo

This repo contains 3 individual skills — each has its own dedicated page.

VoxFlowStudio/skills

Use when the user wants to turn text content into a set of polished, shareable visual CARD IMAGES or narrated card VIDEOS — knowledge cards, quote cards, 小红书图文, carousel cards, poster cards — rendered as HTML/CSS and exported via Playwright at ratios like 1:1 / 3:4 / 9:16; optionally produces a narrated MP4 video from those cards via `voxflow card render` (per-card TTS + FFmpeg static-image clips with optional subtitle bar / intro+outro cards / BGM mix). Triggers: card / 卡片 / 知识卡 / 文字卡片 / 金句卡 / 图文卡片 / 卡片生成 / make cards / card video / 卡片视频. For article → Slice-themed card VIDEO use voxflow:slice; for short videos / AI clips use voxflow:video; for podcasts use voxflow:podcast.

VoxFlowStudio/skills

Use when the user wants to read text aloud (TTS), search VoxFlow voices, sample AI stories, or set up VoxFlow install/auth/quota — the entry-point voice toolkit. For podcasts use voxflow:podcast; for short videos / AI clips use voxflow:video; for article-to-card reels (Slice) use voxflow:slice; for shareable card images or narrated card videos use voxflow:card; for transcription / dubbing / subtitle translation use voxflow:transcribe.

VoxFlowStudio/skills

Use when the user wants to turn a long article / note / report into a vertical 1080×1920 card video — VoxFlow Slice. 13 themes — paper-slide (纸面), editorial-mag (编辑刊), bold-poster (大字海报), notion-card (Notion 卡), brutalist (粗野), glass-dark (玻璃夜), editorial-stencil (编辑·海报), broadsheet (财经刊), blueprint (蓝晒图), daisy-pastel (雏菊), showa-catalog (昭和目录), photo-feature (摄影刊), atmospheric (深夜刊). Triggers — Slice / slice video / 切片视频 / 文章转视频 / 知识卡片视频 / 抖音知识号 / 小红书图文转视频 / 知乎长文转视频 / 公众号转视频 / PaperSlide / paperslide / paper-slide (legacy name).

Related Skills