Communitygithub.com

NeverSight/learn-skills.dev

Generate AI videos with Seedance (ByteDance) via AceDataCloud API. Use when creating videos from text prompts, animating images into motion videos, or driving Seedance 2.0 multimodal generation with real-person / character image references, reference audio, and reference video. Supports multiple models with configurable resolution (up to 4k), aspect ratio, duration, and optional audio generation.

learn-skills.dev 是什麼?

learn-skills.dev is a Claude Code agent skill that generate AI videos with Seedance (ByteDance) via AceDataCloud API. Use when creating videos from text prompts, animating images into motion videos, or driving Seedance 2.0 multimodal generation with real-person / character image references, reference audio, and reference video. Supports multiple models with configurable resolution (up to 4k), aspect ratio, duration, and optional audio generation.

相容平台~Claude Code~Codex CLI~Cursor
npx skills add https://github.com/NeverSight/learn-skills.dev/tree/HEAD/data/skills-md/acedatacloud/skills/seedance-video

在你喜歡的 AI 中提問

開啟一個已預先載入此 Agent Skill 的新對話。

說明文件

Seedance Video Generation

Generate AI dance and motion videos through AceDataCloud's Seedance (ByteDance) API.

Setup: See authentication for token setup.

Quick Start

curl -X POST https://api.acedata.cloud/seedance/videos \
  -H "Authorization: Bearer $ACEDATACLOUD_API_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{"model": "doubao-seedance-2-0-260128", "content": [{"type": "text", "text": "a dancer performing contemporary ballet in a misty forest"}], "callback_url": "https://api.acedata.cloud/health"}'

Async: See async task polling. Poll via POST /seedance/tasks with {"id": "..."}. This returns a task ID immediately. Poll for the result:

curl -X POST https://api.acedata.cloud/seedance/tasks \
  -H "Authorization: Bearer $ACEDATACLOUD_API_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{"id": "<task_id from above>"}'

Models

Seedance 2.0 (current generation — multimodal reference)

The 2.0 series adds multimodal reference inputs: real-person / character image references, reference audio, and reference video (see the workflows below).

ModelBest ForMax resolution
doubao-seedance-2-0-260128Highest quality, real-person/character reference, 4k output4k
doubao-seedance-2-0-fast-260128Faster 2.0 generation720p
doubao-seedance-2-0-mini-260615Lightweight / most cost-effective 2.0720p

Seedance 1.x

ModelTypeBest For
doubao-seedance-1-5-pro-251215Text+Image-to-Video1.5 flagship, audio support
doubao-seedance-1-0-pro-250528Text+Image-to-VideoGeneral-purpose, reliable quality
doubao-seedance-1-0-pro-fast-251015Text+Image-to-VideoFaster Pro generation
doubao-seedance-1-0-lite-t2v-250428Text-to-Video onlyLightweight text-to-video
doubao-seedance-1-0-lite-i2v-250428Image-to-Video onlyLightweight image-to-video

Workflows

1. Text-to-Video

Pass a text content item in the content array.

POST /seedance/videos
{
  "model": "doubao-seedance-1-0-pro-250528",
  "content": [
    {"type": "text", "text": "a street dancer doing breakdancing moves in an urban setting"}
  ],
  "resolution": "1080p",
  "ratio": "16:9",
  "duration": 5
}

2. Image-to-Video

Include an image content item (with an optional role) alongside the text.

POST /seedance/videos
{
  "model": "doubao-seedance-1-5-pro-251215",
  "content": [
    {"type": "text", "text": "the person starts dancing gracefully"},
    {
      "type": "image_url",
      "role": "first_frame",
      "image_url": {"url": "https://example.com/dancer.jpg"}
    }
  ],
  "resolution": "720p",
  "duration": 5
}

Image roles:

  • first_frame — image is used as the opening frame
  • last_frame — image is used as the closing frame
  • reference_image — image is used as a style / subject / real-person reference (Seedance 2.0 keeps the referenced person or character consistent)

Reference media (Seedance 2.0 only):

  • audio_url — reference audio for voice timbre / background music (no role)
  • video_url — reference video for subject, camera movement, motion or overall style (no role)

3. First-frame + Last-frame

Provide both a start and end frame image:

POST /seedance/videos
{
  "model": "doubao-seedance-2-0-260128",
  "content": [
    {"type": "text", "text": "smooth transition between two scenes"},
    {"type": "image_url", "role": "first_frame", "image_url": {"url": "https://example.com/start.jpg"}},
    {"type": "image_url", "role": "last_frame", "image_url": {"url": "https://example.com/end.jpg"}}
  ]
}

4. Real-person / character reference (Seedance 2.0)

Seedance 2.0 models (doubao-seedance-2-0-260128, doubao-seedance-2-0-fast-260128, doubao-seedance-2-0-mini-260615) can keep a specific person or character consistent across a brand-new scene. Pass one or more photos as image_url items with role: "reference_image" — the model preserves that subject's appearance. Up to 9 reference images are accepted.

POST /seedance/videos
{
  "model": "doubao-seedance-2-0-260128",
  "content": [
    {"type": "text", "text": "the same person walking through a neon-lit night market, cinematic"},
    {"type": "image_url", "role": "reference_image", "image_url": {"url": "https://example.com/person.jpg"}}
  ],
  "resolution": "1080p",
  "duration": 8
}

5. Reference audio / video (Seedance 2.0)

2.0 models also accept reference audio (voice timbre, background music) and reference video (subject content, camera movement, motion, overall style). Add audio_url and/or video_url content items. Limits: up to 3 audio and 3 video references per request.

POST /seedance/videos
{
  "model": "doubao-seedance-2-0-260128",
  "content": [
    {"type": "text", "text": "a singer performing on stage, matching the reference voice and motion"},
    {"type": "image_url", "role": "reference_image", "image_url": {"url": "https://example.com/person.jpg"}},
    {"type": "audio_url", "audio_url": {"url": "https://example.com/voice.mp3"}},
    {"type": "video_url", "video_url": {"url": "https://example.com/motion.mp4"}}
  ],
  "generate_audio": true
}

Parameters

ParameterValuesDescription
modelsee Models tableModel to use (required)
contentarrayInput items: text, image_url, audio_url (2.0), video_url (2.0) (required)
resolution"480p", "720p", "1080p", "4k"Output resolution. 4k is doubao-seedance-2-0-260128 (standard) only; 2-0-fast / 2-0-mini max out at 720p (default: 720p for pro/2.0, 480p for lite)
ratio"16:9", "4:3", "1:1", "3:4", "9:16", "21:9", "adaptive"Aspect ratio (default: 16:9)
duration2 – 15Duration in seconds (Seedance 2.0 supports 4–15)
frames29–361 (must satisfy 25+4n)Frame count — mutually exclusive with duration
seed-1 to 4294967295Seed for reproducible results (-1 = random)
generate_audiotrue / falseGenerate audio (supported by doubao-seedance-1-5-pro-251215 and the doubao-seedance-2-0 series; other models ignore it)
camerafixedtrue / falseFix the camera position during generation
watermarktrue / falseAdd a watermark to the generated video
return_last_frametrue / falseReturn the last frame of the generated video
service_tier"default", "flex"Processing tier (default: default)
execution_expires_afternumberTask timeout threshold in seconds

Inline Parameter Syntax

You can also embed generation parameters directly in the text prompt using the --param value syntax:

A kitten yawning at the camera. --rs 720p --rt 16:9 --dur 5 --fps 24 --seed 42

Supported inline params: --rs (resolution), --rt (ratio), --dur (duration), --frames, --fps (24 only), --seed, --cf (camera_fixed), --wm (watermark).

Gotchas

  • Model names use the doubao-* convention (e.g. doubao-seedance-1-0-pro-250528) — old short names like seedance-1.0 are not valid
  • The content array replaces the old prompt + image_url fields; always use content
  • Image and text scenarios are mutually exclusive per content item — each item has either text or image_url, not both
  • first_frame and last_frame may be combined in one request, but reference_image is mutually exclusive with first_frame / last_frame — do not mix a reference image with first/last frames
  • generate_audio: true is supported by doubao-seedance-1-5-pro-251215 and the doubao-seedance-2-0 series; other models ignore this field
  • Lite models are split: *-lite-t2v-* only accepts text, *-lite-i2v-* only accepts image-to-video
  • audio_url and video_url reference items are used by the Seedance 2.0 series only
  • Resolution options are 480p, 720p, 1080p, and 4k (4k is doubao-seedance-2-0-260128 only; 2-0-fast / 2-0-mini max out at 720p) — there is no 360p or 540p
  • service_tier values are "default" and "flex" (not "standard"/"premium")
  • Duration range is 2–15 seconds (Seedance 2.0 supports 4–15) — values outside this range will fail
  • Task states use "succeeded" (not "completed") — check for this value when polling

MCP: pip install mcp-seedance | Hosted: https://seedance.mcp.acedata.cloud/mcp | See all MCP servers

Individual skills in this repo

This repo contains 20 individual skills — each has its own dedicated page.

NeverSight/learn-skills.dev

Use when generating or modifying Remotion video code, creating demo videos, or working with the demo-video/ directory

NeverSight/learn-skills.dev

Create AI avatar and talking head videos via inference.sh CLI. Recommended: P-Video-Avatar (fastest, cheapest, built-in TTS). Also: OmniHuman, Fabric, PixVerse. Audio: Inworld TTS-2 (100+ languages, emotion steering for characters), ElevenLabs, Kokoro. Capabilities: audio-driven avatars, text-to-avatar, lipsync videos, talking head generation, virtual presenters, UGC content. Use for: AI presenters, explainer videos, virtual influencers, dubbing, marketing videos, UGC ads, gaming avatars, NPC dialogue. Triggers: ai avatar, talking head, lipsync, avatar video, virtual presenter, ai spokesperson, audio driven video, heygen alternative, synthesia alternative, talking avatar, lip sync, video avatar, ai presenter, digital human, ugc, ugc video, ugc ad, avatar ugc

NeverSight/learn-skills.dev

Create AI marketing videos for ads, promos, product launches, and brand content. Models: Veo, Seedance, Wan, FLUX for visuals, Kokoro for voiceover. Types: product demos, testimonials, explainers, social ads, brand videos. Use for: Facebook ads, YouTube ads, product launches, brand awareness. Triggers: marketing video, ad video, promo video, commercial, brand video, product video, explainer video, ad creative, video ad, facebook ad video, youtube ad, instagram ad, tiktok ad, promotional video, launch video

NeverSight/learn-skills.dev

Generate AI videos with Google Veo, Seedance 2.0, HappyHorse, Wan, Grok and 40+ models via inference.sh CLI. Models: Veo 3.1, Veo 3, Seedance 2.0, HappyHorse 1.0, Wan 2.5, Grok Imagine Video, OmniHuman, Fabric, HunyuanVideo. Capabilities: text-to-video, image-to-video, reference-to-video, video editing, lipsync, avatar animation, video upscaling, foley sound. Use for: social media videos, marketing content, explainer videos, product demos, AI avatars. Triggers: video generation, ai video, text to video, image to video, veo, animate image, video from image, ai animation, video generator, generate video, t2v, i2v, ai video maker, create video with ai, runway alternative, pika alternative, sora alternative, kling alternative, seedance, happyhorse

NeverSight/learn-skills.dev

ElevenLabs automatic dubbing - translate and dub audio/video into 29 languages while preserving speaker voice via inference.sh CLI. Capabilities: auto speaker detection, voice-preserving translation, video dubbing, audio localization. Use for: content localization, video translation, multilingual content, international distribution. Triggers: dubbing, dub video, translate audio, video translation, audio translation, localize content, elevenlabs dubbing, eleven labs dub, multilingual dub, voice translation, auto dub, language dub, content localization

NeverSight/learn-skills.dev

Explainer video production guide: scripting, voiceover, visuals, and assembly. Covers script formulas, pacing rules, scene planning, and multi-tool pipelines. Use for: product demos, how-it-works videos, onboarding videos, social explainers. Triggers: explainer video, how to make explainer, product video, demo video, video production, video script, animated explainer, product demo video, tutorial video, onboarding video, walkthrough video, video pipeline

NeverSight/learn-skills.dev

Still-to-video conversion guide: model selection, motion prompting, and camera movement. Covers Wan 2.5 i2v, Seedance, Fabric, Grok Video with when to use each. Use for: animating images, creating video from stills, adding motion, product animations. Triggers: image to video, i2v, animate image, still to video, add motion to image, image animation, photo to video, animate still, wan i2v, image2video, bring image to life, animate photo, motion from image

NeverSight/learn-skills.dev

Generate talking head avatar videos with Pruna P-Video-Avatar via inference.sh CLI. Turn a portrait image into a realistic speaking video with built-in TTS. 18x faster and 6x cheaper than competitors. Models: P-Video-Avatar, P-Image (for portrait generation). Capabilities: text-to-avatar, audio-driven avatars, 30 voices, 10 languages, 720p/1080p, built-in TTS, dynamic backgrounds, full-body control. Use for: AI presenters, product demos, explainer videos, virtual influencers, marketing, education, multilingual content, UGC, gaming avatars. Triggers: avatar video, talking head, ai avatar, p-video-avatar, pruna avatar, video avatar, ai presenter, digital human, virtual presenter, lipsync, talking avatar, ai spokesperson, heygen alternative, synthesia alternative, veed alternative, fabric alternative, omnihuman alternative

NeverSight/learn-skills.dev

Generate videos with Pruna P-Video and WAN models via inference.sh CLI. Models: P-Video, WAN-T2V, WAN-I2V. Capabilities: text-to-video, image-to-video, audio support, 720p/1080p, fast inference. Pruna optimizes models for speed without quality loss. Triggers: pruna video, p-video, pruna ai video, fast video generation, optimized video, wan t2v, wan i2v, economic video generation, cheap video generation, pruna text to video, pruna image to video

NeverSight/learn-skills.dev

Render videos from React/Remotion component code via inference.sh. Pass TSX code, get MP4. Supports all Remotion APIs: useCurrentFrame, useVideoConfig, spring, interpolate, AbsoluteFill, Sequence. Configurable resolution, FPS, duration, codec. Use for: programmatic video generation, animated graphics, motion design, data-driven videos, React animations to video. Triggers: remotion, render video from code, tsx to video, react video, programmatic video, remotion render, code to video, animated video, motion graphics code, react animation video

NeverSight/learn-skills.dev

Video ad creation with exact platform-specific specs for TikTok, Instagram, YouTube, Facebook, LinkedIn. Covers dimensions, duration limits, AIDA framework, and caption requirements. Use for: video ads, social media ads, paid media creative, video marketing, ad production. Triggers: video ad, social media ad, tiktok ad, instagram ad, youtube ad, facebook ad, linkedin ad, video creative, ad specs, paid media, video marketing, ad production, reels ad, stories ad, pre roll, bumper ad

NeverSight/learn-skills.dev

Best practices and techniques for writing effective AI video generation prompts. Covers: Veo, Seedance, Wan, Grok, Kling, Runway, Pika, Sora prompting strategies. Learn: shot types, camera movements, lighting, pacing, style keywords, negative prompts. Use for: improving video quality, getting consistent results, professional video prompts. Triggers: video prompt, how to prompt video, veo prompts, video generation tips, better ai video, video prompt engineering, video prompt guide, video prompt template, ai video tips, video prompt best practices, video prompt examples, cinematography prompts

NeverSight/learn-skills.dev

YouTube thumbnail design with specific dimensions, contrast rules, and mobile preview optimization. Covers safe zones, text placement, face expression psychology, and A/B testing. Use for: YouTube thumbnails, video cover images, click-through optimization. Triggers: youtube thumbnail, thumbnail design, video thumbnail, click through rate, ctr optimization, youtube cover, video cover image, thumbnail maker, thumbnail tips, youtube design, video preview image

NeverSight/learn-skills.dev

Landing page conversion optimization with layout rules, hero section design, and CTA psychology. Covers above-the-fold formula, social proof placement, mobile design, and F-pattern reading. Use for: startup landing pages, product pages, SaaS marketing, conversion optimization. Triggers: landing page, hero section, above the fold, conversion optimization, landing page design, cta button, hero image, landing page layout, saas landing page, product page design, conversion rate, landing page best practices

NeverSight/learn-skills.dev

Configure and use the hosted YouTube Data MCP end-to-end with minimal user input. Use when users want the agent to verify Node.js and `npx`, configure MCP server config (Windows/macOS, Cursor/Codex/OpenClaw/OpenCode), request API key at setup time, run post-install capability discovery (`tools/list` and `get_patch_notes`), and then strongly recommend helper skill and Python setup for full local document and spreadsheet workflows.

NeverSight/learn-skills.dev

Creates 120fps GPU-accelerated animations with Motion.dev (Framer Motion successor) for React, Next.js, Svelte, and Astro projects. Use when user requests animation, motion, scroll effects, parallax, hero animations, gestures, drag interactions, spring physics, whileHover effects, whileInView animations, animated UI, micro-interactions, page transitions, or layout animations. Generates production TypeScript/JSX code with accessibility (prefers-reduced-motion) and performance validation (≥60fps). Supports entrance animations, gesture interactions (hover/tap/drag), scroll-based reveals, and layout transitions using spring physics and natural timing. Do NOT use for CSS-only transitions (use native CSS), static sites without JavaScript, Vue animations (use motion-v variant instead), or SVG/Canvas complex animations (GSAP better suited).

NeverSight/learn-skills.dev

Static artifact craft skill for self-contained HTML/CSS/JS documents: docs, sheets, dashboards, explainers, slides, tools, and landing pages. Use when the user asks for a durable, openable, shareable web deliverable they'll keep or hand off — a report, a dashboard, a slide deck, a data table, a page. Local folder first, temporary public link via tunnel (localhost.run), optional durable publish to Surge, GitHub Pages, or Cloudflare. Not for quick look renders, inline snippets, or throwaway scratch. Not for SPA frameworks, backend APIs, database apps, or production product UI.

NeverSight/learn-skills.dev

FFmpeg commands for video/audio conversion, trimming, compression, and processing. Use when user mentions "ffmpeg", "convert video", "compress video", "extract audio", "trim video", "gif from video", "video codec", "transcode", "screen recording", "merge videos", "video to mp4", "reduce file size", or any media processing task.

NeverSight/learn-skills.dev

Vim keybindings, motions, text objects, and operators for efficient text editing. Use when user asks about "vim commands", "vim motions", "text objects", "vim keybindings", "vim cheat sheet", "learn vim", "vim in VS Code", or any Vim editing tasks.

NeverSight/learn-skills.dev

Grok Build ONLY. Turn a 2D character still into smooth animation sprites via image_gen/image_edit base → image_to_video (6s/10s run-in-place) → ffmpeg frames → magenta chroma-key → dense sampled sprites (strip/grid/GIF). Use when the user wants video-to-sprite, motion capture from generated video, smoother run/walk cycles from dense frames, or runs /video2dsprite. Do NOT use on Codex/Claude — only Grok Build has image_to_video. Prefer generate2dsprite for crisp pixel sheets without video.

相關技能