Communitygithub.com

StarAI-2026/ComfyUI-MiniMax-H3-SkillBridge

为 MiniMax H3 图生视频生成高动态运镜 Ref2VA 提示词。基于插件自动加载的参考图(人物、场景、动作、风格),只输出可直接交给 H3 图生视频节点的六段式英文提示词,不生成角色卡或动作首帧文生图提示词。要求两轴位移、主动运镜、前景视差、障碍接触、尺度变化与物理反馈,使用官方镜头指令与自然语言弧线描述。无参考图时允许纯文字描述模式。

Qu'est-ce que ComfyUI-MiniMax-H3-SkillBridge ?

ComfyUI-MiniMax-H3-SkillBridge is a Claude Code agent skill that 为 MiniMax H3 图生视频生成高动态运镜 Ref2VA 提示词。基于插件自动加载的参考图(人物、场景、动作、风格),只输出可直接交给 H3 图生视频节点的六段式英文提示词,不生成角色卡或动作首帧文生图提示词。要求两轴位移、主动运镜、前景视差、障碍接触、尺度变化与物理反馈,使用官方镜头指令与自然语言弧线描述。无参考图时允许纯文字描述模式。.

Compatible avec~Claude Code~Codex CLI~Cursor
npx skills add https://github.com/StarAI-2026/ComfyUI-MiniMax-H3-SkillBridge/tree/HEAD/skills/%E9%AB%98%E5%8A%A8%E6%80%81%E8%BF%90%E9%95%9C

Demander à votre IA préférée

Ouvre une nouvelle conversation avec cette compétence d'agent déjà préchargée.

Documentation

高动态运镜(H3 图生视频提示词)

将插件自动加载的参考图转换为一条完整的 MiniMax H3 图生视频 Ref2VA 提示词。本 Skill 只产出视频提示词,不生成角色卡、不生成动作首帧文生图提示词、不输出分镜表。

输入模式

  • 有图模式(推荐):插件已自动加载 1-6 张参考图。用 h3-ref2va-contract.md 的标签规则定义每张图:<Subject N> 定义可复用内容(人物、场景、道具、风格),<Picture N> 只在图片作为首帧/关键帧锚点时使用。
  • 纯文字模式:没有提供任何参考图时,不得报错或拒绝。改为根据用户文字描述在 subject_definitions 中创建主体定义,例如 <Subject 1> is a ... described by the user as ...。不出现 <Picture N> 首帧锚点,summary 使用 [reference generation];只有用户明确提供动作构图描述并希望作为起始画面时,才把该描述声明为起始画面锚点并加 keyframe completion。
  • 两种模式都不输出角色卡或动作首帧文生图提示词。图片生成任务不在本 Skill 范围内。

参考文件

  • references/h3-ref2va-contract.md:六段式 Ref2VA 结构、标签、保留标记、时间线与声音规则的唯一权威契约。
  • references/high-motion-h3.md:高动态定义、镜头指令语法、运动密度与因果链、拒绝清单。
  • references/generalized-speed-grammar.md:运动原型库、镜头模块库、弹性时间分配与高速词汇。

不读取 visual-case-design.md 与 sample-motion-blueprint.md。除非用户明确要求复刻 hoverboard 样例结构,否则不得套用其固定时间点、起跳、180 度环绕或落地漂移。

输出格式

严格按节点格式输出两段:

[视觉分析]
用 2-3 句中文归纳参考图或文字描述中的主体、场景、动作与可用风格锚点。

[最终视频提示词]
英文六段式,字段名与顺序严格固定:
subject_definitions:
...
summary:
...
retention_analysis:
...
detailed_description:
...
overall_soundscape:
...
non_diegetic_music:
...

detailed_description 必须完整覆盖 H3 需要的构图、主体身份、环境与灯光、动作与状态变化、运镜、同步物理音效,以及参考内容生效的位置。不要为凑字数编造动作,时间可行性优先。

高动态硬性要求

每一条提示词必须同时满足:

  1. 主体至少沿两个轴位移:前进+下降、侧移+上升、前进+横向旋转/压弯等。
  2. 主体视在尺度或机位距离发生明显变化:充满画面、拉开、贴近镜头掠过、快速改变距离。
  3. 前景结构、线缆、栏杆、碎片或粒子横穿画面,形成快速视差。
  4. 主体与物理元素互动:穿门、越障、擦轨、绕车、撞面、偏转推力、超越车辆、擦肩而过。
  5. 加速/减速可见:身体压缩、推力形态、尾流畸变、火花、扬尘、雨水剥离、冲击波、衣物/发丝滞后。
  6. 镜头主动追拍:推近、拉开、横移、升降、变侧、变距离、跟随压弯,不得只是居中静止或匀速平行跟拍。

每个节拍写出因果链:主体发力/动作 → 身体或装备响应 → 环境反应 → 镜头响应。使用速度对比,禁止整条匀速平行跟拍。

镜头指令语法

  • 官方括号指令:[Truck left]、[Truck right]、[Pan left]、[Pan right]、[Push in]、[Pull out]、[Pedestal up]、[Pedestal down]、[Tilt up]、[Tilt down]、[Zoom in]、[Zoom out]、[Shake]、[Tracking shot]、[Static shot]。
  • 同时指令写进同一括号并用逗号分隔,如 [Tracking shot,Push in];一个括号最多 3 个;不同时间点使用不同指令组。
  • 高动态模式禁用 [Static shot],禁止同时推近+拉开等相反指令。
  • 弧线、环绕、摇臂、翻转、贴近掠过等无官方指令的运镜用自然英文描述,不得编造不存在的括号指令。
  • 至少安排 2-3 组顺序镜头指令,覆盖追拍、变向和收尾。

时间分配

按视频时长分配节拍,不套固定时间戳:

  • 约 5 秒:3 拍,0.0-0.5 立即进入高速,中段一次复合穿越或变向,结尾一个决定性空间收束。
  • 6-8 秒:3-4 拍,允许一次额外反转、障碍或环境揭示。
  • 约 10 秒:4-5 拍,包含一次大的速度对比和一次环境过渡。

时间线与声音

  • [Shot 1] 开头不写 At 切割时间;后续镜头用 At MM:SS.mmm 严格递增且不超出总时长。
  • 若以 <Picture 1> 为起始锚点,明确写出 the shot begins from <Picture 1>。
  • 对话/歌词用 <d>[语言] ...</d> 并保留原语言;说话人用稳定的 (S1)、(S2) 标识。
  • detailed_description 放同步对话、人声和镜头内声音事件;overall_soundscape 放持续环境与物理动作声;non_diegetic_music 只放观众可闻配乐,写明乐器、速度、节奏与动态,缺失时写 N/A。三部分不得互相重复。

自查清单(输出前逐项通过)

  • 六字段各出现一次且顺序严格固定。
  • 所有标签在 subject_definitions 定义后再使用,全程含义一致。
  • summary 任务类型与输入模式一致:纯文字参考为 [reference generation];有图首帧锚点为 [keyframe completion + reference generation]。
  • 每个标签在 retention_analysis 恰好一行,使用固定标记(fully_preserved、partially_preserved、attribute_transfer、weak_reference;音频用 fully_copy、partially_copy、reference、weak_reference)。
  • detailed_description 以 [Shot 1] 开场,[Shot 1] 无 At 时间戳,后续镜头时间递增且在时长内。
  • 至少两组顺序官方镜头指令;无 [Static shot];单括号不超过 3 个指令。
  • 满足六项高动态硬性要求:两轴位移、尺度变化、前景视差、障碍接触、可见加速度、主动镜头。
  • 因果链完整:动作 → 身体/装备响应 → 环境反应 → 镜头响应。
  • 至少 3 个明确时间点覆盖开始、穿越与收束。
  • overall_soundscape 与 non_diegetic_music 不重复对话,二者互不重复。
  • 结尾从开头状态在给定时长内可实现,不用静态手势或眨眼充当收尾。
  • 有图模式只用插件提供的参考图;纯文字模式以用户描述为准创建主体,不编造未描述的细节。

Individual skills in this repo

This repo contains 10 individual skills — each has its own dedicated page.

StarAI-2026/ComfyUI-MiniMax-H3-SkillBridge

Create complete stylized 3D animated shorts from a story idea through an ordered production workflow covering project brief, story outline, character and environment cards, standardized shot planning, text or optional pencil storyboards, video-model selection, single-shot generation, assembly, BGM matching, and final review. Use when the user wants an end-to-end narrative animation workflow with strong character consistency, scene continuity, timing, camera, performance, and audio control. Not for single images, simple edits, photorealistic live action, or one standalone clip.

StarAI-2026/ComfyUI-MiniMax-H3-SkillBridge

For marketers and creators producing promotional content for brands, products, websites, apps, shops, or personal projects. Users provide logos, product images, interface screenshots, official links, or other verifiable assets and confirm duration, aspect ratio, audience, and campaign focus. The Skill organizes brand facts and asset provenance, selects a narrative direction, plans precise beats and shots, generates needed imagery, video, voiceover, or music, and completes assembly and pre-delivery review. It outputs a promotional short that highlights product capabilities, use cases, and a call to action. Best for launches, website showcases, and social promotion; not for imitating real brand marks without authorized assets, inventing product claims, or producing long-form narrative films.

StarAI-2026/ComfyUI-MiniMax-H3-SkillBridge

For users creating a two-player co-op game menu or opening animation. Users provide two player names, a game title, a target visual style, and optional character reference images. The Skill locks identity cues, generates an approval image from a fixed menu framework with coordinated color, buttons, icons, and typography, then uses the approved result to rebuild the character, UI-copy, and event timing instructions for the final video. It outputs a co-op game intro featuring two characters, player cards, and menu interaction motion. Best for game concepts, character-led menus, and social content; not for playable game development, complex multi-page UI, exact brand-logo replication, or generic character-free title sequences.

StarAI-2026/ComfyUI-MiniMax-H3-SkillBridge

Write MiniMax H3 video generation prompts for T2VA, I2VA, FL2VA, L2VA, and Ref2VA. Use when rewriting multimodal requests into H3 prompt structures, composing integrated_multimodal_description, overall_soundscape, and non_diegetic_music, aligning keyframes, or defining reference labels for images, videos, and audio.

StarAI-2026/ComfyUI-MiniMax-H3-SkillBridge

For creators making surreal short videos that blend rough glowing hand-drawn animation with live-action spaces. Users provide a scene idea, contact object or hand, desired mood, and optional language or style constraints. The Skill clarifies the physical contact, designs continuous morphing, escape route, and delayed handheld chase movement, then writes a reusable 15-second 16:9 video prompt in the user's language. After user confirmation it recommends MiniMax H3 generation and checks contact realism, camera delay, rough glowing stroke texture, and non-horror tone. Best for single-scene creative clips, not polished CG, horror jump scares, plush characters, or multi-scene cuts.

StarAI-2026/ComfyUI-MiniMax-H3-SkillBridge

Turn user-provided reference images and script into multi-segment H3 Ref2VA video prompts with character presenter, holographic interaction, and full narration coverage. Users provide a character reference image, an environment reference image, and a complete narration script. The Skill auto-splits the script into segments within user-specified max duration, calculates narration budget dynamically, designs at least 2 holographic interactions per segment, and outputs complete self-contained Ref2VA format prompts. No external skill dependencies. Style, aspect ratio, and segment duration are flexible.

StarAI-2026/ComfyUI-MiniMax-H3-SkillBridge

Turn product images and ad requirements into minimalist product ad shorts for e-commerce promotion and product launches. The Skill confirms format and product variants, extracts selling points, writes concise English ad copy, builds product anchors, plans beat-synced typography/storyboards, and generates a clean product film with premium camera language. Not for KOC talking-head ads, general editing, or complex screen demos.

StarAI-2026/ComfyUI-MiniMax-H3-SkillBridge

For musicians, video creators, and social-media editors producing AI music videos or emotional short films with lyric typography. Users provide music, lyrics, references, characters, typography direction, mood, or target platform. The Skill analyzes beat and vocal timing, separates character, scene, and text references, designs beat-reactive spatial typography, decomposes long works into connected shots, audits prompts, and routes generation for H3 or other video tools. It outputs MV concepts, shot prompts, lyric text plans, and stitching guidance. Best for stylized MVs and subtitle-driven music visuals, not ordinary caption cleanup, licensed IP copying, or fully manual post-production editing.

StarAI-2026/ComfyUI-MiniMax-H3-SkillBridge

For creators, educators, and social-video editors who need a tactile paper-collage language for narration, knowledge points, opinions, or abstract topics. Users provide source copy, story beats, or a core concept and may specify aspect ratio, duration, palette, and audio needs. The Skill extracts meaning, proposes visual metaphors, prepares a production plan and storyboard, generates approved halftone collage stills, then creates stop-motion clips with paper movement and tactile sound effects, with optional final assembly. By default it keeps collage SFX and does not add BGM, voiceover, or subtitles unless requested. Best for explainers, viewpoints, story visuals, and social B-roll; not for presenter ads, editable layers, complex typography, or prompt-only tasks.

StarAI-2026/ComfyUI-MiniMax-H3-SkillBridge

For creators explaining science, education, or general knowledge through tactile handmade papercraft visuals. Users provide a topic, core knowledge points, or source material and may specify audience, duration, aspect ratio, and deliverable type. The Skill extracts the learning goal and visual metaphor, proposes creative directions, designs paper characters, layered diorama sets, and props, creates preview concepts plus image and video prompts, and plans storyboards, camera movement, transitions, and sound with staged approvals and review checklists. It outputs a production-ready papercraft stop-motion explainer package, or selected assets such as still prompts, image-series prompts, short-video prompts, or storyboards. Best for cut-paper, pop-up-book, layered diorama, and miniature stop-motion explainers; not for standard 2D cartoons, line doodles, live action, or explainers without a paper-art look.

Skills associés