Communitygithub.com

0xsline/openchatcut

AI image generation via Fal.ai, gpt-image-2, nano-banana, MiniMax image-01, and xAI Grok Imagine. Use when the user wants to generate or create an image / picture / still.

What is openchatcut?

openchatcut is a Claude Code agent skill that aI image generation via Fal.ai, gpt-image-2, nano-banana, MiniMax image-01, and xAI Grok Imagine. Use when the user wants to generate or create an image / picture / still.

Works with~Claude Code~Codex CLI~Cursor
npx skills add https://github.com/0xsline/openchatcut/tree/HEAD/src/agent/skills/image-gen

Ask in your favorite AI

Open a new chat with this agent skill pre-loaded.

Documentation

Image Gen

Generate AI images via submit_image (configured provider keys only). Prefer one clear still per request unless the user asked for variants.

Model Selection

ModelReferenceStrengthsMax refs
fal + falModelreferences/fal.mdExplicit Fal catalog; see tool schema for per-model limitsModel-specific
gpt-image-2references/gpt-image-2.mdBest text rendering, strongest prompt adherence16
nano-bananareferences/nano-banana.mdStrongest reference-image fidelity14
image-01references/image-01.mdMiniMax stills / live style; one subject reference via R21
grok-imaginereferences/grok-imagine.mdxAI Grok Imagine; text-to-image, ≤4 outputs, 1K/2K0
  • If Fal.ai is selected or requested, use model: "fal" and the requested falModel or saved Fal default from capabilities. Ask if none is selected. Native-provider defaults and controls below do not apply to Fal.
  • Default: gpt-image-2 when that key is on.
  • Reference-heavy → nano-banana.
  • User named MiniMax / only MiniMax image key on → image-01.
  • Respect capabilities: do not call a model whose vendor is not configured.

IMPORTANT: Before generating, READ the chosen model's reference.

Tool Params

ParamValuesDefault
aspectRatio1:1, 16:9, 9:16, 4:3, 3:4, 3:2, 2:3, 4:5, 5:4, 21:916:9
imageSize512px, 1K, 2K, 4K (model-specific)1K
width / heightGPT Image: 512–3840, /16; MiniMax: 512–2048, /8—
qualitylow, medium, high, auto (gpt-image-2 only)high
referenceAssetIdsArray of project asset ids — backend resolves bytes server-side—
nameShort descriptive asset name shown in the library—
countNumber of images to generate (1–10; image-01 max 9)1
promptOptimizerMiniMax image-01 only — prompt_optimizerfalse
seedMiniMax image-01 only—
maskAssetId, background, moderation, inputFidelityGPT Image edit/output controls—
outputFormat, outputCompressionGPT Image PNG/JPEG/WebP controlsPNG

Defaults

  • Aspect ratio: 16:9. If the project composition is not 16:9, ASK the user which aspect ratio they want before generating.
  • Size: 1K.

Ask Before Submit

  • Never auto-upgrade size.
  • Only pass imageSize: "2K" or "4K" when the user explicitly asks. Warn that 2K/4K are EXPERIMENTAL and may be slower.

Reference Images

Use when the user provides source material to edit, blend, or use as visual guidance (e.g. "change the background", "combine these into a poster").

  • Pass project asset ids via referenceAssetIds. The backend fetches and encodes them server-side — never pull the asset bytes yourself.
  • When the user @-references an image asset, pass its id directly in referenceAssetIds.
  • Formats accepted by backend: png, jpeg, webp, svg (auto-rasterized to png), heic, heif. Each ≤ 50MB.

Run

// Basic generation
submit_image({
  model: "gpt-image-2",
  prompt: "a cute orange cat",
  name: "Cat",
});

// With quality (gpt-image-2 only)
submit_image({
  model: "gpt-image-2",
  prompt: "hero poster with bold title",
  quality: "high",
  name: "Hero Poster",
});

// With reference images — pass project asset ids; backend resolves bytes
submit_image({
  model: "gpt-image-2",
  prompt: "change background to beach",
  referenceAssetIds: ["<assetId>"],
  name: "Beach Edit",
});

// Reference-heavy with nano-banana
submit_image({
  model: "nano-banana",
  prompt: "composite poster",
  referenceAssetIds: ["<id1>", "<id2>"],
  name: "Composite",
});

// Multiple images
submit_image({
  model: "gpt-image-2",
  prompt: "product shots",
  count: 3,
  name: "Product",
});

// MiniMax (optional single subject reference; R2 must be configured for refs)
submit_image({
  model: "image-01",
  prompt: "matte product bottle on marble, soft studio light",
  name: "Bottle still",
  promptOptimizer: false,
});

OpenChatCut’s submit_image may return completed pool assets synchronously depending on the provider path. If a jobId is returned, use track_progress; otherwise treat the asset ids in the result as done.

Rules

  • Always provide name with a short descriptive asset name.
  • Before submitting, briefly tell the user what you're about to generate — especially when generating multiple images.
  • Only call models whose vendor key is configured (capabilities prompt).

Individual skills in this repo

This repo contains 11 individual skills — each has its own dedicated page.

0xsline/openchatcut

Connect an MCP-capable coding agent to OpenChatCut and edit local video projects. Use when the user asks to install, connect, or set up OpenChatCut; inspect or edit an OpenChatCut project; work with its timeline, transcript, captions, media, generation, motion graphics, audio, color, or export tools; or recover from an OpenChatCut MCP error.

0xsline/openchatcut

Use when acquiring or importing media into a OpenChatCut project asset library for video editing or creation, including local/attached videos, user-provided paths, public media URLs, web video/audio/image assets, upload fallback decisions, and deciding between import_media, download_media, or manual user action.

0xsline/openchatcut

Use when a OpenChatCut video editing or creation workflow needs export, render, download, share, final delivery, subtitle-file export, render choice, local-only asset handling, or export fallback explanation.

0xsline/openchatcut

Use when a OpenChatCut tool call fails or returns an unexpected shape.

0xsline/openchatcut

Music generation via Mureka, MiniMax, Atlas Cloud, and Sonilo. Use for instrumentals, songs, soundtracks, track/stem generation, covers, or video-conditioned scoring of the finished cut through `submit_music`.

0xsline/openchatcut

OpenChatCut product knowledge — UI layout, editor features, and how generation providers are configured. Use when the user asks about the product interface, how to use a feature, where to find something, or needs GUI guidance for something the agent cannot do directly. Also use as fallback when a task fails and the user needs to complete it manually in the UI. NOT for live project-state queries ("where are my folders?", "what's on my timeline?", "where is clip X?") — those are answered by `read_project`, not by this skill.

0xsline/openchatcut

AI shader generator for WebGL video effects, transitions, masks, and color grading (LUT / 调色 / 电影感 / film look). Use when the user wants a video effect (滤镜 / 特效), a transition (转场 / crossfade / wipe / cube / 3d), a mask (蒙版 / 遮罩 / reveal), a zoom / push-in (推近 / 推镜头), or a color grade — try the built-in effects (zoom, builtin LUTs) before generating a new shader.

0xsline/openchatcut

Use when checking whether agent edits are reflected in the OpenChatCut project and editor.

0xsline/openchatcut

AI video generation via Fal.ai, Seedance 2.0, Kling, MiniMax Hailuo, xAI Grok Imagine, and OFox. Use when the user wants to generate a video clip — text-to-video, image-to-video, first/last-frame transitions, reference-guided generation, multi-shot, or generatively editing / extending an existing clip.

0xsline/openchatcut

Text-to-Speech (TTS), voiceover, narration placement/sync, and custom sound effects (SFX) generator. Use when the user wants generated speech from text, wants to add/replace/align narration or voiceover for an existing video/timeline, wants to keep existing voiceover synced after visual retiming edits, needs voice audition/selection, or explicitly wants a newly generated/custom sound effect that is not available in the Sound Effects library.

0xsline/openchatcut

Use when the agent should ask the user for structured input with an in-chat form, including single-select, multi-select, text fields, style pickers, or voice audition choices.

Related Skills