Communitygithub.com

FilippTrigub/abra

Generate videos from text or images using Higgsfield's multi-model cloud platform. Supports kling, seedance, dop, and dop-preview models. Uses a small preset layer for common creative styles. Auto-detects text-to-video or image-to-video based on input. No GPU required.

Qu'est-ce que abra ?

abra is a Claude Code agent skill that generate videos from text or images using Higgsfield's multi-model cloud platform. Supports kling, seedance, dop, and dop-preview models. Uses a small preset layer for common creative styles. Auto-detects text-to-video or image-to-video based on input. No GPU required.

Compatible avec✓Claude Code~Codex CLI~Cursor
npx skills add https://github.com/FilippTrigub/abra/tree/HEAD/skills/video-generator

Demander à votre IA préférée

Ouvre une nouvelle conversation avec cette compétence d'agent déjà préchargée.

Documentation

video-generator — Multi-Model Cloud Video (Higgsfield)

Generates video clips via Higgsfield — a platform that aggregates 6+ leading models under one API and billing account. Choose the right model for cost vs. quality tradeoff; the rest of the interface is identical.

The skill directory (where this SKILL.md lives) is referred to as $SKILL_DIR below.

Cloud-based. No local GPU required. Requires a Higgsfield account and API key from cloud.higgsfield.ai.


Modes

ModeTriggerWhat happens
text-to-videoNo images in input/Generates video from prompt alone
image-to-videoImages present in input/Animates each image into a clip
auto (default)—Picks the mode above automatically

Presets

The skill uses presets as thin overlays on top of the same video engine. A preset can set the default model, aspect ratio, duration, and prompt prefix. Explicit config values still win.

PresetBest forDefault modelDefault ratio
cinematic (default)General-purpose cinematic videokling16:9
social-hookScroll-stopping short-formseedance9:16
motion-design-adSaaS/product motion adskling16:9
ecommerce-adProduct promo clipskling9:16
brand-storyFounder/brand narrativekling16:9
product-360Isolated product showcasekling1:1

Workflow rules:

  • choose a preset before editing config.json
  • verify the preset matches the desired output style
  • if the user does not choose, keep the default preset and record it in config
  • do not stack the same style twice; if the prompt already matches the preset, keep the prompt as-is

Models

KeyFull model IDBest for
kling (default)kling-video/v2.1/pro/image-to-videoRealistic motion, cost-efficient
seedancebytedance/seedance/v1/pro/image-to-videoNative audio sync
dophiggsfield-ai/dop/standardCinema Studio optical physics
dop-previewhiggsfield-ai/dop/previewCinema Studio (preview tier)

These are the model IDs confirmed in the official docs. Additional models (Sora, Veo, Wan, etc.) can be found at cloud.higgsfield.ai/explore. To add one, insert it into MODEL_IDS in generate.py and add a --model alias.

Credits are billed per successful generation only; failed and NSFW requests are refunded.


Setup (first run only)

cd "$SKILL_DIR" && uv sync

Get credentials from cloud.higgsfield.ai (API Key + API Secret), then export them using one of these forms:

# Option A — combined (recommended)
export HF_KEY="your-api-key:your-api-secret"

# Option B — separate vars
export HF_API_KEY="your-api-key"
export HF_API_SECRET="your-api-secret"

Agent Workflow

1. Ask the user

Before I generate the video, I need to know:

🎬 Mode
   - text-to-video (prompt only) or image-to-video (drop images in input/)
   - default: auto-detect

🎛️ Preset (default: cinematic)
   - cinematic · social-hook · motion-design-ad · ecommerce-ad · brand-story · product-360
   - verify the choice before writing config.json
   - if omitted, keep the default preset and record that choice

🤖 Model  (default: kling)
   kling · seedance · dop · dop-preview
   Recommendation: kling for speed/cost, seedance for quality

💬 Prompt  (required)
   Describe the motion or scene, e.g.:
   "slow cinematic push-in, warm golden hour light, subtle camera drift"

⚙️  Settings
   - duration     — video length in seconds (default: 6, range: 3–16)
   - aspect_ratio — 16:9 · 9:16 · 1:1  (default: 16:9)
   - resolution   — 720p · 1080p  (default: 1080p)

Wait for user response before proceeding.

2. Edit config.json

Write or update $SKILL_DIR/config.json based on the user's choices. Set preset explicitly, even when using the default.

3. (image-to-video only) Place images

Copy source images to $SKILL_DIR/input/. Supported: .jpg .jpeg .png .webp

4. Run

cd "$SKILL_DIR" && uv run python scripts/generate.py --config config.json

5. Report results

Tell the user:

  • Output file path(s) in output/
  • Model used and approximate credit cost
  • Video duration and resolution

Config Reference

KeyValuesDefaultDescription
input_dirpath./inputSource images (image-to-video only)
output_dirpath./outputWhere MP4s are written
modeauto · text-to-video · image-to-videoautoGeneration mode
presetcinematic · social-hook · motion-design-ad · ecommerce-ad · brand-story · product-360cinematicCreative preset overlay
modelkling · seedance · dop · dop-previewklingWhich model to use
promptstring(required)Motion or scene description
durationinteger 3–165Video length in seconds
aspect_ratio16:9 · 9:16 · 1:116:9Output aspect ratio
extra_paramsobject{}Pass-through params to the model API

Common Invocations

# Text-to-video with default model (kling)
cd "$SKILL_DIR" && uv run python scripts/generate.py \
  --prompt "drone shot over misty mountains at sunrise" --duration 8

# Image-to-video with Kling (drop images in input/ first)
cd "$SKILL_DIR" && uv run python scripts/generate.py \
  --prompt "slow cinematic push-in, golden hour light" --model kling

# Higher quality: Seedance
cd "$SKILL_DIR" && uv run python scripts/generate.py \
  --prompt "aerial pull-back over a foggy city at dawn" --model seedance

# Vertical video for Reels/TikTok
cd "$SKILL_DIR" && uv run python scripts/generate.py \
  --prompt "talking head, natural bokeh background" \
  --aspect-ratio 9:16 --model kling

# Fast draft: lower resolution, shorter clip
cd "$SKILL_DIR" && uv run python scripts/generate.py \
  --prompt "product reveal on dark background" \
  --resolution 720p --duration 4 --model dop-preview

Output

Each job writes one .mp4 to output_dir:

  • text-to-video: output.mp4
  • image-to-video: <image_stem>.mp4 per input image

vs animate-image

animate-imagevideo-generator
Providerfal.aiHiggsfield
Modeimage→video onlytext→video + image→video
ModelLTX-2.3 Fast (fixed)Kling, Sora, Veo, Wan, Seedance, MiniMax
SpeedVery fastModel-dependent
Max resolution4K1080p
Cinematic controlsNoVia extra_params

Use animate-image for fast, cheap, high-resolution clips from images. Use video-generator when you need text-to-video, model choice, or Higgsfield's Cinema Studio physics controls.


API Notes

The skill talks directly to the Higgsfield REST API (cloud.higgsfield.ai). Model endpoint paths are defined in MODEL_ENDPOINTS at the top of generate.py. If Higgsfield updates a route, override via config "extra_params" or update the constant — no other changes needed.

Individual skills in this repo

This repo contains 7 individual skills — each has its own dedicated page.

FilippTrigub/abra

>- Animate a still image into a short video clip using fal.ai's LTX-2.3 Fast image-to-video model in the cloud. No GPU required - runs entirely on fal.ai serverless infrastructure.

FilippTrigub/abra

>- Auto-describe and caption images using a local vision-language model. Writes a JSON sidecar per image containing a one-sentence description, a suggested Instagram caption with hashtags, and detected content tags.

FilippTrigub/abra

>- Animated caption pipeline. Use this skill when the user wants to burn word-by-word animated captions into videos — using Whisper for transcription and pycaps for rendering. Supports default minimalist style or a futuristic CSS theme with alternating gold/magenta glowing words.

FilippTrigub/abra

>- Cut videos into segments, rearrange them, and produce an output video with a specific cuts-per-second rate. Uses MoviePy for video manipulation. Prioritizes audio transcription for timestamped cutting, falls back to adaptive scene detection.

FilippTrigub/abra

>- Edit a region of a video using a text prompt via Wan2.1-VACE inpainting. Supports two modes: background (auto-segment via rembg) or region (rectangle defined by fractions). Requires a CUDA GPU with at least 8 GB free VRAM.

FilippTrigub/abra

>- Video enhancement pipeline. Use this skill when the user wants to sharpen, colour grade, warm, or normalise the audio of videos — including presets for natural, cinematic, or vivid looks.

FilippTrigub/abra

>- Remove the background from every frame of a video using AI (BiRefNet-general via rembg). Outputs transparent-background video or composites onto a solid colour or image. Requires a CUDA GPU with at least 3 GB free VRAM.

Skills associés