Communitygithub.com

PrunaAI/pruna-skills

Use when someone wants an original AI song with vocals — sung lyrics, a style prompt track, or source audio for a music video.

What is pruna-skills?

pruna-skills is a Gemini CLI agent skill that use when someone wants an original AI song with vocals — sung lyrics, a style prompt track, or source audio for a music video.

Works with~Claude Code~Codex CLI~Cursor✓Gemini CLI
npx skills add https://github.com/PrunaAI/pruna-skills/tree/HEAD/skills/audio/music-2.5

Ask in your favorite AI

Open a new chat with this agent skill pre-loaded.

Documentation

Prerequisites

Install and load these skills before generating (skip if already in context via @pruna):

SkillDescriptionInstall
generation-diversityUse when writing any generative prompt — ritual seed, explicit structure, scenario axes, and quality gates before paid API calls.npx skills add PrunaAI/pruna-skills@generation-diversity -y
audio-promptingUse when crafting TTS, music, or bed prompts for any generative audio model — director style, song structure, and post-production layering.npx skills add PrunaAI/pruna-skills@audio-prompting -y
pruna-apiUse before any Pruna or Replicate HTTP call — credentials, upload/poll/download, parallel batches, and agent safety.npx skills add PrunaAI/pruna-skills@pruna-api -y

Or install the full suite once: npx skills add PrunaAI/pruna-skills@pruna -y

Follow each skill's Before generating / craft sections — do not restate guide content here.

Agent habit

In the first reply, name `music-2.5` in backticks, confirm REPLICATE_API_TOKEN (or stop with signup links from pruna-api), then ask for required inputs. Open intake → generation-diversity clarification intake before the first POST. Redirect when When NOT to use fits better.

When NOT to use

Use a different skill instead:

SkillDescriptionInstall
gemini-3.1-flash-ttsUse when someone needs spoken narration or voiceover — explainer tracks, documentary lines, or voice to pair with generated video.npx skills add PrunaAI/[email protected] -y
stable-audio-2.5Use when someone wants light instrumental background music — an ambient bed under dialogue or underscore for reels and explainers.npx skills add PrunaAI/[email protected] -y

Environment

export REPLICATE_API_TOKEN=r8_...

Requires ffmpeg / ffprobe for slicing and assembly in the music-video workflow.

HTTP (curl)

curl -s -X POST \
  -H "Authorization: Bearer ${REPLICATE_API_TOKEN}" \
  -H "Content-Type: application/json" \
  -d '{
    "input": {
      "lyrics": "[Verse]\nWe built it line by line\nEvery skill a stepping stone\n\n[Chorus]\nRun the pipeline, watch it grow\nPruna models, let them flow",
      "prompt": "Indie pop, uplifting, warm female vocal, 92 BPM, acoustic guitar and mellow synth pads, no harsh distortion",
      "sample_rate": 44100,
      "bitrate": 256000,
      "audio_format": "mp3"
    }
  }' \
  "https://api.replicate.com/v1/models/minimax/music-2.5/predictions"

Poll urls.get until status is succeeded; download output.

Before generating

  1. Complete Prerequisites guide reading order.
  2. Confirm lyrics (with structure tags) and optional style prompt. When listing required fields, name REPLICATE_API_TOKEN (Replicate — not PRUNA_API_KEY).
  3. Model notes: structure tags on their own lines — [Intro] [Verse] [Pre Chorus] [Chorus] [Hook] [Bridge] [Solo] [Inst] [Build Up] [Drop] [Interlude] [Break] [Transition] [Outro]. \n = line break (also a safe video cut boundary); \n\n = pause. Max ~5 minutes per generation. English and Mandarin have strongest pronunciation. Data is sent to MiniMax via Replicate — see their privacy policy.

Required input

  • lyrics (string) — 1–3,500 characters

Common optional fields

  • prompt — genre, mood, tempo, vocal timbre, instruments (up to ~2,000 chars)
  • sample_rate: 16000 · 24000 · 32000 · 44100 (default)
  • bitrate: 32000 · 64000 · 128000 · 256000 (default)
  • audio_format: mp3 (default) · wav · pcm

Typical next steps

Common follow-ons after this skill:

SkillDescriptionInstall
music-videoUse when someone wants a full music video — original song or vocals, performance clips, B-roll, and lyric-synced edits.npx skills add PrunaAI/pruna-skills@music-video -y
whisperxUse when someone needs word-level timestamps from audio — lyric alignment, cut-safe line boundaries, or caption source timing before burn-in with video-editing.npx skills add PrunaAI/pruna-skills@whisperx -y
p-videoUse when someone wants a simple short clip from text or images — quick B-roll, drafts, or start/end frame animation. Not when the brief needs cinematic generation, highest quality, tight lip-sync, or imported audio at 1080p.npx skills add PrunaAI/pruna-skills@p-video -y

Individual skills in this repo

This repo contains 11 individual skills — each has its own dedicated page.

PrunaAI/pruna-skills

Use when someone needs word-level timestamps from audio — lyric alignment, cut-safe line boundaries, or caption source timing before burn-in with video-editing.

PrunaAI/pruna-skills

Use before any Pruna or Replicate HTTP call — credentials, upload/poll/download, parallel batches, and agent safety.

PrunaAI/pruna-skills

Use when assembling or polishing already-rendered clips with ffmpeg — concat, crossfades, burned captions and subtitles, text/logo overlays, before/after sliders, background music beds, platform export — or when composing a multi-layer HTML combination video with Hyperframes. Not for AI video generation, prompt craft, or model-based video edits.

PrunaAI/pruna-skills

Use when someone wants to edit an existing photo — change outfits or backgrounds, compose from reference images, or apply prompt-driven edits.

PrunaAI/pruna-skills

Use when someone explicitly wants the fastest, cheapest photo generation — mood boards, bulk panels, or quick iterations — not when controlled photoreal or in-image text is needed.

PrunaAI/pruna-skills

Use when someone wants virtual try-on — dress a person in clothes from reference photos for fashion or ecommerce.

PrunaAI/pruna-skills

Use when installing the full Pruna generative media suite — all guides, tools, and workflows in one package.

PrunaAI/pruna-skills

Use when someone wants a cinematic clip from text or start/end frames — product ads, documentary shots, or dialogue with generated audio. Not for 1080p, imported audio tracks, or talking-head-only hosts.

PrunaAI/pruna-skills

Use when someone wants a polished short clip from text, images, or imported audio — 1080p B-roll, start/end frame animation, or a motion shot with a mixed track. Not for cinematic generated-audio clips or talking-head-only hosts.

PrunaAI/pruna-skills

Use when someone wants to edit an existing video with a text instruction — recolor, restyle, remove or add objects, change environment or lighting, update on-screen text, or apply optional reference-guided product and accessory edits. Not for a new clip from scratch or ffmpeg assembly.

PrunaAI/pruna-skills

Use when someone wants a simple short clip from text or images — quick B-roll, drafts, or start/end frame animation. Not when the brief needs cinematic generation, highest quality, tight lip-sync, or imported audio at 1080p.

Related Skills