Communitygithub.com

PrunaAI/pruna-skills

Use when someone needs spoken narration or voiceover — explainer tracks, documentary lines, or voice to pair with generated video.

¿Qué es pruna-skills?

pruna-skills is a Gemini CLI agent skill that use when someone needs spoken narration or voiceover — explainer tracks, documentary lines, or voice to pair with generated video.

Compatible con~Claude Code~Codex CLI~Cursor✓Gemini CLI
npx skills add https://github.com/PrunaAI/pruna-skills/tree/HEAD/skills/audio/gemini-3.1-flash-tts

Preguntar en tu IA favorita

Abre un nuevo chat con esta habilidad de agente ya precargada.

Documentación

Prerequisites

Install and load these skills before generating (skip if already in context via @pruna):

SkillDescriptionInstall
generation-diversityUse when writing any generative prompt — ritual seed, explicit structure, scenario axes, and quality gates before paid API calls.npx skills add PrunaAI/pruna-skills@generation-diversity -y
audio-promptingUse when crafting TTS, music, or bed prompts for any generative audio model — director style, song structure, and post-production layering.npx skills add PrunaAI/pruna-skills@audio-prompting -y
video-promptingUse when crafting video or motion prompts for any generative model — dramaturgy, camera, physics-safe motion, frame anchors, and clip chaining.npx skills add PrunaAI/pruna-skills@video-prompting -y
pruna-apiUse before any Pruna or Replicate HTTP call — credentials, upload/poll/download, parallel batches, and agent safety.npx skills add PrunaAI/pruna-skills@pruna-api -y

Or install the full suite once: npx skills add PrunaAI/pruna-skills@pruna -y

Follow each skill's Before generating / craft sections — do not restate guide content here.

Agent habit

In the first reply, name `gemini-3.1-flash-tts` in backticks, confirm REPLICATE_API_TOKEN (or stop with signup links from pruna-api), then ask for required inputs. Open intake → generation-diversity clarification intake (locale, voice, script) before the first POST. Redirect when When NOT to use fits better.

When NOT to use

Use a different skill instead:

SkillDescriptionInstall
p-video-avatarUse when someone wants a person on camera speaking a script — lip-synced host, spokesperson, or narrated avatar from a portrait photo.npx skills add PrunaAI/pruna-skills@p-video-avatar -y
music-2.5Use when someone wants an original AI song with vocals — sung lyrics, a style prompt track, or source audio for a music video.npx skills add PrunaAI/[email protected] -y
stable-audio-2.5Use when someone wants light instrumental background music — an ambient bed under dialogue or underscore for reels and explainers.npx skills add PrunaAI/[email protected] -y

Environment

export REPLICATE_API_TOKEN=r8_...

Requires ffmpeg / ffprobe when trimming, concatenating scene VO, or mixing with a bed.

HTTP (curl)

curl -s -X POST \
  -H "Authorization: Bearer ${REPLICATE_API_TOKEN}" \
  -H "Content-Type: application/json" \
  -d '{
    "input": {
      "text": "[warmly] The plush went flying. [short pause] And then it was gone.",
      "voice": "Sulafat",
      "prompt": "Warm storybook narrator, gentle pace, empathetic, no announcer voice.",
      "language_code": "en-US"
    }
  }' \
  "https://api.replicate.com/v1/models/google/gemini-3.1-flash-tts/predictions"

Poll urls.get until status is succeeded; download output (audio URL). Shared client: follow pruna-api (Replicate HTTP in the tool skill).

Before generating

  1. Complete Prerequisites guide reading order.
  2. Confirm text, voice, prompt, and language_code with the user. text, prompt, and inline [tags] must align — same emotional direction (see audio-prompting tts-style-prompting). When listing fields, name REPLICATE_API_TOKEN (Replicate — not PRUNA_API_KEY).
  3. Model notes: combined text + prompt ≤ ~8,000 bytes; output capped ~655s. When TTS feeds p-video as input.audio, keep each line ≤ ~19s (ffprobe) — P-API clips audio at 20s. Common voices: Kore, Aoede, Sulafat, Achird, Charon, Puck, Vindemiatrix — full list on the Replicate readme.

Required input

  • text (string) — spoken copy; supports inline [tags]. Max ~4,000 bytes.

Common optional fields

  • voice (default Kore)
  • prompt — style / director notes (max ~4,000 bytes)
  • language_code — BCP-47 (default en-US)

Inline tags (examples): [sigh] [laughing] [whispering] [short pause] [medium pause] [long pause] [excitedly].

Typical next steps

Common follow-ons after this skill:

SkillDescriptionInstall
p-videoUse when someone wants a simple short clip from text or images — quick B-roll, drafts, or start/end frame animation. Not when the brief needs cinematic generation, highest quality, tight lip-sync, or imported audio at 1080p.npx skills add PrunaAI/pruna-skills@p-video -y
p-video-avatarUse when someone wants a person on camera speaking a script — lip-synced host, spokesperson, or narrated avatar from a portrait photo.npx skills add PrunaAI/pruna-skills@p-video-avatar -y
stable-audio-2.5Use when someone wants light instrumental background music — an ambient bed under dialogue or underscore for reels and explainers.npx skills add PrunaAI/[email protected] -y
narrated-multi-sceneUse when someone wants a multi-part story with voiceover — episodic B-roll, chaptered promo, or several linked video scenes without on-camera dialogue.npx skills add PrunaAI/pruna-skills@narrated-multi-scene -y
video-editingUse when assembling or polishing already-rendered clips with ffmpeg — concat, crossfades, burned captions and subtitles, text/logo overlays, before/after sliders, background music beds, platform export — or when composing a multi-layer HTML combination video with Hyperframes. Not for AI video generation, prompt craft, or model-based video edits.npx skills add PrunaAI/pruna-skills@video-editing -y

Individual skills in this repo

This repo contains 13 individual skills — each has its own dedicated page.

PrunaAI/pruna-skills

Use when someone wants an original AI song with vocals — sung lyrics, a style prompt track, or source audio for a music video.

PrunaAI/pruna-skills

Use when someone needs word-level timestamps from audio — lyric alignment, cut-safe line boundaries, or caption source timing before burn-in with video-editing.

PrunaAI/pruna-skills

Use before any Pruna or Replicate HTTP call — credentials, upload/poll/download, parallel batches, and agent safety.

PrunaAI/pruna-skills

Use when assembling or polishing already-rendered clips with ffmpeg — concat, crossfades, burned captions and subtitles, text/logo overlays, before/after sliders, background music beds, platform export — or when composing a multi-layer HTML combination video with Hyperframes. Not for AI video generation, prompt craft, or model-based video edits.

PrunaAI/pruna-skills

Use when crafting video or motion prompts for any generative model — dramaturgy, camera, physics-safe motion, frame anchors, and clip chaining.

PrunaAI/pruna-skills

Use when someone wants to edit an existing photo — change outfits or backgrounds, compose from reference images, or apply prompt-driven edits.

PrunaAI/pruna-skills

Use when someone explicitly wants the fastest, cheapest photo generation — mood boards, bulk panels, or quick iterations — not when controlled photoreal or in-image text is needed.

PrunaAI/pruna-skills

Use when someone wants virtual try-on — dress a person in clothes from reference photos for fashion or ecommerce.

PrunaAI/pruna-skills

Use when installing the full Pruna generative media suite — all guides, tools, and workflows in one package.

PrunaAI/pruna-skills

Use when someone wants a cinematic clip from text or start/end frames — product ads, documentary shots, or dialogue with generated audio. Not for 1080p, imported audio tracks, or talking-head-only hosts.

PrunaAI/pruna-skills

Use when someone wants a polished short clip from text, images, or imported audio — 1080p B-roll, start/end frame animation, or a motion shot with a mixed track. Not for cinematic generated-audio clips or talking-head-only hosts.

PrunaAI/pruna-skills

Use when someone wants to edit an existing video with a text instruction — recolor, restyle, remove or add objects, change environment or lighting, update on-screen text, or apply optional reference-guided product and accessory edits. Not for a new clip from scratch or ffmpeg assembly.

PrunaAI/pruna-skills

Use when someone wants a simple short clip from text or images — quick B-roll, drafts, or start/end frame animation. Not when the brief needs cinematic generation, highest quality, tight lip-sync, or imported audio at 1080p.

Skills relacionados