Communitygithub.com

broomva/skills

Local TTS, voice cloning, voice design, and video dubbing via the OmniVoice Studio MCP server (open-source ElevenLabs alternative; nothing leaves the machine, runs on MPS/CUDA/CPU). Use when: (1) generating speech from text in any of 646 languages, (2) cloning a voice from a 3-second reference clip, (3) designing a voice by gender/age/accent/pitch/style, (4) dubbing a video into another language, (5) listing voice profiles or personality presets, (6) producing narration where privacy, cost, or absent API keys matter, (7) non-English narration where Edge TTS/kokoro fall short, (8) batch audio for blog posts or content pipelines. Triggers: 'omnivoice', 'voice clone', 'clone this voice', 'tts', 'narrate', 'generate speech', 'voice synthesis', 'dub video', 'voice design', 'local tts', 'multilingual voice', 'narrate this post', 'elevenlabs alternative'.

skills とは?

skills is a Antigravity agent skill that local TTS, voice cloning, voice design, and video dubbing via the OmniVoice Studio MCP server (open-source ElevenLabs alternative; nothing leaves the machine, runs on MPS/CUDA/CPU). Use when: (1) generating speech from text in any of 646 languages, (2) cloning a voice from a 3-second reference clip, (3) designing a voice by gender/age/accent/pitch/style, (4) dubbing a video into another language, (5) listing voice profiles or personality presets, (6) producing narration where privacy, cost, or absent API keys matter, (7) non-English narration where Edge TTS/kokoro fall short, (8) batch audio for blog posts or content pipelines. Triggers: 'omnivoice', 'voice clone', 'clone this voice', 'tts', 'narrate', 'generate speech', 'voice synthesis', 'dub video', 'voice design', 'local tts', 'multilingual voice', 'narrate this post', 'elevenlabs alternative'.

対応~Claude Code~Codex CLI~Cursor✓Antigravity
npx skills add https://github.com/broomva/skills/tree/HEAD/skills/audio/omnivoice

お気に入りのAIに質問する

このエージェントスキルを事前に読み込んだ状態で新しいチャットを開きます。

ドキュメント

OmniVoice

Overview

Generate audio locally via the OmniVoice Studio MCP server. Tools: generate_speech, list_voices, list_personalities, list_languages, check_health. Resources: voice://{id}, history://recent.

Prerequisites — Backend Must Be Running

The MCP tools all hit $OMNIVOICE_API_URL (default http://localhost:3900). If the backend is down, every tool returns a connection error. Install + boot:

git clone https://github.com/debpalash/OmniVoice-Studio.git "$OMNIVOICE_HOME"
cd "$OMNIVOICE_HOME"
uv sync
VIRTUAL_ENV="$(pwd)/.venv" uv pip install 'mcp[cli]'

Then:

scripts/check-health.sh        # exit 0 if up
scripts/start-backend.sh       # boot in background (MPS/CUDA auto-detected)

First synthesis call lazy-downloads the k2-fsa/OmniVoice model (~2.4 GB) from HuggingFace — cached on subsequent boots.

Task Index — Pick the Right Tool

TaskToolNotes
Verify backend is upcheck_healthReturns `{"status":"ok","device":"mps
Text → audio with a saved voicegenerate_speech(text, profile_id)Returns base64 WAV. profile_id="demo0001" is the bundled demo voice
Text → audio without a clone (voice design)generate_speech(text, instruct="…")Omit profile_id; pass an instruct like "warm middle-aged female narrator, calm pace"
Multilingual narrationgenerate_speech(text, language="es")Any ISO 639 code or "Auto"
List existing voiceslist_voicesReturns id, name, type, personality
List personality presetslist_personalitiesReturns narrator / casual / news-anchor / etc. with their instruct strings
List supported languageslist_languages646 total; returns 20 popular + the full count

For non-trivial decisions (which engine to use, when to pick OmniVoice over kokoro / Edge TTS / ElevenLabs), see references/engines-comparison.md.

For MCP wiring details, backend lifecycle, troubleshooting, and a clean teardown, see references/mcp-setup.md.

Common Workflows

1. One-shot narration with the demo voice

# As called through the MCP client (your agent will do this for you):
result = generate_speech(
    text="Hello — this is OmniVoice generating speech locally.",
    profile_id="demo0001",
    language="English",
    steps=16,                   # 8 = fast/draft · 16 = balanced · 32 = quality
)
# result is JSON with audio_id, generation_time_s, audio_duration_s, format, wav_base64

Benchmark: 4.2 s of audio in ~24 s server-side on Apple Silicon MPS at 16 diffusion steps.

2. Save the WAV to disk and play

Tool returns base64 PCM WAV (16-bit, mono, 24 kHz). Decode + write:

import base64, json
payload = json.loads(result_text)            # parse JSON the tool returns
open("out.wav","wb").write(base64.b64decode(payload["wav_base64"]))

On macOS: afplay out.wav. Convert to MP3 with ffmpeg -i out.wav -codec:a libmp3lame -b:a 128k out.mp3.

3. Voice clone — end-to-end recipe

Cloning needs a 3-10 second reference clip the model will use as a speaker embedding. The MCP server does NOT expose profile creation — it only reads existing profiles. Two paths to create one:

Path A — bundled helper (macOS, recommended for fresh clones):

scripts/record-reference.sh ~/Downloads/my-ref.wav 12 1
# args: output_path raw_duration_sec mic_index
# Default mic_index=1 (MacBook built-in); list devices via:
#   ffmpeg -f avfoundation -list_devices true -i ""

The script gives audible countdown + start/stop cues via macOS say + /System/Library/Sounds/Ping.aiff so the user knows when to speak (terminal stdout is buffered — text "speak now" prompts arrive too late). It records a longer raw window, then trims to ~10 seconds of speech via silenceremove + atrim, plays back for verification, and prints the next-step curl command.

Path B — manual:

# 1. Record (mono, 24 kHz native — matches model's internal rate)
ffmpeg -f avfoundation -i ":1" -t 12 -ac 1 -ar 24000 raw.wav

# 2. Trim leading silence + take first 10 sec of speech
ffmpeg -i raw.wav \
  -af "silenceremove=start_periods=1:start_silence=0.05:start_threshold=-40dB,atrim=end=10" \
  -ac 1 -ar 24000 ref.wav

# 3. Verify
ffmpeg -i ref.wav -af volumedetect -f null - 2>&1 | grep volume   # max should be > -20 dB
afplay ref.wav

POST to /profiles (multipart/form-data — required fields: name, ref_audio):

curl -X POST http://127.0.0.1:3900/profiles \
  -F "name=carlos-clone" \
  -F "[email protected]" \
  -F "ref_text=The exact text spoken in the clip" \
  -F "language=English" \
  | python3 -m json.tool
# returns { "id": "abc12345", "name": "carlos-clone" }

Once created, pass profile_id to generate_speech (via MCP) or directly via POST /generate. Profiles persist in SQLite + reference-audio files at ~/Library/Application Support/OmniVoice/voices/<id>.<ext> (the backend preserves the uploaded extension — .wav if you uploaded a WAV, .mp3 if MP3, etc.). State persists across backend restarts.

Reference clip tips that materially affect quality:

FactorWhy it matters
Single speakerMixed speakers blur the embedding
Clean speech, no music/noiseModel embeds the noise too
Natural prosody (avoid pangrams)Diffusion samples replicate prosody, not just timbre
3-10 sec is the sweet spot< 3 s lacks information; > 10 s adds compute without quality gain
Match ref_text to what's spokenImproves alignment, especially on noisy refs
language correctWrong language → cross-lingual transfer artifacts
Loudness peak ≥ -15 dBQuiet refs work but normalize poorly

4. Voice design (no reference clip)

Skip profile_id; provide an instruct string describing the desired voice:

generate_speech(
    text="Welcome to the future of agentic systems.",
    instruct="warm middle-aged female narrator, calm authoritative pace, documentary style",
)

Get pre-made instructs via list_personalities and copy the one matching the brief (narrator, casual, news-anchor, etc.).

5. Video dubbing (web UI only)

The MCP server does not expose the dubbing endpoint. The full transcribe → translate → re-voice → mux pipeline lives behind the desktop UI (bun run desktop in $OMNIVOICE_HOME) and the /dub/* REST routes. When the user asks to dub a video, point them to the UI; surface this skill only for the synthesis primitives above.

When NOT to use OmniVoice

  • Fast English-only narration on weak hardware → kokoro-tts is ~10× smaller and 2× realtime on CPU (see references/engines-comparison.md)
  • Lowest-friction one-off TTS → Edge TTS needs no install or backend
  • Highest possible quality regardless of cost → ElevenLabs still wins on English narration polish; OmniVoice ties or wins on multilingual + cloning
  • Real-time streaming dictation → use the OmniVoice desktop widget (⌘+⇧+Space), not the MCP server

Resources

Backend Swagger / OpenAPI: http://127.0.0.1:3900/docs (when backend is up).

Upstream: github.com/debpalash/OmniVoice-Studio — FSL-1.1-ALv2 (free for personal/internal/non-commercial; auto-converts to Apache-2.0 two years after each release).

Individual skills in this repo

This repo contains 7 individual skills — each has its own dedicated page.

broomva/skills

>- Speak an explanation out loud while working in any project — tiered text-to-speech with a pluggable backend (ElevenLabs by default and quota-guarded, macOS `say` via `--fast` for free instant local speech, local OmniVoice as an unlimited private tier). Markdown-aware, so code fences, URLs and deep paths collapse to short spoken placeholders instead of being dictated character by character, while snake_case identifiers survive intact so the listener can still search for them. Every utterance is saved to disk for later replay. Also carries a **talk mode** toggle: turn it on and the agent speaks a full readback of every turn, for as long as that session lasts — the whole response, not a summary of it, with `brief` and `marker` levels for when you want less. Talk mode is off by default and scoped to the single session that enabled it, so parallel agents in other worktrees stay silent. Use when the user asks to hear something rather than read it — an explanation of a change, a walkthrough of what just happen...

broomva/skills

Shop Tiendas D1 (Colombia, d1.com.co) from the command line — search the catalogue, resolve your nearest physical store, price a basket against that store's real stock, and quote delivery. D1 runs VTEX IO (account `d1tiendas`), so this drives its public storefront API with no admin key at all — catalogue and cart work fully anonymously, and a one-time emailed code unlocks order history. Handles the two traps that make naive D1 automation wrong — availability is regionalized (an unregioned query reports a national catalogue nobody can actually buy from) and prices arrive in two different units (search reports whole pesos, checkout reports hundredths, a silent 100x). Builds and prices baskets; it deliberately cannot pay, handing a checkout URL to a human instead. USE WHEN the user wants to find D1 products or prices, check whether D1 delivers somewhere, build or cost a D1 grocery basket, compare D1 items, or review their D1 orders. NOT FOR other Colombian retailers (Éxito, Jumbo, Ara, Alkosto), and not for c...

broomva/skills

Stateful, local-first household toxics inventory + swap engine. Identify the items in a home that carry endocrine disruptors and persistent chemicals (BPA/BPS, phthalates, PFAS/PTFE, parabens, flame retardants, VOCs, microplastics), score each by *real* exposure (severity x presence x how it's used x condition), and track the swap to a safer alternative from "flagged" -> "sourced" -> "swapped". Ships a grounded, cited knowledge graph of ~20 hazards, ~40 item-classes, and ~40 alternatives. Hands sourcing off to the `procurer` skill. The skill's state is the source of truth — the agent is the app.

broomva/skills

Tekton — the shared architecture-intent substrate for co-designing systems with the agent. One typed graph across six tiers (system / journey / data / infra / decisions / qualities); views are queries, not separate diagrams. The canonical artifact is a diff-friendly YAML model both human and agent read and write; it renders to Mermaid (agent-legible, GitHub-native) and a self-contained tabbed HTML viewer (human-visual). Cross-tier traceability (`tekton query <from> <to>`) answers "which infra does this user-journey step touch?" as a path query. USE WHEN: designing or thinking deeply about architecture, a system, a data model, user journeys/flows, or a technical plan WITH the agent; when a Category-C HTML doc isn't enough because you need to see AND edit AND traverse the design across tiers; "let's design X", "architect this", "model the system", "draw the flow", "how does this fit together", "diagram this", "/tekton". NOT FOR: a one-off throwaway diagram (use Mermaid inline); prose-only ADRs (write the ADR...

broomva/skills

Generate a polished Remotion video and X thread showcasing the full agent skills inventory. Use when creating social media content about skills, rendering category-based skill visualizations, or producing animated showcases of agent capabilities. Triggers on "skills video", "showcase skills", "skills thread", "render skills", "social content for skills", or requests to visualize the skills inventory.

broomva/skills

OpenCaptions extension for Content Engine — adds intent-driven CWI (Caption With Intent) captions to the post-production pipeline. Hooks into the grade → caption → final stage. Generates captions that understand video intent (pitch, volume, emotion, emphasis) and style themselves accordingly with variable font weight, size, and color. Uses the OpenCaptions CLI or MCP server. Triggers on: 'add captions', 'opencaptions', 'CWI captions', 'intent captions'.

broomva/skills

Produce polished product launch videos using the Liquid Glass aesthetic — dark void backgrounds, 3D perspective floating UI panels, particle effects, spring animations, and cinematic pacing. Built on Remotion with Imagen 4.0 for frames and Veo 3.1 for B-roll. Use when: (1) creating a product demo or launch video, (2) showcasing a UI/app/tool with cinematic polish, (3) building a social-ready video from screenshots and renders, (4) applying the liquid glass floating panel style, (5) composing Remotion videos with 3D transforms and spring animations. Triggers on: 'launch video', 'product video', 'liquid glass video', 'demo video', 'showcase video', 'remotion video', 'floating panel', 'glass aesthetic'.

関連スキル