Communitygithub.com

delorenj/skills

Generate speech using the self-hosted voxxy (vox) TTS service at https://vox.delo.sh. Use when the user asks to speak, say, narrate, synthesize speech, clone a voice, create a voice, add or register a voice, pipe TTS, or control voice qualities by description (e.g. "a young woman with a cheerful voice"). Handles HTTP API usage, voice profile management, description-based voice design, cloning, MCP registration, and integration patterns for new platforms.

skills 是什么?

skills is a Claude Code agent skill that generate speech using the self-hosted voxxy (vox) TTS service at https://vox.delo.sh. Use when the user asks to speak, say, narrate, synthesize speech, clone a voice, create a voice, add or register a voice, pipe TTS, or control voice qualities by description (e.g. "a young woman with a cheerful voice"). Handles HTTP API usage, voice profile management, description-based voice design, cloning, MCP registration, and integration patterns for new platforms.

兼容平台~Claude Code~Codex CLI~Cursor
npx skills add https://github.com/delorenj/skills/tree/HEAD/vox-tts

在你喜欢的 AI 中提问

打开一个已预加载此 Agent Skill 的新对话。

文档

vox-tts

A self-hosted TTS service at https://vox.delo.sh wrapping VoxCPM2 with a postgres-backed voice profile store and support for multiple engines. Deployed at ~/code/voxxy (core: compose.yml, GPU engines: compose.engines.yml — both engine containers must be up or clones degrade silently, see below).

Quick reference

ActionHow
One-off synthesis (inline WAV bytes)POST /synthesize { text, voice?, cfg?, steps? } → audio/wav
Synthesis for delivery to Telegram / browser / HA / DiscordPOST /synthesize-url → {audio_url, engine, duration_s, bytes}
MCP tool: inline bytes (base64 WAV)speak(text, voice?)
MCP tool: delivery URL (OGG/Opus, Telegram-ready)speak_url(text, voice?)
MCP tool: list voiceslist_voices_tool()
List voices (HTTP)GET /voices
Add a voicePOST /voices (multipart: name, display_name, audio)
Interactive voice cloning (user says "clone my voice as…")See Interactive voice cloning workflow below
Upgrade an engine (VoxCPM / VibeVoice)See Engine upgrade workflow below
Speak from a Hermes agentNative tts.provider: vox (fleet base); the fleet does NOT load the vox MCP
Speak from a shell / any agent with a terminalvoxxy speak "text" (see voxxy speak --help)
Register with an MCP client that has no native vox pathMCP server at https://vox.delo.sh/mcp/ (trailing slash required)
Node-REDnode-red-contrib-vox at ~/docker/stacks/utils/vox/node-red-contrib-vox/
Health + engine statusGET /healthz

Trailing slash on /mcp/ is mandatory. Without it, FastAPI 307-redirects and HTTPX drops the POST body.

speak vs speak_url: pick the right one

If the audio will be...UseWhy
Sent to Telegram / Discord / Slackspeak_urlChannel APIs accept a URL; their servers fetch it. Zero byte-bloat on the agent wire.
Piped into a browser <audio> tagspeak_urlBrowsers stream URLs; no base64 round-trip.
Handed to Home Assistant media_player.play_mediaspeak_urlHA wants a URL for media_content_id.
Processed inline by the agent (splice, analyze, loop)speakBytes are already local; a URL fetch would add a hop.
Written to a local file in a shell scripteitherspeak_url + curl -o is easier than base64 + base64 -d.

Default to speak_url. It costs the agent nothing in tokens (the response is small JSON) and works across every delivery surface except raw inline byte processing.

Engine fallback

GET /healthz reports which engines are registered and whether each is available:

{
  "status": "ok",
  "model_loaded": true,
  "engines": [
    { "name": "voxcpm", "available": true },
    { "name": "elevenlabs", "available": true }
  ]
}

The orchestrator tries them in order. Every speak_url / speak response includes engine: "voxcpm" or engine: "elevenlabs" so you can detect when fallback engaged. ElevenLabs auto-disables when ELEVENLABS_API_KEY is unset.

Per-voice ElevenLabs mapping lives in the voices.elevenlabs_voice_id column. NULL falls back to the global default (ELEVENLABS_DEFAULT_VOICE, Adam by default).

Cloned voices are engine-bound. A voice cloned via POST /voices stores its reference in vibevoice_ref_path and is served by the vibevoice engine first (per-request preferred routing in app/engines.py; everything else follows VOX_ENGINES order, voxcpm first). If vibevoice is down, voxcpm still receives the reference clip — app/voices.py:46-48 falls back to wav_path for any engine with no specific override — so the clone degrades, it is not erased. Symptom is a recognisable-but-off read, not a stranger. Fix: docker compose -f ~/code/voxxy/compose.engines.yml up -d voxxy-engine-vibevoice, then confirm "name":"vibevoice","ready":true in /healthz.

The one that does replace the voice outright is ElevenLabs fallback: if both GPU engines are down and ELEVENLABS_API_KEY is set, vox answers 200 in a stranger's voice, logging vox: voice clone BYPASSED, and delivers it anyway. Check the engine field on every clone synthesis — vibevoice or voxcpm is the clone, elevenlabs is not your voice.

The reference budget is 10 seconds, not 30. Two different caps stack, and only the second one matters. POST /voices trims ingest to VOX_REF_AUDIO_MAX_SECONDS = 30s (compose.yml:16, applied at app/main.py:578-584, keeping the first 30s). But the vibevoice container sets the same variable to 10 (compose.engines.yml:69), and engines/vibevoice/engine/synth.py:183,208 reloads the reference with duration=10.0. Its /healthz says so directly: "max_ref_seconds": 10.0. So seconds 10–30 of any reference clip are stored, backed up, and never seen by the model. Write clone passages for 8–10 clean seconds and put the best, most neutral speech first — everything after is dead weight.

Detailed procedures

Read voice and integration workflows for the relevant implementation or diagnosis. Voice cloning and external delivery require the user to request those actions.

Defaults cheat sheet

ParamDefaultNotes
cfg2.0Classifier-free guidance; higher = more faithful, less variation
steps10Diffusion steps; 4-6 for speed, 15-20 for max quality
normalizefalseText normalization (numbers → words etc.)
denoisefalseDead field. Declared at app/main.py:98, read nowhere; POST /voices has no denoise param at all. Clean the audio before upload instead.
Cache TTL (audio URLs)3600sVOX_AUDIO_TTL_SECONDS env; 1h is plenty for Telegram
Fallback voice (ElevenLabs)Adam (pNInz6obpgDQGcFmaJgB)ELEVENLABS_DEFAULT_VOICE env

Synthesis runs at roughly 21 characters per second on an RTX 3090 with VOX_OPTIMIZE=1 — measured, not estimated: a 60-char line takes ~2s, a 1,665-char line takes ~35s. Budget by length, not by the 2s headline. First call after a container restart adds ~15s (JIT compile). OGG/Opus transcode adds <100ms via ffmpeg.

This matters for any caller with a timeout. The Hermes vox TTS plugin hardcodes a 60s httpx timeout (plugins/tts/vox/__init__.py:29,316) and ignores tts.vox.timeout, so a single chunk past roughly 2,500 characters fails outright. Cap it with tts.vox.max_text_length (read at tools/tts_tool.py:436-441), which chunks instead of failing.

Engine upgrade workflow

The three-container topology decouples core orchestration (voxxy-core, container vox, CPU-only) from GPU-bound sidecar engines (voxxy-engine-voxcpm, voxxy-engine-vibevoice). Each engine container maintains its own pyproject.toml, uv.lock, and Dockerfile under engines/<name>/.

Step-by-step upgrade procedure

1. Version discovery & runtime audit

  • Upstream releases: Check PyPI and GitHub releases:
    curl -s https://pypi.org/pypi/voxcpm/json | jq -r '.info.version'
    
  • Running container version: Check installed package version via importlib.metadata:
    docker exec voxxy-engine-voxcpm /opt/venv/bin/python -c \
      "import importlib.metadata; print(importlib.metadata.version('voxcpm'))"
    
    (Never use voxcpm.__version__; the module does not define it and raises AttributeError).
  • Model weights on Hugging Face: Check if the model weights snapshot has updated:
    cat ~/.cache/huggingface/hub/models--openbmb--VoxCPM2/refs/main
    curl -s https://huggingface.co/api/models/openbmb/VoxCPM2 | jq -r '.sha'
    
    Package updates (inference code, streaming VAE decoders, device handling) are distinct from model weight revisions.

2. Dependency bumping & lockfile update

  • Bump the version pin in engines/<name>/pyproject.toml (e.g. voxcpm==2.0.3).
  • Re-lock dependencies with uv:
    cd engines/voxcpm && uv lock
    
  • Review the diff (git diff engines/voxcpm/uv.lock) to ensure PyTorch and torchaudio stay pinned to the explicit pytorch-cu124 index and do not drift to PyPI CPU wheels.

3. Rebuild the engine container image

Rebuild the engine image via compose:

docker compose -f compose.yml -f compose.engines.yml build voxxy-engine-voxcpm
# Or via mise alias:
mise run build:voxcpm

The Dockerfile provisions Python 3.12 via uv and runs uv sync --frozen --no-install-project into /opt/venv.

4. Recreate the engine container with compose profile

Engine containers are defined with compose profiles in compose.engines.yml (profiles: ["voxcpm"] / profiles: ["vibevoice"]). You must pass --profile or compose will ignore the service:

docker compose -f compose.yml -f compose.engines.yml --profile voxcpm up -d --no-build voxxy-engine-voxcpm

Verify the container picked up the new version:

docker exec voxxy-engine-voxcpm /opt/venv/bin/python -c \
  "import importlib.metadata; print(importlib.metadata.version('voxcpm'))"

5. Verification & smoke testing

  • Container startup: Inspect logs for successful model load:
    docker logs --tail 30 voxxy-engine-voxcpm
    # Look for: "Loaded VoxCPM2Model" and "Application startup complete."
    
  • Engine healthz:
    docker exec vox curl -s http://voxxy-engine-voxcpm:8000/healthz
    # Expects: {"engine":"voxcpm","ready":true,"model_loaded":true,...}
    
  • Direct synthesis test: Send a raw synthesis payload directly from the vox core container over the Docker network:
    docker exec vox curl -s -X POST http://voxxy-engine-voxcpm:8000/v1/synthesize \
      -H 'content-type: application/json' \
      -d '{"text":"Testing synthesis."}' | jq -r '{engine, sample_rate, duration_s, bytes, has_wav: (.wav_b64 != null)}'
    
  • Voice cloning verification: Test with a reference audio sample (/data/voices/rick.wav):
    docker exec vox python3 -c '
    import base64, urllib.request, json
    with open("/data/voices/rick.wav", "rb") as f:
        b64 = base64.b64encode(f.read()).decode("ascii")
    data = json.dumps({"text": "Test clone.", "reference_audio_b64": b64}).encode("utf-8")
    req = urllib.request.Request("http://voxxy-engine-voxcpm:8000/v1/synthesize", data=data, headers={"Content-Type": "application/json"})
    print(json.loads(urllib.request.urlopen(req).read().decode("utf-8")))
    '
    
  • Contract and CLI tests:
    bash scripts/verify-engine-contract.sh
    (cd cli && uv run pytest)
    

6. Delivery

Commit both pyproject.toml and uv.lock in the engine directory:

git add engines/voxcpm/pyproject.toml engines/voxcpm/uv.lock
git commit -m "feat(voxcpm): upgrade voxcpm to <version>"
git push origin main

Lessons learned & engine gotchas

  1. voxcpm package version inspection: voxcpm does NOT expose __version__. Calling python -c "import voxcpm; print(voxcpm.__version__)" throws AttributeError. Always inspect installed metadata with importlib.metadata.version('voxcpm').
  2. Compose profiles requirement: In compose.engines.yml, each engine has a profile (profiles: ["voxcpm"]). If you run docker compose up voxxy-engine-voxcpm without --profile voxcpm, Docker Compose will silently skip the service.
  3. Standby engines are directly testable: Local engines on a single GPU card are mutually exclusive (VoxCPM ~5 GB VRAM, VibeVoice ~7.5 GB VRAM). Even if an engine is in standby in voxxy daemon status / .voxxy.state.json, its container remains running and healthy on the internal Docker network (http://voxxy-engine-<name>:8000). You can test it directly via curl from vox without switching the active engine.
  4. Pytest test suite location: Running pytest from repo root fails because plugins/tts/vox/ is a Hermes plugin expecting agent.tts_provider. The CLI and integration test suite lives in cli/tests/ and should be executed via cd cli && uv run pytest.
  5. Weights cache persistence: Engine containers bind-mount /home/delorenj/.cache/huggingface to /cache/huggingface. Never change this to an ephemeral volume, or image rebuilds will re-download gigabytes of weights on startup.

Individual skills in this repo

This repo contains 14 individual skills — each has its own dedicated page.

delorenj/skills

Plan and generate terminal ASCII animations/screensaver-style output (FPS, refresh rules, loop policy, low-flicker guidance), with a static poster frame and an optional local demo script.

delorenj/skills

ASCII video: convert video/audio to colored ASCII MP4/GIF.

delorenj/skills

Plan, build and quality-check a premium short commercial for a real business (roofing, property, ecommerce, B2B software) using AI-generated footage, code-built motion (HTML/GSAP), selective Three.js and an independent-critic "Gauntlet" loop. Use when asked to make a launch-style / SaaS-style video, business ad, explainer, sample reel or pitch video, or to review and improve one. Encodes motion principles distilled from 28 professional launch films, a quality bar, audio rules and a business-offer playbook.

delorenj/skills

CSS animation adapter patterns for HyperFrames. Use when authoring CSS keyframes, animation-delay based timing, animation-fill-mode, animation-play-state, or CSS-only motion that HyperFrames must seek deterministically during preview and rendering.

delorenj/skills

Generate professional voiceovers using ElevenLabs AI. Use when the user needs to create voiceovers for videos, audio narration, or text-to-speech content. Supports multiple voices with character presets (narrator, salesperson, expert) for natural delivery. Includes single scene regeneration for fine-tuning.

delorenj/skills

HyperFrames CLI dev loop — `npx hyperframes` for scaffolding (init), validation (lint, inspect), preview, render, and environment troubleshooting (doctor, browser, info, upgrade). Use when running any of these commands or troubleshooting the HyperFrames build/render environment. For asset preprocessing commands (`tts`, `transcribe`, `remove-background`), invoke the `hyperframes-media` skill instead.

delorenj/skills

Asset preprocessing for HyperFrames compositions — text-to-speech narration (Kokoro), audio/video transcription (Whisper), and background removal for transparent overlays (u2net). Use when generating voiceover from text, transcribing speech for captions, removing the background from a video or image to use as a transparent overlay, choosing a TTS voice or whisper model, or chaining these (TTS → transcribe → captions). Each command downloads its own model on first run.

delorenj/skills

Install and wire registry blocks and components into HyperFrames compositions. Use when running hyperframes add, installing a block or component, wiring an installed item into index.html, or working with hyperframes.json. Covers the add command, install locations, block sub-composition wiring, component snippet merging, and registry discovery.

delorenj/skills

Create video compositions, animations, title cards, overlays, captions, voiceovers, audio-reactive visuals, and scene transitions in HyperFrames HTML. Use when asked to build any HTML-based video content, add captions or subtitles synced to audio, generate text-to-speech narration, create audio-reactive animation (beat sync, glow, pulse driven by music), add animated text highlighting (marker sweeps, hand-drawn circles, burst lines, scribble, sketchout), or add transitions between scenes (crossfades, wipes, reveals, shader transitions). Covers composition authoring, timing, media, and the full video production workflow. For dev-loop CLI commands (init, lint, inspect, preview, render) see the hyperframes-cli skill; for asset preprocessing commands (tts, transcribe, remove-background) see the hyperframes-media skill.

delorenj/skills

Manim CE animations: 3Blue1Brown math/algo videos.

delorenj/skills

Join a Google Meet or Zoom call as a video meeting agent via PikaStreaming. Trigger: user drops a Google Meet or Zoom link, or asks to join a meeting.

delorenj/skills

Translate an existing Remotion (React-based) video composition into a HyperFrames HTML composition. Use ONLY when the user explicitly asks to port, convert, migrate, translate, or rewrite a Remotion composition as HyperFrames (e.g. "port my Remotion project to HyperFrames"). Do NOT use when (a) authoring a NEW HyperFrames composition (even if A/B-testing a Remotion video); (b) Remotion is mentioned in passing; (c) Remotion code is shared as reference, not for translation; (d) the user wants "the same video as my Remotion one" without explicitly asking to migrate the source — treat as a fresh HyperFrames build. When in doubt, default to the `hyperframes` skill. Detects unsupported patterns (useState, useEffect side effects, async calculateMetadata, third-party React component libraries, `@remotion/lambda`) and recommends the runtime interop escape hatch instead of a lossy translation.

delorenj/skills

Capture a website and create a HyperFrames video from it. Use when: (1) a user provides a URL and wants a video, (2) someone says "capture this site", "turn this into a video", "make a promo from my site", (3) the user wants a social ad, product tour, or any video based on an existing website, (4) the user shares a link and asks for any kind of video content. Even if the user just pastes a URL — this is the skill to use.

delorenj/skills

YouTube transcripts to summaries, threads, blogs.

相关技能