Communitygithub.com

Zuhair-01/claude-skills-and-systems

Generate Mandarin and multilingual narration with Volcengine Doubao Speech 2.0. Use when creating Chinese voiceovers, when the user prefers Doubao/Volcengine/火山引擎/豆包 TTS, or when narration needs character-level timestamp metadata for subtitles.

O que é claude-skills-and-systems?

claude-skills-and-systems is a Claude Code agent skill that generate Mandarin and multilingual narration with Volcengine Doubao Speech 2.0. Use when creating Chinese voiceovers, when the user prefers Doubao/Volcengine/火山引擎/豆包 TTS, or when narration needs character-level timestamp metadata for subtitles.

Funciona com✓Claude Code~Codex CLI~Cursor
npx skills add https://github.com/Zuhair-01/claude-skills-and-systems/tree/HEAD/skills/doubao-tts

Perguntar na sua IA favorita

Abre um novo chat com esta habilidade de agente já pré-carregada.

Documentação

Doubao TTS

Requires DOUBAO_SPEECH_API_KEY in .env. Set DOUBAO_SPEECH_VOICE_TYPE for the default voice, or pass voice_id to the tool.

Current API

Use the new-console API key flow:

X-Api-Key: ${DOUBAO_SPEECH_API_KEY}
X-Api-Resource-Id: seed-tts-2.0

Do not use X-Api-App-Id and X-Api-Access-Key with a new-console API Key. If the API returns load grant: requested grant not found, the key type or auth header is probably wrong.

For long-form video narration, prefer the async endpoint:

POST https://openspeech.bytedance.com/api/v3/tts/submit
POST https://openspeech.bytedance.com/api/v3/tts/query

This returns audio_url plus sentences[].words[] timing metadata that can be used to build subtitles.

OpenMontage Usage

Generate with the TTS selector:

from tools.audio.tts_selector import TTSSelector

result = TTSSelector().execute({
    "preferred_provider": "doubao",
    "text": "如果 AI 真的会改变未来,普通人到底该怎么参与?",
    "voice_id": "zh_female_vv_uranus_bigtts",
    "output_path": "projects/my-video/assets/audio/narration.mp3",
    "speech_rate": 0,
    "enable_timestamp": True,
})

Or call the provider directly:

from tools.audio.doubao_tts import DoubaoTTS

result = DoubaoTTS().execute({
    "text": "短样本试听文本。",
    "voice_id": "zh_female_vv_uranus_bigtts",
    "output_path": "projects/my-video/assets/audio/doubao_sample.mp3",
})

The provider writes:

  • output_path: downloaded audio file
  • metadata_path: full query response JSON, defaulting to <output_path>.json

Recommended Workflow

  1. Generate a 10-15 second sample before a full paid narration.
  2. Ask the user to approve voice naturalness, accent, and speed.
  3. Generate the full narration only after approval.
  4. Keep the query JSON. It is the source of truth for subtitle timing.
  5. Build captions from sentences[].words[], not from estimated text length.
  6. Group captions by Chinese semantic phrases before applying timestamps. Do not split only by fixed character count; it can break phrases like "在不押单个公司的情况下" or "可能会被慢慢稀释" and hurt comprehension.
  7. Let the video duration follow the approved voice rhythm unless the user explicitly asks to match a prior runtime.

Parameters

  • voice_id: Doubao speaker / voice type. Defaults to DOUBAO_SPEECH_VOICE_TYPE.
  • resource_id: use seed-tts-2.0 for Doubao Speech 2.0 voices.
  • speech_rate: 0 is normal, 100 is 2x, -50 is 0.5x.
  • sample_rate: default 24000.
  • enable_timestamp: default true.
  • return_usage: default true, requests usage metadata when available.

Do not pass additions.explicit_language by default. Some endpoint/key combinations reject zh-cn with unsupported additions explicit language zh-cn.

For calm Mandarin explainers, start with speech_rate: 0. If the result is too long for the approved format, make a short comparison sample with speech_rate: 25 or 50 before regenerating the full narration. Do not speed up only to match a previous provider's duration if the user prefers Doubao's natural pace.

Troubleshooting

  • load grant: requested grant not found: wrong key type or wrong auth header. Use X-Api-Key for new-console API Keys.
  • speaker permission denied: voice id is wrong or not authorized for the selected resource.
  • quota exceeded: quota, lifetime characters, or concurrency exceeded.
  • Missing timestamps: verify enable_timestamp: true, keep the query JSON, and confirm the selected endpoint returned sentences.

Safety

Never print or write the API key to logs, metadata, patches, or project artifacts. .env.example should contain only empty variable names.

Individual skills in this repo

This repo contains 7 individual skills — each has its own dedicated page.

Zuhair-01/claude-skills-and-systems

Create AI avatar videos with precise control over avatars, voices, scripts, scenes, and backgrounds using HeyGen's v2 API. Use when: (1) Choosing a specific avatar and voice for a video, (2) Writing exact scripts for an avatar to speak, (3) Building multi-scene videos with different backgrounds per scene, (4) Creating transparent WebM videos for compositing, (5) Using talking photos as video presenters, (6) Integrating HeyGen avatars with Remotion, (7) Batch video generation with exact specs, (8) Brand-consistent production videos with precise control.

Zuhair-01/claude-skills-and-systems

Edit existing videos using AI — remix style, upscale, remove background, and add audio via fal.ai's hosted video models.

Zuhair-01/claude-skills-and-systems

Video and audio processing with FFmpeg. Use for format conversion, resizing, compression, audio extraction, and preparing assets for Remotion. Triggers include converting GIF to MP4, resizing video, extracting audio, compressing files, or any media transformation task.

Zuhair-01/claude-skills-and-systems

Author or edit a custom HyperFrames composition when no specialized workflow fits, or when BRIEF.md sets flow: companion. Use for longer or multi-scene pieces, brand and sizzle reels, montages, static loops, static title cards, footage remixes, and freeform builds. Use motion-graphics instead for a short unnarrated motion-first unit, including an animated title. Route fresh creation through hyperframes before using this skill.

Zuhair-01/claude-skills-and-systems

Generate HeyGen presenter videos via the v3 Video Agent pipeline — handles Frame Check (aspect ratio correction), prompt engineering, avatar resolution, and voice selection. Required for any HeyGen video generation. Replaces deprecated endpoints with v3. Use when: (1) generating any HeyGen video (via API or otherwise), (2) sending a personalized video message (outreach, update, announcement, pitch, knowledge), (3) creating a HeyGen presenter-led explainer, tutorial, or product demo with a human face, (4) "make a video of me saying...", "send a video to my leads", "record an update for my team", "create a video pitch", "make a loom-style message", "I want to appear in this video", "generate a HeyGen video", "make a talking head video". Accepts avatar_id from heygen-avatar for identity-first HeyGen videos, or uses a stock presenter. Returns video share URL + HeyGen session URL for iteration. Chain signal: when the user wants to create/design an avatar AND make a video in the same request, run heygen-avatar f...

Zuhair-01/claude-skills-and-systems

All animation knowledge for HyperFrames — atomic motion rules, multi-phase scene blueprints, scene transitions, broader motion-design techniques, AND the seven runtime adapters (GSAP default, plus Lottie, Three.js, Anime.js, CSS keyframes, Web Animations API, TypeGPU). Use for any motion or animation task: pick 2-4 rules and compose, or load a blueprint, or look up runtime-specific API (e.g. GSAP eases / Lottie player / Three.js mixer). Also covers auditing an existing composition's choreography (animation map) and 24 named text-animation effects. HyperFrames-native: single paused timeline, seek-safe, deterministic.

Zuhair-01/claude-skills-and-systems

The HyperFrames composition contract — build one renderable project. Use for composition structure, the `data-*` timing attributes, `class="clip"`, tracks, sub-compositions, variables, framework-owned media playback, deterministic-render rules, and validation. Read before writing composition HTML.

Habilidades Relacionadas