Communitygithub.com

David-Li0406/meta-skill-evloving

Streams Edge TTS audio in real-time using PyAudio's callback mechanism to eliminate gaps, handling MP3 to PCM conversion and queue-based buffering.

What is meta-skill-evloving?

meta-skill-evloving is a Claude Code agent skill that streams Edge TTS audio in real-time using PyAudio's callback mechanism to eliminate gaps, handling MP3 to PCM conversion and queue-based buffering.

Works with~Claude Code~Codex CLI~Cursor
npx skills add https://github.com/David-Li0406/meta-skill-evloving/tree/HEAD/AutoSkill/SkillBank/ConvSkill/english_gpt4_8/edge_tts_pyaudio_gapless_streaming

Ask in your favorite AI

Open a new chat with this agent skill pre-loaded.

Documentation

edge_tts_pyaudio_gapless_streaming

Streams Edge TTS audio in real-time using PyAudio's callback mechanism to eliminate gaps, handling MP3 to PCM conversion and queue-based buffering.

Prompt

Role & Objective

You are a Python Audio Engineer. Your task is to implement real-time Text-to-Speech (TTS) streaming using edge_tts and pyaudio. The primary goal is to eliminate audio gaps and popping by employing a callback-based playback mechanism.

Operational Rules & Constraints

  1. Architecture: You MUST use a stream_callback mechanism with PyAudio (as opposed to blocking stream.write() calls) to ensure continuous playback and eliminate voids between audio chunks.
  2. TTS Streaming: Use edge_tts.Communicate to generate audio. Iterate over communicate.stream() to retrieve audio chunks.
  3. Format Conversion: The incoming data from edge_tts is MP3. You must convert these chunks to PCM/WAV format using pydub.AudioSegment before they can be played by pyaudio.
  4. Buffering: Implement a buffer (e.g., queue.Queue) to hold the converted PCM data. The callback function should read from this buffer to feed the audio stream continuously.
  5. Concurrency: Handle the asynchronous nature of edge_tts alongside the synchronous PyAudio callback. Use asyncio and threading to keep the buffer filled without blocking the audio playback.
  6. Error Handling:
    • Initialize pcm_data to None or b'' at the start of conversion functions to prevent UnboundLocalError.
    • Handle exceptions during MP3 to PCM conversion (log error, set data to empty bytes to prevent crashes).
    • Handle IOError (buffer underrun) inside the callback by logging warnings and returning silence or pausing.

Communication & Style Preferences

  • Provide complete, runnable code snippets for the integration logic.
  • Ensure imports (asyncio, edge_tts, queue, pyaudio, pydub, io, logging, threading) are included.
  • Use clear comments explaining the data flow from TTS generation to the audio callback.

Anti-Patterns

  • Do not use blocking stream.write() calls inside the main loop if they cause gaps.
  • Do not save the audio to a file before playing; it must be streamed.
  • Do not ignore the requirement to convert MP3 chunks to PCM/WAV.
  • Do not assume specific sample rates or bit depths; use the values defined in the audio configuration.
  • Do not modify the internal logic of core audio classes unless explicitly requested.

Triggers

  • stream edge_tts with pyaudio
  • fix audio gaps in python tts
  • real-time text to speech streaming
  • pyaudio callback for tts
  • edge_tts pyaudio player

Individual skills in this repo

This repo contains 20 individual skills — each has its own dedicated page.

David-Li0406/meta-skill-evloving

Browser automation with persistent page state. Use when users ask to navigate websites, fill forms, take screenshots, extract web data, test web apps, or automate browser workflows. Trigger phrases include "go to [url]", "click on", "fill out the form", "take a screenshot", "scrape", "automate", "test the website", "log into", or any browser interaction request.

David-Li0406/meta-skill-evloving

Create technical diagrams using Mermaid syntax for architecture, sequences, ERDs, flowcharts, and state machines. Use for visualizing system design, data flows, and processes. Triggers: diagram, mermaid, architecture diagram, sequence diagram, flowchart, ERD, entity relationship, state diagram, C4 model, component diagram, visualize, draw.

David-Li0406/meta-skill-evloving

Comprehensive document creation, editing, and analysis with support for tracked changes, comments, formatting preservation, and text extraction. When Claude needs to work with professional documents (.docx files) for: (1) Creating new documents, (2) Modifying or editing content, (3) Working with tracked changes, (4) Adding comments, or any other document tasks

David-Li0406/meta-skill-evloving

Complete color manipulation and green screen effects system. PROACTIVELY activate for: (1) Green screen/chromakey removal, (2) Color grading with LUTs, (3) Color correction (levels, curves, white balance), (4) Colorkey/color removal effects, (5) Hue/saturation manipulation, (6) Color space conversions (BT.709, BT.2020, HDR), (7) Color isolation effects, (8) Teal and orange look, (9) Vintage/film looks, (10) Color keying for transparency. Provides: chromakey and colorkey filters, LUT application (lut3d), curves and levels adjustment, color balance, selective color manipulation, color space handling, HDR tone mapping, professional color grading chains.

David-Li0406/meta-skill-evloving

Convert files and office documents to Markdown. Supports PDF, DOCX, PPTX, XLSX, images (with OCR), audio (with transcription), HTML, CSV, JSON, XML, ZIP, YouTube URLs, EPubs and more.

David-Li0406/meta-skill-evloving

Guide for creating high-quality MCP (Model Context Protocol) servers that enable LLMs to interact with external services through well-designed tools. Use when building MCP servers to integrate external APIs or services, whether in Python (FastMCP) or Node/TypeScript (MCP SDK).

David-Li0406/meta-skill-evloving

Build with OpenAI's stateless APIs - Chat Completions (GPT-5, GPT-4o), Embeddings, Images (DALL-E 3), Audio (Whisper + TTS), and Moderation. Includes Node.js SDK and fetch-based approaches for Cloudflare Workers. Use when: implementing chat completions with GPT-5/GPT-4o, streaming responses with SSE, using function calling/tools, creating structured outputs with JSON schemas, generating embeddings for RAG (text-embedding-3-small/large), generating images with DALL-E 3, editing images with GPT-Image-1, transcribing audio with Whisper, synthesizing speech with TTS (11 voices), moderating content (11 safety categories), or troubleshooting rate limits (429), invalid API keys (401), function calling failures, streaming parse errors, embeddings dimension mismatches, or token limit exceeded.

David-Li0406/meta-skill-evloving

Comprehensive PDF manipulation toolkit for extracting text and tables, creating new PDFs, merging/splitting documents, and handling forms. When Claude needs to fill in a PDF form or programmatically process, generate, or analyze PDF documents at scale.

David-Li0406/meta-skill-evloving

Presentation creation, editing, and analysis. When Claude needs to work with presentations (.pptx files) for: (1) Creating new presentations, (2) Modifying or editing content, (3) Working with layouts, (4) Adding comments or speaker notes, or any other presentation tasks

David-Li0406/meta-skill-evloving

Comprehensive spreadsheet creation, editing, and analysis with support for formulas, formatting, data analysis, and visualization. When Claude needs to work with spreadsheets (.xlsx, .xlsm, .csv, .tsv, etc) for: (1) Creating new spreadsheets with formulas and formatting, (2) Reading or analyzing data, (3) Modify existing spreadsheets while preserving formulas, (4) Data analysis and visualization in spreadsheets, or (5) Recalculating formulas

David-Li0406/meta-skill-evloving

Download YouTube videos with customizable quality and format options. Use this skill when the user asks to download, save, or grab YouTube videos. Supports various quality settings (best, 1080p, 720p, 480p, 360p), multiple formats (mp4, webm, mkv), and audio-only downloads as MP3.

David-Li0406/meta-skill-evloving

Best practices for Remotion - Video creation in React

David-Li0406/meta-skill-evloving

Provides bidirectional translation (Chinese/English), grammar checking, and sentence polishing. Supports specific tense constraints, context-aware tone adjustment, and natural expression optimization. Outputs direct text without formatting.

David-Li0406/meta-skill-evloving

指导如何通过自定义注解(如@Unicom)标识Dubbo接口,利用BeanDefinitionRegistryPostProcessor在启动时扫描并动态注册ReferenceBean,实现RPC调用的封装。

David-Li0406/meta-skill-evloving

生成Python脚本,利用FFmpeg将图片序列按3x3布局合并,或将合并图拆分。要求使用subprocess模块执行命令,并支持用户交互式输入路径。

David-Li0406/meta-skill-evloving

SOP for extracting and analyzing promotional activity details, handling time formatting, location mapping, payment logic, and step-by-step process validation.

David-Li0406/meta-skill-evloving

Specializes in creating TikTok short video scripts using an anthropomorphic 'Chef' perspective to explain wildlife food chains and survival environments, utilizing real field footage.

David-Li0406/meta-skill-evloving

针对TikTok平台,以拟人化视角(如动物作为厨师)创作系列短视频脚本。重点介绍野生动物、食物链及生态环境,要求使用野外实拍素材,风格幽默且具有教育意义。

David-Li0406/meta-skill-evloving

Generates KWL (Know, Want to know, Learned) charts for educational videos, specifically BrainPOP, using simple 7th-grade language and short, concise bullet points.

David-Li0406/meta-skill-evloving

Tracks complex, multi-session work using the Beads issue tracker and dependency graphs, and provides persistent memory that survives conversation compaction. Use when work spans multiple sessions, has complex dependencies, or needs persistent context across compaction cycles. Trigger with phrases like "create task for", "what's ready to work on", "show task", "track this work", "what's blocking", or "update status".

Related Skills