Communitygithub.com

elevenlabs/speech-to-text

使用 ElevenLabs Scribe v2 將音頻轉錄為文字。適用於將音頻/視頻轉換為文字、生成字幕、轉錄會議或處理語音內容。

speech-to-text 是什麼?

這個技能利用 ElevenLabs Scribe v2 模型高效地將音頻或視頻內容轉錄為文字。它支援多種場景,如生成準確的會議記錄、自動建立影片字幕、以及處理播客或錄音中的語音內容。該技能適用於 Claude Code、Cursor 和 Codex 等 AI 代理平台,幫助開發者快速從語音中提取資訊。透過直接呼叫 ElevenLabs 的 API,它能夠處理長時間音頻並保持高精度,同時支援多種語言。對於需要將非結構化語音資料轉化為可搜尋或可編輯文字的工作流程,這個技能能顯著提升效率和自動化程度。

相容平台~Claude Code~Codex CLI~Cursor
npx skills add https://github.com/elevenlabs/skills/tree/main/skills/speech-to-text

在你喜歡的 AI 中提問

開啟一個已預先載入此 Agent Skill 的新對話。

說明文件

speech-to-text 是做什麼的?

Transcribe audio to text using ElevenLabs Scribe v2. Use when converting audio/video to text, generating subtitles, transcribing meetings, or processing spoken content.

Individual skills in this repo

This repo contains 7 individual skills — each has its own dedicated page.

elevenlabs/agents

Build voice AI agents with ElevenLabs. Use when creating voice assistants, customer service bots, interactive voice characters, or any real-time voice conversation experience.

elevenlabs/music

Generate music using ElevenLabs Music API. Use when creating instrumental tracks, songs with lyrics, background music, jingles, or any AI-generated music composition. Supports prompt-based generation, composition plans for granular control, and detailed output with metadata.

elevenlabs/setup-api-key

Guides users through setting up an ElevenLabs API key for ElevenLabs MCP tools. Use when the user needs to configure an ElevenLabs API key, when ElevenLabs tools fail due to missing API key, or when the user mentions needing access to ElevenLabs. First checks whether ELEVENLABS_API_KEY is already configured and valid, and only runs full setup when needed.

elevenlabs/sound-effects

Generate sound effects from text descriptions using ElevenLabs. Use when creating sound effects, generating audio textures, producing ambient sounds, cinematic impacts, UI sounds, or any audio that isn't speech. Supports looping, duration control, and prompt influence tuning.

elevenlabs/text-to-speech

Convert text to speech using ElevenLabs voice AI. Use when generating audio from text, creating voiceovers, building voice apps, or synthesizing speech in 70+ languages.

elevenlabs/voice-changer

Transform the voice in an audio recording into a different target voice while preserving emotion, timing, and delivery using the ElevenLabs Voice Changer (speech-to-speech) API. Use when converting one voice to another, changing the speaker/narrator of an existing recording, dubbing a voice-over in a different voice, creating character voices from a scratch performance, anonymizing a speaker, or any "voice conversion / voice transfer / speech-to-speech" task. Make sure to use this skill whenever the user mentions voice changing, voice conversion, speech-to-speech, swapping a voice in audio, re-voicing a clip, or applying a different voice to an existing recording — even if they don't explicitly say "voice changer".

elevenlabs/voice-isolator

Remove background noise and isolate vocals/speech from audio using ElevenLabs Voice Isolator (audio isolation) API. Use when cleaning up noisy recordings, removing music or background ambience from dialogue, isolating speech from field recordings, preparing audio for transcription, extracting vocals, or any "denoise / clean up / isolate voice" task.

相關技能