by sickn33
使用异步工作器管道、流式转录、LLM 代理和 TTS 合成功能创建生产级对话式 AI 语音引擎。本技能提供完整的语音助手构建指南,包括中断处理和多云服务商支持。
1. 打开 Claude 聊天界面
2. 点击下方 "📋 复制" 按钮
3. 粘贴到 Claude 聊天框中并发送
4. 输入 "使用 voice-ai-engine-development 技能" 开始使用
=== voice-ai-engine-development 技能 === 作者: sickn33 描述: 使用异步工作器管道、流式转录、LLM 代理和 TTS 合成功能创建生产级对话式 AI 语音引擎。本技能提供完整的语音助手构建指南,包括中断处理和多云服务商支持。 使用方法: 1. 调用技能: "使用 voice-ai-engine-development 技能" 2. 提供相关信息: 根据技能要求提供必要参数 3. 查看结果: 技能会返回处理结果 示例: "使用 voice-ai-engine-development 技能,帮我分析一下这段代码"
这种方法适用于所有 Claude 用户,不需要安装额外工具。
coding
safe
Build production-ready real-time conversational AI voice engines with async worker pipelines, streaming transcription, LLM agents, and TTS synthesis.
This skill provides comprehensive guidance for building voice AI engines that enable natural, bidirectional conversations between users and AI agents. It covers the complete architecture from audio input to audio output, including:
# Use the skill in your AI assistant
@voice-ai-engine-development I need to build a voice assistant that can handle real-time conversations with interrupts
SKILL.md - Comprehensive guide to voice AI engine developmentcomplete_voice_engine.py - Full working implementationgemini_agent_example.py - LLM agent with proper response bufferinginterrupt_system_example.py - Interrupt handling demonstrationbase_worker_template.py - Template for creating new workersmulti_provider_factory_template.py - Multi-provider factory patterncommon_pitfalls.md - Common issues and solutionsprovider_comparison.md - Comparison of transcription, LLM, and TTS providersEvery voice AI engine follows this pipeline:
Audio In → Transcriber → Agent → Synthesizer → Audio Out
(Worker 1) (Worker 2) (Worker 3)
Each worker:
class BaseWorker:
async def _run_loop(self):
while self.active:
item = await self.input_queue.get()
await self.process(item)
# User interrupts bot mid-sentence
if stop_event.is_set():
partial_message = get_message_up_to(seconds_spoken)
return partial_message, True # cut_off = True
factory = VoiceComponentFactory()
transcriber = factory.create_transcriber(config) # Deepgram, AssemblyAI, etc.
agent = factory.create_agent(config) # OpenAI, Gemini, etc.
synthesizer = factory.create_synthesizer(config) # ElevenLabs, Azure, etc.
The skill includes examples for:
See references/common_pitfalls.md for detailed solutions to:
This skill is part of the Antigravity Awesome Skills repository. Contributions are welcome!
@websocket-patterns - WebSocket implementation@async-python - Asyncio patterns@streaming-apis - Streaming API integration@audio-processing - Audio format conversionMIT License - See repository LICENSE file
Built with ❤️ for the Antigravity community
View Count
0
Download Count
0
Favorite Count
0
Quality Score
75