第七章 语音能力
7.1 概述
Mastra Agent 可以被赋予语音能力——既能"说"(TTS,Text-to-Speech),也能"听"(STT,Speech-to-Text),甚至支持实时语音对话(Speech-to-Speech)。
语音能力的典型应用:
- 智能客服语音交互
- 语音笔记助手
- 实时语音翻译
- 播客/有声内容生成
7.2 支持的语音提供商
| 提供商 | 包名 | 能力 |
|---|---|---|
| OpenAI | @mastra/voice-openai | TTS + STT |
| OpenAI Realtime | @mastra/voice-openai-realtime | 实时语音对话 |
| ElevenLabs | @mastra/voice-elevenlabs | 高品质 TTS |
| PlayAI | @mastra/voice-playai | TTS |
@mastra/voice-google | TTS + STT | |
| Deepgram | @mastra/voice-deepgram | STT |
| Azure | @mastra/voice-azure | TTS + STT |
| Cloudflare | @mastra/voice-cloudflare | TTS |
7.3 基本用法:单一提供商
typescript
import { Agent } from '@mastra/core/agent'
import { OpenAIVoice } from '@mastra/voice-openai'
const voice = new OpenAIVoice()
export const voiceAgent = new Agent({
id: 'voice-agent',
name: 'Voice Agent',
instructions: '你是一个具有语音能力的助手。',
model: 'openai/gpt-4.1',
voice, // 绑定语音提供商
})文字转语音(TTS)
typescript
import { createWriteStream } from 'fs'
import path from 'path'
// 生成语音流
const audio = await voiceAgent.voice.speak('你好,我是你的 AI 助手!')
// 保存为音频文件
const filePath = path.join(process.cwd(), 'output.mp3')
const writer = createWriteStream(filePath)
audio.pipe(writer)
await new Promise((resolve, reject) => {
writer.on('finish', resolve)
writer.on('error', reject)
})语音转文字(STT)
typescript
import { createReadStream } from 'fs'
import path from 'path'
const audioFilePath = path.join(process.cwd(), 'input.m4a')
const audioStream = createReadStream(audioFilePath)
const transcription = await voiceAgent.voice.listen(audioStream, {
filetype: 'm4a',
})
console.log('转写结果:', transcription)7.4 混合提供商:CompositeVoice
实际项目中,你可能想用不同提供商处理不同方向——比如用 OpenAI 做语音识别(性价比高),用 ElevenLabs 做语音合成(效果更自然):
typescript
import { CompositeVoice } from '@mastra/core/voice'
import { OpenAIVoice } from '@mastra/voice-openai'
import { PlayAIVoice } from '@mastra/voice-playai'
const agent = new Agent({
id: 'hybrid-voice-agent',
name: 'Hybrid Voice Agent',
instructions: '你是一个双向语音助手。',
model: 'openai/gpt-4.1',
voice: new CompositeVoice({
input: new OpenAIVoice(), // STT 用 OpenAI
output: new PlayAIVoice(), // TTS 用 PlayAI
}),
})也可以与 AI SDK 提供商混用
typescript
import { CompositeVoice } from '@mastra/core/voice'
import { openai } from '@ai-sdk/openai'
import { elevenlabs } from '@ai-sdk/elevenlabs'
const voice = new CompositeVoice({
input: openai.transcription('whisper-1'), // AI SDK STT
output: elevenlabs.speech('eleven_turbo_v2'), // AI SDK TTS
})7.5 实时语音对话(Speech-to-Speech)
最酷的模式——通过 WebSocket 实现实时双向语音交互:
typescript
import { Agent } from '@mastra/core/agent'
import { getMicrophoneStream } from '@mastra/node-audio'
import { OpenAIRealtimeVoice } from '@mastra/voice-openai-realtime'
const voice = new OpenAIRealtimeVoice({
apiKey: process.env.OPENAI_API_KEY,
model: 'gpt-5.1-realtime',
speaker: 'alloy', // 音色选择
})
const agent = new Agent({
id: 'realtime-agent',
name: 'Realtime Agent',
instructions: '你是一个实时语音助手。',
model: 'openai/gpt-4.1',
tools: { searchTool, calculateTool }, // 工具也会传给语音提供商
voice,
})
// 建立 WebSocket 连接
await agent.voice.connect()
// 开始对话
agent.voice.speak("你好,有什么可以帮你的?")
// 接收麦克风输入
const microphoneStream = getMicrophoneStream()
agent.voice.send(microphoneStream)
// 监听事件
agent.voice.on('speaking', ({ audio }) => {
// audio 是 ReadableStream 或 Int16Array,播放它
})
agent.voice.on('writing', ({ text, role }) => {
console.log(`${role}: ${text}`)
})
agent.voice.on('error', (error) => {
console.error('语音错误:', error)
})
// 结束对话
agent.voice.close()7.6 完整示例:双 Agent 语音交互
一个有趣的例子——两个 Agent 之间的语音对话:
typescript
import { Agent } from '@mastra/core/agent'
import { CompositeVoice } from '@mastra/core/voice'
import { OpenAIVoice } from '@mastra/voice-openai'
import { Mastra } from '@mastra/core'
// Agent A:提问者
const questionAgent = new Agent({
id: 'questioner',
name: 'Questioner',
model: 'openai/gpt-4.1',
instructions: '你专门提出有趣的哲学问题。',
voice: new CompositeVoice({
input: new OpenAIVoice(),
output: new OpenAIVoice(),
}),
})
// Agent B:回答者
const answerAgent = new Agent({
id: 'answerer',
name: 'Answerer',
instructions: '你用简洁深刻的方式回答哲学问题。',
model: 'openai/gpt-4.1',
voice: new OpenAIVoice(),
})
const mastra = new Mastra({
agents: { questionAgent, answerAgent },
})
// 对话流程
const questioner = mastra.getAgent('questionAgent')
const answerer = mastra.getAgent('answerer')
// 1. Agent A 语音提问
const questionAudio = await questioner.voice.speak('生命的意义是什么?')
await saveToFile(questionAudio, 'question.mp3')
// 2. Agent B 听取并转写
const questionText = await answerer.voice.listen(
createReadStream('question.mp3')
)
// 3. Agent B 生成回答
const answer = await answerer.generate(questionText)
// 4. Agent B 说出回答
const answerAudio = await answerer.voice.speak(answer.text)
await saveToFile(answerAudio, 'answer.mp3')7.7 实际应用建议
- 延迟优化:实时对话对延迟敏感,优先选择低延迟的提供商
- 音色选择:不同场景用不同音色,客服用温和的,播报用清晰的
- 错误处理:语音流容易因网络问题中断,务必加好 error 处理
- 成本控制:TTS 按字符计费,STT 按时长计费,生产环境注意监控用量
- 中文支持:并非所有提供商对中文的支持都一样好,实际测试很重要
7.8 本章小结
| 模式 | 方法 | 适用场景 |
|---|---|---|
| TTS | agent.voice.speak() | 语音播报、有声内容 |
| STT | agent.voice.listen() | 语音输入、会议转写 |
| 实时对话 | OpenAIRealtimeVoice | 语音助手、电话客服 |
| 混合模式 | CompositeVoice | 分别优化输入和输出 |
核心理念:语音能力不是独立功能,而是 Agent 的一个属性。同一个 Agent 的 Tools、Instructions、Memory 在语音模式下同样有效。
下一章我们将学习评估与可观测性,这是让 AI 应用从原型走向生产的关键。