Skip to content

第七章 语音能力 ​

7.1 概述 ​

Mastra Agent 可以被赋予语音能力——既能"说"(TTS,Text-to-Speech),也能"听"(STT,Speech-to-Text),甚至支持实时语音对话(Speech-to-Speech)。

语音能力的典型应用:

  • 智能客服语音交互
  • 语音笔记助手
  • 实时语音翻译
  • 播客/有声内容生成

7.2 支持的语音提供商 ​

提供商包名能力
OpenAI@mastra/voice-openaiTTS + STT
OpenAI Realtime@mastra/voice-openai-realtime实时语音对话
ElevenLabs@mastra/voice-elevenlabs高品质 TTS
PlayAI@mastra/voice-playaiTTS
Google@mastra/voice-googleTTS + STT
Deepgram@mastra/voice-deepgramSTT
Azure@mastra/voice-azureTTS + STT
Cloudflare@mastra/voice-cloudflareTTS

7.3 基本用法:单一提供商 ​

typescript
import { Agent } from '@mastra/core/agent'
import { OpenAIVoice } from '@mastra/voice-openai'

const voice = new OpenAIVoice()

export const voiceAgent = new Agent({
  id: 'voice-agent',
  name: 'Voice Agent',
  instructions: '你是一个具有语音能力的助手。',
  model: 'openai/gpt-4.1',
  voice, // 绑定语音提供商
})

文字转语音(TTS) ​

typescript
import { createWriteStream } from 'fs'
import path from 'path'

// 生成语音流
const audio = await voiceAgent.voice.speak('你好,我是你的 AI 助手!')

// 保存为音频文件
const filePath = path.join(process.cwd(), 'output.mp3')
const writer = createWriteStream(filePath)
audio.pipe(writer)
await new Promise((resolve, reject) => {
  writer.on('finish', resolve)
  writer.on('error', reject)
})

语音转文字(STT) ​

typescript
import { createReadStream } from 'fs'
import path from 'path'

const audioFilePath = path.join(process.cwd(), 'input.m4a')
const audioStream = createReadStream(audioFilePath)

const transcription = await voiceAgent.voice.listen(audioStream, {
  filetype: 'm4a',
})
console.log('转写结果:', transcription)

7.4 混合提供商:CompositeVoice ​

实际项目中,你可能想用不同提供商处理不同方向——比如用 OpenAI 做语音识别(性价比高),用 ElevenLabs 做语音合成(效果更自然):

typescript
import { CompositeVoice } from '@mastra/core/voice'
import { OpenAIVoice } from '@mastra/voice-openai'
import { PlayAIVoice } from '@mastra/voice-playai'

const agent = new Agent({
  id: 'hybrid-voice-agent',
  name: 'Hybrid Voice Agent',
  instructions: '你是一个双向语音助手。',
  model: 'openai/gpt-4.1',
  voice: new CompositeVoice({
    input: new OpenAIVoice(),    // STT 用 OpenAI
    output: new PlayAIVoice(),   // TTS 用 PlayAI
  }),
})

也可以与 AI SDK 提供商混用 ​

typescript
import { CompositeVoice } from '@mastra/core/voice'
import { openai } from '@ai-sdk/openai'
import { elevenlabs } from '@ai-sdk/elevenlabs'

const voice = new CompositeVoice({
  input: openai.transcription('whisper-1'),       // AI SDK STT
  output: elevenlabs.speech('eleven_turbo_v2'),   // AI SDK TTS
})

7.5 实时语音对话(Speech-to-Speech) ​

最酷的模式——通过 WebSocket 实现实时双向语音交互:

typescript
import { Agent } from '@mastra/core/agent'
import { getMicrophoneStream } from '@mastra/node-audio'
import { OpenAIRealtimeVoice } from '@mastra/voice-openai-realtime'

const voice = new OpenAIRealtimeVoice({
  apiKey: process.env.OPENAI_API_KEY,
  model: 'gpt-5.1-realtime',
  speaker: 'alloy', // 音色选择
})

const agent = new Agent({
  id: 'realtime-agent',
  name: 'Realtime Agent',
  instructions: '你是一个实时语音助手。',
  model: 'openai/gpt-4.1',
  tools: { searchTool, calculateTool }, // 工具也会传给语音提供商
  voice,
})

// 建立 WebSocket 连接
await agent.voice.connect()

// 开始对话
agent.voice.speak("你好,有什么可以帮你的?")

// 接收麦克风输入
const microphoneStream = getMicrophoneStream()
agent.voice.send(microphoneStream)

// 监听事件
agent.voice.on('speaking', ({ audio }) => {
  // audio 是 ReadableStream 或 Int16Array,播放它
})

agent.voice.on('writing', ({ text, role }) => {
  console.log(`${role}: ${text}`)
})

agent.voice.on('error', (error) => {
  console.error('语音错误:', error)
})

// 结束对话
agent.voice.close()

7.6 完整示例:双 Agent 语音交互 ​

一个有趣的例子——两个 Agent 之间的语音对话:

typescript
import { Agent } from '@mastra/core/agent'
import { CompositeVoice } from '@mastra/core/voice'
import { OpenAIVoice } from '@mastra/voice-openai'
import { Mastra } from '@mastra/core'

// Agent A:提问者
const questionAgent = new Agent({
  id: 'questioner',
  name: 'Questioner',
  model: 'openai/gpt-4.1',
  instructions: '你专门提出有趣的哲学问题。',
  voice: new CompositeVoice({
    input: new OpenAIVoice(),
    output: new OpenAIVoice(),
  }),
})

// Agent B:回答者
const answerAgent = new Agent({
  id: 'answerer',
  name: 'Answerer',
  instructions: '你用简洁深刻的方式回答哲学问题。',
  model: 'openai/gpt-4.1',
  voice: new OpenAIVoice(),
})

const mastra = new Mastra({
  agents: { questionAgent, answerAgent },
})

// 对话流程
const questioner = mastra.getAgent('questionAgent')
const answerer = mastra.getAgent('answerer')

// 1. Agent A 语音提问
const questionAudio = await questioner.voice.speak('生命的意义是什么?')
await saveToFile(questionAudio, 'question.mp3')

// 2. Agent B 听取并转写
const questionText = await answerer.voice.listen(
  createReadStream('question.mp3')
)

// 3. Agent B 生成回答
const answer = await answerer.generate(questionText)

// 4. Agent B 说出回答
const answerAudio = await answerer.voice.speak(answer.text)
await saveToFile(answerAudio, 'answer.mp3')

7.7 实际应用建议 ​

  1. 延迟优化:实时对话对延迟敏感,优先选择低延迟的提供商
  2. 音色选择:不同场景用不同音色,客服用温和的,播报用清晰的
  3. 错误处理:语音流容易因网络问题中断,务必加好 error 处理
  4. 成本控制:TTS 按字符计费,STT 按时长计费,生产环境注意监控用量
  5. 中文支持:并非所有提供商对中文的支持都一样好,实际测试很重要

7.8 本章小结 ​

模式方法适用场景
TTSagent.voice.speak()语音播报、有声内容
STTagent.voice.listen()语音输入、会议转写
实时对话OpenAIRealtimeVoice语音助手、电话客服
混合模式CompositeVoice分别优化输入和输出

核心理念:语音能力不是独立功能,而是 Agent 的一个属性。同一个 Agent 的 Tools、Instructions、Memory 在语音模式下同样有效。

下一章我们将学习评估与可观测性,这是让 AI 应用从原型走向生产的关键。

📖本文阅读--次|📊全站访问--次|👥访客--人