第 6 章:Node.js 对接大模型
这是 AI 时代 Node 最高频的实战场景——用 Node 服务调用大模型 API,实现流式输出。这一章把前面学的 HTTP、stream、异步全部串起来。
6.1 整体架构
Node 在大模型应用里的典型定位是中间编排层:前端不直接连大模型,而是连 Node 服务,由 Node 统一调用、流式转发、记日志、做鉴权和计费:
┌──────────┐ SSE 流式 ┌───────────────┐ HTTP 流式 ┌─────────────┐
│ 前端页面 │ ◄──────────► │ Node 服务层 │ ◄──────────► │ 大模型 API │
│ (浏览器) │ │ (Express/Koa) │ │ (OpenAI等) │
└──────────┘ └───────┬───────┘ └─────────────┘
│
┌───────▼───────┐
│ 日志/鉴权/计费 │
└───────────────┘为什么中间加一层 Node:
- 大模型的 API Key 不能暴露给前端;
- 前端直接调大模型有 CORS 限制,且无法统一计费、限流、日志;
- Node 的流式能力天然适合转发大模型的流式输出。
6.2 普通调用:一次性拿完整结果
最简单的方式是 fetch 调大模型 API(OpenAI 兼容协议):
javascript
async function chat(prompt) {
const res = await fetch('https://api.openai.com/v1/chat/completions', {
method: 'POST',
headers: {
'Content-Type': 'application/json',
'Authorization': `Bearer ${process.env.OPENAI_API_KEY}`,
},
body: JSON.stringify({
model: 'gpt-4o-mini',
messages: [{ role: 'user', content: prompt }],
}),
});
const data = await res.json();
return data.choices[0].message.content; // 完整回复文本
}问题:大模型生成是「一个字一个字」出来的,一次性 res.json() 要等它全部生成完才返回,用户要等很久才看到第一个字。体验差。
6.3 流式调用:SSE 边生成边返回
大模型 API 支持 stream: true,返回的是 SSE(Server-Sent Events) 格式——服务端不断推送数据块,客户端边收边显示:
时间轴 ─────────────────────────────────────────►
大模型生成: "你" → "好" → "," → "我" → "是" → ...
│ │ │
▼ ▼ ▼
SSE 推送: data: 你 data: 好 data: ,... (一个个 token 流过来)
│
▼
前端/Node 边收边显示,不用等全部生成完用 Node 流式接收:
javascript
async function chatStream(prompt, onChunk) {
const res = await fetch('https://api.openai.com/v1/chat/completions', {
method: 'POST',
headers: {
'Content-Type': 'application/json',
'Authorization': `Bearer ${process.env.OPENAI_API_KEY}`,
},
body: JSON.stringify({
model: 'gpt-4o-mini',
stream: true, // 关键:开启流式
messages: [{ role: 'user', content: prompt }],
}),
});
// 用 stream 边读边解析 SSE
const reader = res.body.getReader();
const decoder = new TextDecoder();
let buffer = '';
while (true) {
const { done, value } = await reader.read();
if (done) break;
buffer += decoder.decode(value, { stream: true });
const lines = buffer.split('\n');
buffer = lines.pop(); // 最后一个可能不完整,留到下次
for (const line of lines) {
if (line.startsWith('data: ')) {
const payload = line.slice(6).trim();
if (payload === '[DONE]') break;
const json = JSON.parse(payload);
const token = json.choices[0]?.delta?.content || '';
if (token) onChunk(token); // 每收到一个字就回调
}
}
}
}SSE 数据格式(每行一个 data:):
data: {"choices":[{"delta":{"content":"你"}}]}
data: {"choices":[{"delta":{"content":"好"}}]}
data: [DONE]6.4 完整落地:Express 流式转发给前端
把大模型的流式输出,通过 Express 用 SSE 转发给前端(前端用 EventSource 或 fetch 流式接收):
javascript
import express from 'express';
const app = express();
app.use(express.json());
app.post('/api/chat', async (req, res) => {
const { prompt } = req.body;
// 设置 SSE 响应头
res.setHeader('Content-Type', 'text/event-stream');
res.setHeader('Cache-Control', 'no-cache');
res.setHeader('Connection', 'keep-alive');
try {
// 复用上一节的 chatStream,把每个 token 转成 SSE 推给前端
await chatStream(prompt, (token) => {
res.write(`data: ${JSON.stringify({ content: token })}\n\n`);
});
res.write('data: [DONE]\n\n');
res.end();
} catch (err) {
res.write(`data: ${JSON.stringify({ error: err.message })}\n\n`);
res.end();
}
});
app.listen(3000);前端 ──POST /api/chat──► Node Express ──stream:true──► 大模型
▲ │ │
│ SSE 流式返回 │◄──── token 流 ───────────┘
└──────────────────────────┘6.5 工程化要点
生产级的大模型服务,除了「调通」,还要处理这些:
| 要点 | 做法 |
|---|---|
| 超时控制 | AbortController 给 fetch 设超时,防止大模型卡死拖垮服务 |
| 错误兜底 | 捕获网络/解析错误,返回友好提示,别让进程崩溃 |
| 重试 | 网络抖动时对非流式请求做有限重试(流式一般不重试) |
| 并发限流 | 大模型有 QPS 限制,用队列/信号量控制并发 |
| 多模型切换 | 统一封装接口,OpenAI/通义/DeepSeek 等自由切换 |
| Token 计费 | 从返回的 usage 里统计 token,做成本核算 |
一个超时 + 多模型的封装骨架:
javascript
async function callLLM({ model = 'deepseek-chat', prompt, signal }) {
const endpoints = {
'deepseek-chat': 'https://api.deepseek.com/v1/chat/completions',
'gpt-4o-mini': 'https://api.openai.com/v1/chat/completions',
};
const keys = {
'deepseek-chat': process.env.DEEPSEEK_API_KEY,
'gpt-4o-mini': process.env.OPENAI_API_KEY,
};
const res = await fetch(endpoints[model], {
method: 'POST',
headers: {
'Content-Type': 'application/json',
'Authorization': `Bearer ${keys[model]}`,
},
body: JSON.stringify({ model, messages: [{ role: 'user', content: prompt }] }),
signal, // 传入超时信号
});
return res.json();
}6.6 本章小结
| 要点 | 说明 |
|---|---|
| 架构定位 | Node 是前端和大模型之间的编排层 |
| 普通调用 | fetch + json,等完整结果 |
| 流式调用 | stream: true + SSE,边生成边返回,体验好 |
| 关键能力 | 读响应流(getReader)+ 解析 SSE(data: 行)+ 转发 |
| 工程化 | 超时、重试、限流、多模型、计费一个都不能少 |
下一章预告
功能写完了,怎么调试、怎么压性能、怎么用 pm2 和 Docker 稳定部署?最后一章讲工程化落地。