前端通过 SSE / ReadableStream 接收大模型流式输出时,UI 渲染慢、内存容易暴涨,应该如何优化?

流式场景(如 AI 对话)的核心是「边读边处理,不 accumulate 全量」。

1. 用 ReadableStream 逐块消费,而不是等全部读完

const res = await fetch('/api/stream')
const reader = res.body.getReader()
const decoder = new TextDecoder()

while (true) {
  const { done, value } = await reader.read()
  if (done) break
  const chunk = decoder.decode(value, { stream: true })
  appendToUI(chunk)        // 逐块追加,不保存全量字符串
}

2. 避免内存暴涨的关键点

3. 多 Session / 切换标签页的状态隔离

4. 内存泄漏排查

一句话:流式渲染 = 增量消费 + 合并更新 + 及时释放,不要等「全部到齐」再处理。

同分类其他题目