一、你的流式输出已经够快了,但用户还是在等
你已经上了 SSE 流式输出,TTFT 稳定在 800ms,P99 也控制在 1.2 秒以内。你觉得延迟优化到头了。
然后你做了一次用户研究,发现有个反复出现的抱怨:
"点击生成按钮之后感觉要等很久才有反应。"
用户描述的"等很久",就是你精心优化过的 800ms。
这个问题不是延迟本身,而是感知延迟。用户点击按钮的那一刻,神经系统进入等待模式,超过 300ms 没有视觉反馈就开始焦虑。你的 800ms 在服务器监控上是绿色,在用户脑子里是红色。
解决感知延迟有两个思路:
- 让响应开始得更快(优化 TTFT)
- 让响应在用户感知到之前就已经开始(Prefetching)
第一个思路你已经做得很好了。第二个思路,大多数团队没碰过。
这篇文章讲的就是第二个:LLM Prefetching,在用户明确发出请求之前,预测性地提前发起 LLM 调用,把感知延迟从 800ms 压到不足 300ms。
二、Prefetching 不是新概念,但 LLM 场景有特殊挑战
浏览器预取是老技术了。<link rel="prefetch"> 告诉浏览器提前加载下一页的资源。DNS prefetch 提前解析域名。Chrome 的 Speculation Rules API 甚至可以预渲染整个页面。
CPU 里的分支预测器每个时钟周期都在猜你接下来要执行哪条指令。CDN 的 Edge Worker 在用户请求到来之前就已经把缓存预热好了。
这些技术的核心逻辑一致:利用信号预测未来需求,提前准备,消灭等待。
但把 Prefetching 搬进 LLM 场景,有三个静态资源没有的特殊挑战:
挑战 1:预取有真实成本
浏览器预取失败,浪费的是带宽。LLM 预取失败,浪费的是真实 API Token 费用。一次大模型调用输出 500 tokens,约几分钱。看起来小,但如果预测准确率只有 40%,相当于每次有效请求你额外付了 1.5 倍的钱。
挑战 2:响应不可缓存(通常)
静态资源有确定的内容,可以长时间缓存。大模型响应是 prompt + model + temperature 的函数,prompt 里任何一个字变了,缓存就失效了。预取的响应有效期通常只有几十秒。
挑战 3:流式响应的中断与复用
LLM 响应通常是流式的。预取发起后,服务器开始向你传输 token stream。但如果用户没有真正触发请求,这个 stream 需要被取消;如果用户触发了,你需要能直接"接管"这个已经在流动的 stream,而不是重新请求。这比取消一个 HTTP GET 请求复杂得多。
三、四种 Prefetching 策略,按成本从低到高
策略 1:Context Warm-up(上下文预热)
原理:不预取最终响应,只预热大模型需要的 context。
如果你用了国内主流大模型厂商(如通义千问、DeepSeek)的 Prompt Caching 能力,system prompt 和常用 context 可以缓存在 KV Cache 中。预热的意思是:在用户进入某个功能页面时,发一次带 cache_control 标记的"空请求"(或者最轻量的探测请求),让 Provider 的 KV Cache 热起来。
当用户真正触发请求时,system prompt 命中 KV Cache,TTFT 可以降低 50-85%。
// 用户进入文档编辑页面时触发
async function warmupDocumentContext(docId: string) {
const doc = await fetchDocument(docId);
// 发一次最小 prompt,目的是让 system prompt + doc context 进入 KV Cache
// 以 DeepSeek / 通义千问兼容接口为例
await llmClient.messages.create({
model: "deepseek-chat",
max_tokens: 1, // 只要 1 个 token,成本极低
system: [
{
type: "text",
text: buildSystemPrompt(),
cache_control: { type: "ephemeral" }
},
{
type: "text",
text: `Document context:\n${doc.content}`,
cache_control: { type: "ephemeral" }
}
],
messages: [{ role: "user", content: "ready" }]
});
}
成本:极低。一次 warm-up 消耗约 1 output token + cache write tokens(通常比 input token 便宜 75%)。
收益:后续真实请求命中 KV Cache,TTFT 降低 50-85%。
适用场景:用户进入编辑器、打开文档、进入某个功能模块时触发。
策略 2:Hover-triggered Prefetch(悬停预取)
原理:用户将鼠标悬停在触发元素上 150-200ms 时,预先发起 LLM 请求。
核心洞察:用户 hover 一个按钮超过 150ms,通常意味着他打算点击。这个时间窗口足够发起 LLM 请求并拿到前几十个 token。
// React hook:悬停预取
function useHoverPrefetch(
prompt: string,
options: { hoverDelayMs?: number; ttlMs?: number } = {}
) {
const { hoverDelayMs = 150, ttlMs = 30000 } = options;
const prefetchRef = useRef<PrefetchSession | null>(null);
const timerRef = useRef<NodeJS.Timeout>();
const onMouseEnter = useCallback(() => {
timerRef.current = setTimeout(() => {
// 150ms 后用户还没移走,开始预取
prefetchRef.current = startPrefetch(prompt, { ttlMs });
}, hoverDelayMs);
}, [prompt, hoverDelayMs, ttlMs]);
const onMouseLeave = useCallback(() => {
clearTimeout(timerRef.current);
// 用户移走了:取消正在进行的预取(如果还没完成)
prefetchRef.current?.cancel();
prefetchRef.current = null;
}, []);
const consumePrefetch = useCallback(() => {
const session = prefetchRef.current;
prefetchRef.current = null;
return session; // 调用者接管这个 stream
}, []);
return { onMouseEnter, onMouseLeave, consumePrefetch };
}
PrefetchSession 的实现:
interface PrefetchSession {
stream: ReadableStream<string>;
buffered: string[]; // 已收到但未消费的 tokens
status: 'pending' | 'streaming' | 'complete' | 'cancelled';
cancel: () => void;
}
function startPrefetch(prompt: string, options: { ttlMs: number }): PrefetchSession {
const controller = new AbortController();
const buffered: string[] = [];
let status: PrefetchSession['status'] = 'pending';
// 启动低优先级预取请求
const responsePromise = fetch('/api/llm/prefetch', {
method: 'POST',
headers: { 'Content-Type': 'application/json' },
body: JSON.stringify({ prompt, priority: 'low' }),
signal: controller.signal,
});
// 开始缓冲 tokens
responsePromise.then(async (res) => {
status = 'streaming';
const reader = res.body!.getReader();
const decoder = new TextDecoder();
while (true) {
const { done, value } = await reader.read();
if (done) { status = 'complete'; break; }
buffered.push(decoder.decode(value));
}
}).catch((err) => {
if (err.name !== 'AbortError') console.error('Prefetch failed:', err);
});
// TTL 到期自动取消
const ttlTimer = setTimeout(() => {
if (status !== 'complete') {
controller.abort();
status = 'cancelled';
}
}, options.ttlMs);
return {
stream: createReadableFromBuffer(buffered, responsePromise),
buffered,
get status() { return status; },
cancel: () => {
clearTimeout(ttlTimer);
controller.abort();
status = 'cancelled';
}
};
}
命中后如何复用:
function GenerateButton({ prompt }: { prompt: string }) {
const { onMouseEnter, onMouseLeave, consumePrefetch } = useHoverPrefetch(prompt);
const handleClick = async () => {
const prefetched = consumePrefetch();
if (prefetched && prefetched.status !== 'cancelled') {
// 命中!直接复用已有 stream
// 用户感知到的延迟 = 0(已有 buffered tokens 直接渲染)
streamFromPrefetch(prefetched);
} else {
// 未命中,走正常请求
await streamFromFresh(prompt);
}
};
return (
<button
onMouseEnter={onMouseEnter}
onMouseLeave={onMouseLeave}
onClick={handleClick}
>
生成
</button>
);
}
实测数据:在 Chat UI 场景,hover prefetch 把感知延迟从约 1.1s 压到 450ms 左右。命中率约 55-70%(取决于用户是否犹豫)。
策略 3:Debounce Prefetch(防抖预取)
原理:用户在输入框中停止输入 300-500ms 后,预测用户可能的完整输入并预取响应。
核心假设:用户停止打字超过 300ms,通常意味着他在思考如何结束这句话,而不是还在输入。
class DebouncePrefetcher {
private currentPrefetch: PrefetchSession | null = null;
private debounceTimer: NodeJS.Timeout | null = null;
private lastInput = '';
constructor(
private readonly debounceMs: number = 350,
private readonly intentClassifier: IntentClassifier,
) {}
onInputChange(input: string) {
this.lastInput = input;
// 清除旧的 debounce 和预取
if (this.debounceTimer) clearTimeout(this.debounceTimer);
this.currentPrefetch?.cancel();
this.currentPrefetch = null;
// 输入太短,不值得预取
if (input.length < 10) return;
this.debounceTimer = setTimeout(async () => {
await this.maybePrefetch(input);
}, this.debounceMs);
}
private async maybePrefetch(input: string) {
// 第一步:用轻量模型做意图分类(避免直接用昂贵模型预取)
const intent = await this.intentClassifier.classify(input);
// 只有置信度足够高才值得预取
if (intent.confidence < 0.72) {
console.debug(`[prefetch] skip: low confidence ${intent.confidence}`);
return;
}
console.debug(`[prefetch] start: "${input.slice(0, 30)}..." (confidence: ${intent.confidence})`);
this.currentPrefetch = startPrefetch(input, { ttlMs: 45000 });
}
consumePrefetch(actualInput: string): PrefetchSession | null {
if (!this.currentPrefetch) return null;
// 验证实际输入与预取时的输入足够相似
const similarity = cosineSimilarity(actualInput, this.lastInput);
if (similarity < 0.85) {
this.currentPrefetch.cancel();
this.currentPrefetch = null;
return null;
}
const session = this.currentPrefetch;
this.currentPrefetch = null;
return session;
}
}
意图分类器(轻量版),使用国产轻量模型成本更低:
class IntentClassifier {
// 用 Qwen-Turbo 或 DeepSeek-V3 做意图分类,成本极低
async classify(input: string): Promise<{ intent: string; confidence: number }> {
const response = await llmClient.chat.completions.create({
model: "qwen-turbo",
max_tokens: 20,
messages: [{
role: "user",
content: `Rate the completeness of this user query on a scale 0-1.
Only output a JSON: {"confidence": 0.X}
Query: "${input}"`
}]
});
try {
return JSON.parse(response.choices[0].message.content!);
} catch {
return { intent: 'unknown', confidence: 0 };
}
}
}
策略 4:Workflow Prefetch(工作流预取)
原理:已知用户会经历的固定流程(如 onboarding wizard、多步骤表单),在用户完成当前步骤时,预取下一步的 LLM 响应。
// Onboarding 向导:步骤 2 完成时,提前为步骤 3 准备内容
class WorkflowPrefetcher {
private prefetchCache = new Map<string, PrefetchSession>();
onStepComplete(step: number, formData: Record<string, unknown>) {
const nextStep = step + 1;
const nextPrompt = this.buildStepPrompt(nextStep, formData);
if (!nextPrompt) return; // 没有下一步了
const cacheKey = `step-${nextStep}-${hashObject(formData)}`;
// 取消旧的预取,开始新的
this.prefetchCache.get(cacheKey)?.cancel();
this.prefetchCache.set(cacheKey, startPrefetch(nextPrompt, { ttlMs: 120000 }));
console.debug(`[workflow-prefetch] pre-loading step ${nextStep}`);
}
consumeForStep(step: number, formData: Record<string, unknown>): PrefetchSession | null {
const cacheKey = `step-${step}-${hashObject(formData)}`;
const session = this.prefetchCache.get(cacheKey) ?? null;
this.prefetchCache.delete(cacheKey);
return session;
}
}
工作流预取的命中率通常最高(80-95%),因为下一步的 prompt 几乎是确定的。代价是提前消耗 API 配额,适合高价值用户流程。
四、服务端:预取请求的优先级隔离
客户端的预取策略再好,如果服务端不区分预取请求和正式请求,预取会挤占正式请求的资源。
4.1 请求标记
在 HTTP 头里标记预取请求:
// 客户端发送预取请求时
const response = await fetch('/api/llm/stream', {
method: 'POST',
headers: {
'Content-Type': 'application/json',
'X-Request-Priority': 'low', // 低优先级
'X-Prefetch': 'true', // 明确标记为预取
'X-Prefetch-TTL': '30000', // 预取响应的 TTL
},
body: JSON.stringify({ prompt }),
signal: controller.signal,
});
4.2 服务端队列隔离
// Express/Fastify 中间件:分离预取流量
const prefetchQueue = new PQueue({ concurrency: 2 }); // 最多 2 个并发预取
const normalQueue = new PQueue({ concurrency: 10 }); // 正式请求 10 并发
app.post('/api/llm/stream', async (req, res) => {
const isPrefetch = req.headers['x-prefetch'] === 'true';
const queue = isPrefetch ? prefetchQueue : normalQueue;
await queue.add(async () => {
// 如果是预取请求且队列已满,直接 429
if (isPrefetch && queue.size > 5) {
res.status(429).json({ error: 'prefetch_capacity_exceeded' });
return;
}
await handleLLMStream(req, res);
});
});
4.3 Token Budget 隔离
预取请求不计入用户的 Token 配额,但要计入服务端的账单预算:
class PrefetchBudgetGuard {
private hourlyPrefetchTokens = 0;
private readonly hourlyBudget: number;
constructor(hourlyBudget: number) {
this.hourlyBudget = hourlyBudget;
// 每小时重置
setInterval(() => { this.hourlyPrefetchTokens = 0; }, 3600 * 1000);
}
async checkAndReserve(estimatedTokens: number): Promise<boolean> {
if (this.hourlyPrefetchTokens + estimatedTokens > this.hourlyBudget) {
console.warn(`[prefetch-budget] budget exhausted: ${this.hourlyPrefetchTokens}/${this.hourlyBudget}`);
return false; // 拒绝预取
}
this.hourlyPrefetchTokens += estimatedTokens;
return true;
}
onPrefetchComplete(actualTokens: number, estimated: number) {
// 用实际 token 数修正
this.hourlyPrefetchTokens += (actualTokens - estimated);
}
}
五、5 个生产陷阱
陷阱 1:预取命中率不监控,成本悄悄翻倍
不监控命中率是最常见的错误。预取系统上线两周后,某团队发现 API 账单增加了 80%,查了半天才发现预取命中率只有 18%——相当于每次有效请求额外付了 4.5 倍的钱。
解法:在每次 consumePrefetch 调用时记录 hit/miss,用监控系统追踪命中率,设置告警阈值(建议 < 35% 时告警)。
// 在 consumePrefetch 中埋点
const session = prefetcher.consumePrefetch(input);
metrics.increment('prefetch.attempt');
if (session && session.status !== 'cancelled') {
metrics.increment('prefetch.hit');
} else {
metrics.increment('prefetch.miss');
if (session?.status === 'cancelled') metrics.increment('prefetch.cancelled');
}
陷阱 2:预取请求消耗 Provider 并发槽位
国内主流大模型 Provider 都有 per-key 的并发 in-flight 限制。如果你的预取系统在高峰期发出 50 个并发预取请求,正式请求可能会被 rate limit 掉。
解法:预取请求必须使用独立的 API Key(或独立的账号),与正式请求的 Key 完全隔离。
// 使用独立的预取专用 client
const prefetchClient = createLLMClient({
apiKey: process.env.LLM_PREFETCH_API_KEY, // 专用 key
});
const normalClient = createLLMClient({
apiKey: process.env.LLM_API_KEY, // 正式 key
});
陷阱 3:流式复用的缓冲区溢出
预取后用户迟迟不点击,缓冲区一直在累积 tokens。一个 4096 token 的响应,如果全部缓冲在内存里,每个预取 session 占约 16KB(token 平均 4 bytes × 4096)。100 个并发预取 session = 1.6MB,听起来不多,但如果 token 是以 JSON SSE 格式缓冲的,实际内存占用可能是 10-15 倍。
解法:缓冲区设上限,超过阈值后停止缓冲(但不取消请求,依然接收并丢弃 tokens):
const MAX_BUFFER_TOKENS = 200; // 只缓冲前 200 个 tokens
// 在 startPrefetch 中
if (buffered.length >= MAX_BUFFER_TOKENS) {
// 超过上限:丢弃 token 但保持连接
// 用户命中时从这里开始直接传流
continue;
}
buffered.push(token);
陷阱 4:预取响应在上下文变化后仍被消费
用户在文档 A 上悬停了按钮,触发了预取。然后用户切换到文档 B,点击了按钮。如果消费层不校验上下文,会把文档 A 的响应显示给文档 B 的用户。
解法:预取 session 必须绑定上下文指纹:
interface PrefetchSession {
contextFingerprint: string; // 上下文的哈希
// ...
}
function consumePrefetch(currentContext: unknown): PrefetchSession | null {
if (!this.session) return null;
const currentFingerprint = hashObject(currentContext);
if (currentFingerprint !== this.session.contextFingerprint) {
// 上下文变了,丢弃预取
this.session.cancel();
this.session = null;
return null;
}
return this.session;
}
陷阱 5:预取在用户 AbortController 触发后继续烧钱
用户点击"停止生成"时,你的代码会 controller.abort() 取消正式请求。但预取请求有自己的 AbortController,如果没有在用户取消时一并处理,预取还在继续跑。
解法:维护一个全局预取注册表,用户取消时清理所有活跃预取:
class GlobalPrefetchRegistry {
private activeSessions = new Set<PrefetchSession>();
register(session: PrefetchSession) {
this.activeSessions.add(session);
// session 完成时自动清理
session.onComplete(() => this.activeSessions.delete(session));
}
cancelAll() {
this.activeSessions.forEach(s => s.cancel());
this.activeSessions.clear();
}
}
// 用户点击停止时
userAbortButton.addEventListener('click', () => {
normalRequestController.abort();
globalPrefetchRegistry.cancelAll(); // 同时取消所有预取
});
六、什么时候不该用 Prefetching
Prefetching 不是银弹,以下场景不适合:
1. prompt 与上下文强绑定:如果每次请求的 prompt 都包含大量唯一状态(如实时数据、动态 context),预测准确率会极低,净成本为负。
2. 用户行为不可预测:B2B 工具的用户操作路径差异极大,hover-triggered prefetch 命中率可能只有 10-20%,不值得引入复杂度。
3. API 费用敏感阶段:如果你的产品还在 PMF 验证阶段,API 成本每一分都要精打细算,不应该把钱花在命中率不确定的预取上。
4. Provider 并发限制很低:如果你的 API Key 只有 10 RPM,预取会严重影响正式请求。
七、实测效果总结
在一个 Chat UI + 文档生成应用中,综合应用上述四种策略后的实测结果:
| 策略 | 感知延迟改善 | 命中率 | 额外成本 | 适用场景 |
|---|---|---|---|---|
| Context Warm-up | 50-85% | ~100% | <5% | 所有 LLM 应用 |
| Hover Prefetch | 55-70% | 55-70% | 30-45% | 按钮触发的生成 |
| Debounce Prefetch | 40-60% | 40-65% | 35-60% | Chat 输入框 |
| Workflow Prefetch | 80-95% | 80-95% | 5-20% | 多步骤向导 |
Context Warm-up 应该是所有团队的第一选择:改一行代码,加一个 cache_control,零风险地把 TTFT 砍掉一半。然后根据你的产品形态选择 hover 预取或 debounce 预取。
Hover Prefetch 是第二推荐:实现相对简单,对 Chat 和生成类应用效果明显。
Debounce Prefetch 需要配套的意图分类器,引入了额外的复杂度和成本,建议在 hover prefetch 之后再考虑。
Workflow Prefetch 命中率最高,但只在已知流程中有效,不通用。
八、给你的 Checklist
在上 Prefetching 之前,确认以下几点:
- 已用
cache_control开启 Prompt Caching(最低成本的预热) - 预取请求使用独立 API Key,与正式请求完全隔离
- 客户端对每次预取命中/未命中有埋点
- 服务端有预取专用并发队列,上限不超过正式队列的 20%
- 预取 session 绑定了上下文指纹,防止跨上下文污染
- 有每小时预取 Token 预算上限,超出自动关闭预取
- 用户点击停止时,同步取消所有活跃预取 session
- 缓冲区有 token 上限(建议 200 token),防止内存积累
- 告警:命中率 < 35% 时发出告警,额外成本 > 50% 时触发降级
总结
LLM Prefetching 是感知延迟优化的最后一公里。流式输出解决了"响应过程中的等待",Prefetching 解决了"触发请求到收到第一个 token 的等待"。
起点永远是 Context Warm-up:改一行代码,加一个 cache_control,零风险地把 TTFT 砍掉一半。然后根据你的产品形态选择 hover 预取或 debounce 预取。
关键不是实现了多少种策略,而是把命中率监控做好,让数据告诉你哪种策略值得继续投入。