LLM 应用的预取工程实践:用预测性调用把首 Token 延迟砍掉 60%

0 阅读10分钟

一、你的流式输出已经够快了,但用户还是在等

你已经上了 SSE 流式输出,TTFT 稳定在 800ms,P99 也控制在 1.2 秒以内。你觉得延迟优化到头了。

然后你做了一次用户研究,发现有个反复出现的抱怨:

"点击生成按钮之后感觉要等很久才有反应。"

用户描述的"等很久",就是你精心优化过的 800ms。

这个问题不是延迟本身,而是感知延迟。用户点击按钮的那一刻,神经系统进入等待模式,超过 300ms 没有视觉反馈就开始焦虑。你的 800ms 在服务器监控上是绿色,在用户脑子里是红色。

解决感知延迟有两个思路:

  1. 让响应开始得更快(优化 TTFT)
  2. 让响应在用户感知到之前就已经开始(Prefetching)

第一个思路你已经做得很好了。第二个思路,大多数团队没碰过。

这篇文章讲的就是第二个:LLM Prefetching,在用户明确发出请求之前,预测性地提前发起 LLM 调用,把感知延迟从 800ms 压到不足 300ms。


二、Prefetching 不是新概念,但 LLM 场景有特殊挑战

浏览器预取是老技术了。<link rel="prefetch"> 告诉浏览器提前加载下一页的资源。DNS prefetch 提前解析域名。Chrome 的 Speculation Rules API 甚至可以预渲染整个页面。

CPU 里的分支预测器每个时钟周期都在猜你接下来要执行哪条指令。CDN 的 Edge Worker 在用户请求到来之前就已经把缓存预热好了。

这些技术的核心逻辑一致:利用信号预测未来需求,提前准备,消灭等待。

但把 Prefetching 搬进 LLM 场景,有三个静态资源没有的特殊挑战:

挑战 1:预取有真实成本

浏览器预取失败,浪费的是带宽。LLM 预取失败,浪费的是真实 API Token 费用。一次大模型调用输出 500 tokens,约几分钱。看起来小,但如果预测准确率只有 40%,相当于每次有效请求你额外付了 1.5 倍的钱。

挑战 2:响应不可缓存(通常)

静态资源有确定的内容,可以长时间缓存。大模型响应是 prompt + model + temperature 的函数,prompt 里任何一个字变了,缓存就失效了。预取的响应有效期通常只有几十秒。

挑战 3:流式响应的中断与复用

LLM 响应通常是流式的。预取发起后,服务器开始向你传输 token stream。但如果用户没有真正触发请求,这个 stream 需要被取消;如果用户触发了,你需要能直接"接管"这个已经在流动的 stream,而不是重新请求。这比取消一个 HTTP GET 请求复杂得多。


三、四种 Prefetching 策略,按成本从低到高

策略 1:Context Warm-up(上下文预热)

原理:不预取最终响应,只预热大模型需要的 context。

如果你用了国内主流大模型厂商(如通义千问、DeepSeek)的 Prompt Caching 能力,system prompt 和常用 context 可以缓存在 KV Cache 中。预热的意思是:在用户进入某个功能页面时,发一次带 cache_control 标记的"空请求"(或者最轻量的探测请求),让 Provider 的 KV Cache 热起来。

当用户真正触发请求时,system prompt 命中 KV Cache,TTFT 可以降低 50-85%。

// 用户进入文档编辑页面时触发
async function warmupDocumentContext(docId: string) {
  const doc = await fetchDocument(docId);
  
  // 发一次最小 prompt,目的是让 system prompt + doc context 进入 KV Cache
  // 以 DeepSeek / 通义千问兼容接口为例
  await llmClient.messages.create({
    model: "deepseek-chat",
    max_tokens: 1,  // 只要 1 个 token,成本极低
    system: [
      {
        type: "text",
        text: buildSystemPrompt(),
        cache_control: { type: "ephemeral" }
      },
      {
        type: "text", 
        text: `Document context:\n${doc.content}`,
        cache_control: { type: "ephemeral" }
      }
    ],
    messages: [{ role: "user", content: "ready" }]
  });
}

成本:极低。一次 warm-up 消耗约 1 output token + cache write tokens(通常比 input token 便宜 75%)。

收益:后续真实请求命中 KV Cache,TTFT 降低 50-85%。

适用场景:用户进入编辑器、打开文档、进入某个功能模块时触发。


策略 2:Hover-triggered Prefetch(悬停预取)

原理:用户将鼠标悬停在触发元素上 150-200ms 时,预先发起 LLM 请求。

核心洞察:用户 hover 一个按钮超过 150ms,通常意味着他打算点击。这个时间窗口足够发起 LLM 请求并拿到前几十个 token。

// React hook:悬停预取
function useHoverPrefetch(
  prompt: string,
  options: { hoverDelayMs?: number; ttlMs?: number } = {}
) {
  const { hoverDelayMs = 150, ttlMs = 30000 } = options;
  const prefetchRef = useRef<PrefetchSession | null>(null);
  const timerRef = useRef<NodeJS.Timeout>();

  const onMouseEnter = useCallback(() => {
    timerRef.current = setTimeout(() => {
      // 150ms 后用户还没移走,开始预取
      prefetchRef.current = startPrefetch(prompt, { ttlMs });
    }, hoverDelayMs);
  }, [prompt, hoverDelayMs, ttlMs]);

  const onMouseLeave = useCallback(() => {
    clearTimeout(timerRef.current);
    // 用户移走了:取消正在进行的预取(如果还没完成)
    prefetchRef.current?.cancel();
    prefetchRef.current = null;
  }, []);

  const consumePrefetch = useCallback(() => {
    const session = prefetchRef.current;
    prefetchRef.current = null;
    return session; // 调用者接管这个 stream
  }, []);

  return { onMouseEnter, onMouseLeave, consumePrefetch };
}

PrefetchSession 的实现:

interface PrefetchSession {
  stream: ReadableStream<string>;
  buffered: string[];         // 已收到但未消费的 tokens
  status: 'pending' | 'streaming' | 'complete' | 'cancelled';
  cancel: () => void;
}

function startPrefetch(prompt: string, options: { ttlMs: number }): PrefetchSession {
  const controller = new AbortController();
  const buffered: string[] = [];
  let status: PrefetchSession['status'] = 'pending';
  
  // 启动低优先级预取请求
  const responsePromise = fetch('/api/llm/prefetch', {
    method: 'POST',
    headers: { 'Content-Type': 'application/json' },
    body: JSON.stringify({ prompt, priority: 'low' }),
    signal: controller.signal,
  });

  // 开始缓冲 tokens
  responsePromise.then(async (res) => {
    status = 'streaming';
    const reader = res.body!.getReader();
    const decoder = new TextDecoder();
    
    while (true) {
      const { done, value } = await reader.read();
      if (done) { status = 'complete'; break; }
      buffered.push(decoder.decode(value));
    }
  }).catch((err) => {
    if (err.name !== 'AbortError') console.error('Prefetch failed:', err);
  });

  // TTL 到期自动取消
  const ttlTimer = setTimeout(() => {
    if (status !== 'complete') {
      controller.abort();
      status = 'cancelled';
    }
  }, options.ttlMs);

  return {
    stream: createReadableFromBuffer(buffered, responsePromise),
    buffered,
    get status() { return status; },
    cancel: () => {
      clearTimeout(ttlTimer);
      controller.abort();
      status = 'cancelled';
    }
  };
}

命中后如何复用:

function GenerateButton({ prompt }: { prompt: string }) {
  const { onMouseEnter, onMouseLeave, consumePrefetch } = useHoverPrefetch(prompt);

  const handleClick = async () => {
    const prefetched = consumePrefetch();
    
    if (prefetched && prefetched.status !== 'cancelled') {
      // 命中!直接复用已有 stream
      // 用户感知到的延迟 = 0(已有 buffered tokens 直接渲染)
      streamFromPrefetch(prefetched);
    } else {
      // 未命中,走正常请求
      await streamFromFresh(prompt);
    }
  };

  return (
    <button
      onMouseEnter={onMouseEnter}
      onMouseLeave={onMouseLeave}
      onClick={handleClick}
    >
      生成
    </button>
  );
}

实测数据:在 Chat UI 场景,hover prefetch 把感知延迟从约 1.1s 压到 450ms 左右。命中率约 55-70%(取决于用户是否犹豫)。


策略 3:Debounce Prefetch(防抖预取)

原理:用户在输入框中停止输入 300-500ms 后,预测用户可能的完整输入并预取响应。

核心假设:用户停止打字超过 300ms,通常意味着他在思考如何结束这句话,而不是还在输入。

class DebouncePrefetcher {
  private currentPrefetch: PrefetchSession | null = null;
  private debounceTimer: NodeJS.Timeout | null = null;
  private lastInput = '';

  constructor(
    private readonly debounceMs: number = 350,
    private readonly intentClassifier: IntentClassifier,
  ) {}

  onInputChange(input: string) {
    this.lastInput = input;
    
    // 清除旧的 debounce 和预取
    if (this.debounceTimer) clearTimeout(this.debounceTimer);
    this.currentPrefetch?.cancel();
    this.currentPrefetch = null;

    // 输入太短,不值得预取
    if (input.length < 10) return;

    this.debounceTimer = setTimeout(async () => {
      await this.maybePrefetch(input);
    }, this.debounceMs);
  }

  private async maybePrefetch(input: string) {
    // 第一步:用轻量模型做意图分类(避免直接用昂贵模型预取)
    const intent = await this.intentClassifier.classify(input);
    
    // 只有置信度足够高才值得预取
    if (intent.confidence < 0.72) {
      console.debug(`[prefetch] skip: low confidence ${intent.confidence}`);
      return;
    }

    console.debug(`[prefetch] start: "${input.slice(0, 30)}..." (confidence: ${intent.confidence})`);
    this.currentPrefetch = startPrefetch(input, { ttlMs: 45000 });
  }

  consumePrefetch(actualInput: string): PrefetchSession | null {
    if (!this.currentPrefetch) return null;
    
    // 验证实际输入与预取时的输入足够相似
    const similarity = cosineSimilarity(actualInput, this.lastInput);
    if (similarity < 0.85) {
      this.currentPrefetch.cancel();
      this.currentPrefetch = null;
      return null;
    }
    
    const session = this.currentPrefetch;
    this.currentPrefetch = null;
    return session;
  }
}

意图分类器(轻量版),使用国产轻量模型成本更低:

class IntentClassifier {
  // 用 Qwen-Turbo 或 DeepSeek-V3 做意图分类,成本极低
  async classify(input: string): Promise<{ intent: string; confidence: number }> {
    const response = await llmClient.chat.completions.create({
      model: "qwen-turbo",
      max_tokens: 20,
      messages: [{
        role: "user",
        content: `Rate the completeness of this user query on a scale 0-1. 
        Only output a JSON: {"confidence": 0.X}
        Query: "${input}"`
      }]
    });
    
    try {
      return JSON.parse(response.choices[0].message.content!);
    } catch {
      return { intent: 'unknown', confidence: 0 };
    }
  }
}

策略 4:Workflow Prefetch(工作流预取)

原理:已知用户会经历的固定流程(如 onboarding wizard、多步骤表单),在用户完成当前步骤时,预取下一步的 LLM 响应。

// Onboarding 向导:步骤 2 完成时,提前为步骤 3 准备内容
class WorkflowPrefetcher {
  private prefetchCache = new Map<string, PrefetchSession>();

  onStepComplete(step: number, formData: Record<string, unknown>) {
    const nextStep = step + 1;
    const nextPrompt = this.buildStepPrompt(nextStep, formData);
    
    if (!nextPrompt) return; // 没有下一步了
    
    const cacheKey = `step-${nextStep}-${hashObject(formData)}`;
    
    // 取消旧的预取,开始新的
    this.prefetchCache.get(cacheKey)?.cancel();
    this.prefetchCache.set(cacheKey, startPrefetch(nextPrompt, { ttlMs: 120000 }));
    
    console.debug(`[workflow-prefetch] pre-loading step ${nextStep}`);
  }

  consumeForStep(step: number, formData: Record<string, unknown>): PrefetchSession | null {
    const cacheKey = `step-${step}-${hashObject(formData)}`;
    const session = this.prefetchCache.get(cacheKey) ?? null;
    this.prefetchCache.delete(cacheKey);
    return session;
  }
}

工作流预取的命中率通常最高(80-95%),因为下一步的 prompt 几乎是确定的。代价是提前消耗 API 配额,适合高价值用户流程。


四、服务端:预取请求的优先级隔离

客户端的预取策略再好,如果服务端不区分预取请求和正式请求,预取会挤占正式请求的资源。

4.1 请求标记

在 HTTP 头里标记预取请求:

// 客户端发送预取请求时
const response = await fetch('/api/llm/stream', {
  method: 'POST',
  headers: {
    'Content-Type': 'application/json',
    'X-Request-Priority': 'low',       // 低优先级
    'X-Prefetch': 'true',              // 明确标记为预取
    'X-Prefetch-TTL': '30000',         // 预取响应的 TTL
  },
  body: JSON.stringify({ prompt }),
  signal: controller.signal,
});

4.2 服务端队列隔离

// Express/Fastify 中间件:分离预取流量
const prefetchQueue = new PQueue({ concurrency: 2 });   // 最多 2 个并发预取
const normalQueue = new PQueue({ concurrency: 10 });    // 正式请求 10 并发

app.post('/api/llm/stream', async (req, res) => {
  const isPrefetch = req.headers['x-prefetch'] === 'true';
  const queue = isPrefetch ? prefetchQueue : normalQueue;
  
  await queue.add(async () => {
    // 如果是预取请求且队列已满,直接 429
    if (isPrefetch && queue.size > 5) {
      res.status(429).json({ error: 'prefetch_capacity_exceeded' });
      return;
    }
    
    await handleLLMStream(req, res);
  });
});

4.3 Token Budget 隔离

预取请求不计入用户的 Token 配额,但要计入服务端的账单预算:

class PrefetchBudgetGuard {
  private hourlyPrefetchTokens = 0;
  private readonly hourlyBudget: number;
  
  constructor(hourlyBudget: number) {
    this.hourlyBudget = hourlyBudget;
    // 每小时重置
    setInterval(() => { this.hourlyPrefetchTokens = 0; }, 3600 * 1000);
  }

  async checkAndReserve(estimatedTokens: number): Promise<boolean> {
    if (this.hourlyPrefetchTokens + estimatedTokens > this.hourlyBudget) {
      console.warn(`[prefetch-budget] budget exhausted: ${this.hourlyPrefetchTokens}/${this.hourlyBudget}`);
      return false; // 拒绝预取
    }
    this.hourlyPrefetchTokens += estimatedTokens;
    return true;
  }
  
  onPrefetchComplete(actualTokens: number, estimated: number) {
    // 用实际 token 数修正
    this.hourlyPrefetchTokens += (actualTokens - estimated);
  }
}

五、5 个生产陷阱

陷阱 1:预取命中率不监控,成本悄悄翻倍

不监控命中率是最常见的错误。预取系统上线两周后,某团队发现 API 账单增加了 80%,查了半天才发现预取命中率只有 18%——相当于每次有效请求额外付了 4.5 倍的钱。

解法:在每次 consumePrefetch 调用时记录 hit/miss,用监控系统追踪命中率,设置告警阈值(建议 < 35% 时告警)。

// 在 consumePrefetch 中埋点
const session = prefetcher.consumePrefetch(input);
metrics.increment('prefetch.attempt');
if (session && session.status !== 'cancelled') {
  metrics.increment('prefetch.hit');
} else {
  metrics.increment('prefetch.miss');
  if (session?.status === 'cancelled') metrics.increment('prefetch.cancelled');
}

陷阱 2:预取请求消耗 Provider 并发槽位

国内主流大模型 Provider 都有 per-key 的并发 in-flight 限制。如果你的预取系统在高峰期发出 50 个并发预取请求,正式请求可能会被 rate limit 掉。

解法:预取请求必须使用独立的 API Key(或独立的账号),与正式请求的 Key 完全隔离。

// 使用独立的预取专用 client
const prefetchClient = createLLMClient({
  apiKey: process.env.LLM_PREFETCH_API_KEY,  // 专用 key
});

const normalClient = createLLMClient({
  apiKey: process.env.LLM_API_KEY,           // 正式 key
});

陷阱 3:流式复用的缓冲区溢出

预取后用户迟迟不点击,缓冲区一直在累积 tokens。一个 4096 token 的响应,如果全部缓冲在内存里,每个预取 session 占约 16KB(token 平均 4 bytes × 4096)。100 个并发预取 session = 1.6MB,听起来不多,但如果 token 是以 JSON SSE 格式缓冲的,实际内存占用可能是 10-15 倍。

解法:缓冲区设上限,超过阈值后停止缓冲(但不取消请求,依然接收并丢弃 tokens):

const MAX_BUFFER_TOKENS = 200; // 只缓冲前 200 个 tokens

// 在 startPrefetch 中
if (buffered.length >= MAX_BUFFER_TOKENS) {
  // 超过上限:丢弃 token 但保持连接
  // 用户命中时从这里开始直接传流
  continue; 
}
buffered.push(token);

陷阱 4:预取响应在上下文变化后仍被消费

用户在文档 A 上悬停了按钮,触发了预取。然后用户切换到文档 B,点击了按钮。如果消费层不校验上下文,会把文档 A 的响应显示给文档 B 的用户。

解法:预取 session 必须绑定上下文指纹:

interface PrefetchSession {
  contextFingerprint: string; // 上下文的哈希
  // ...
}

function consumePrefetch(currentContext: unknown): PrefetchSession | null {
  if (!this.session) return null;
  
  const currentFingerprint = hashObject(currentContext);
  if (currentFingerprint !== this.session.contextFingerprint) {
    // 上下文变了,丢弃预取
    this.session.cancel();
    this.session = null;
    return null;
  }
  
  return this.session;
}

陷阱 5:预取在用户 AbortController 触发后继续烧钱

用户点击"停止生成"时,你的代码会 controller.abort() 取消正式请求。但预取请求有自己的 AbortController,如果没有在用户取消时一并处理,预取还在继续跑。

解法:维护一个全局预取注册表,用户取消时清理所有活跃预取:

class GlobalPrefetchRegistry {
  private activeSessions = new Set<PrefetchSession>();

  register(session: PrefetchSession) {
    this.activeSessions.add(session);
    // session 完成时自动清理
    session.onComplete(() => this.activeSessions.delete(session));
  }

  cancelAll() {
    this.activeSessions.forEach(s => s.cancel());
    this.activeSessions.clear();
  }
}

// 用户点击停止时
userAbortButton.addEventListener('click', () => {
  normalRequestController.abort();
  globalPrefetchRegistry.cancelAll(); // 同时取消所有预取
});

六、什么时候不该用 Prefetching

Prefetching 不是银弹,以下场景不适合:

1. prompt 与上下文强绑定:如果每次请求的 prompt 都包含大量唯一状态(如实时数据、动态 context),预测准确率会极低,净成本为负。

2. 用户行为不可预测:B2B 工具的用户操作路径差异极大,hover-triggered prefetch 命中率可能只有 10-20%,不值得引入复杂度。

3. API 费用敏感阶段:如果你的产品还在 PMF 验证阶段,API 成本每一分都要精打细算,不应该把钱花在命中率不确定的预取上。

4. Provider 并发限制很低:如果你的 API Key 只有 10 RPM,预取会严重影响正式请求。


七、实测效果总结

在一个 Chat UI + 文档生成应用中,综合应用上述四种策略后的实测结果:

策略感知延迟改善命中率额外成本适用场景
Context Warm-up50-85%~100%<5%所有 LLM 应用
Hover Prefetch55-70%55-70%30-45%按钮触发的生成
Debounce Prefetch40-60%40-65%35-60%Chat 输入框
Workflow Prefetch80-95%80-95%5-20%多步骤向导

Context Warm-up 应该是所有团队的第一选择:改一行代码,加一个 cache_control,零风险地把 TTFT 砍掉一半。然后根据你的产品形态选择 hover 预取或 debounce 预取。

Hover Prefetch 是第二推荐:实现相对简单,对 Chat 和生成类应用效果明显。

Debounce Prefetch 需要配套的意图分类器,引入了额外的复杂度和成本,建议在 hover prefetch 之后再考虑。

Workflow Prefetch 命中率最高,但只在已知流程中有效,不通用。


八、给你的 Checklist

在上 Prefetching 之前,确认以下几点:

  • 已用 cache_control 开启 Prompt Caching(最低成本的预热)
  • 预取请求使用独立 API Key,与正式请求完全隔离
  • 客户端对每次预取命中/未命中有埋点
  • 服务端有预取专用并发队列,上限不超过正式队列的 20%
  • 预取 session 绑定了上下文指纹,防止跨上下文污染
  • 有每小时预取 Token 预算上限,超出自动关闭预取
  • 用户点击停止时,同步取消所有活跃预取 session
  • 缓冲区有 token 上限(建议 200 token),防止内存积累
  • 告警:命中率 < 35% 时发出告警,额外成本 > 50% 时触发降级

总结

LLM Prefetching 是感知延迟优化的最后一公里。流式输出解决了"响应过程中的等待",Prefetching 解决了"触发请求到收到第一个 token 的等待"。

起点永远是 Context Warm-up:改一行代码,加一个 cache_control,零风险地把 TTFT 砍掉一半。然后根据你的产品形态选择 hover 预取或 debounce 预取。

关键不是实现了多少种策略,而是把命中率监控做好,让数据告诉你哪种策略值得继续投入。