LLM 应用的 HTTP 连接池工程实践:配置错了,高并发下 AI 服务会悄悄变慢

0 阅读6分钟

一句话摘要:LLM API 调用不是普通 HTTP 请求,它的长超时、流式响应和高并发特性会让默认连接池配置在生产中悄悄制造 P99 延迟劣化。本文从真实事故出发,系统讲解连接池核心参数、Keep-Alive 陷阱、连接泄漏检测与高并发调优实践。


背景:一次"莫名其妙"的 P99 劣化

某团队的 AI 写作助手在日活 5 万后开始出现奇怪的现象:平均延迟 1.2s,P99 却飙到 18s,而且这个问题只在早晚高峰出现,低峰期完全正常。

他们排查了一圈:模型 API 端的延迟监控显示正常,自己服务的 CPU/Memory 没异常,日志里没有明显报错。最后用 netstat 一看:

$ netstat -an | grep 443 | grep -E 'ESTABLISHED|CLOSE_WAIT|TIME_WAIT' | wc -l
847

CLOSE_WAIT 状态的连接积压了 600+ 个。问题找到了——连接池配置完全没针对 LLM 场景调整过,用的是框架默认值。

这类问题的典型特征:

  • 表现为 P99/P999 劣化,平均延迟看起来正常
  • 高并发时才出现,压测低并发不复现
  • 监控里看不到明显错误,只是"慢"

为什么 LLM 调用对连接池特别敏感

LLM API 调用和普通 REST 请求有三个核心差异,每一个都会影响连接池行为:

1. 超时时间极长

普通 API:超时 1-5 秒
LLM API(非流式):超时 30-120 秒
LLM API(流式,等第一个 token):超时 30-60 秒
LLM API(流式,总持续时间):可能 3-10 分钟

超时越长,连接被"占用"的时间越长。如果连接池上限是 10,10 个并发请求就能把连接池打满,后续请求开始排队等待。

2. 流式响应占用连接更久

普通请求:连接占用时间 ≈ 服务端处理时间(100ms~2s)
流式请求:连接占用时间 ≈ 服务端处理时间 + 全部 token 传输时间(可能 20-60s)

一个流式请求,从发出到读完最后一个 token,连接始终被占用。如果你的应用有 20% 的用户在用流式模式,这 20% 的请求会消耗不成比例的连接资源。

3. 并发峰值集中

AI 写作、AI 搜索这类应用有明显的早晚高峰。用户集中在同一时间触发请求,同时打满连接池的概率远高于分散均匀的微服务场景。


连接池的核心参数

以 Python httpx 为例(Node.js undici、Go net/http 的参数名不同,但概念一致):

import httpx

# 默认配置(危险!)
client = httpx.AsyncClient()  # 等价于下方注释的默认值

# 显式配置(推荐)
client = httpx.AsyncClient(
    limits=httpx.Limits(
        max_connections=100,        # 总连接上限(默认 100,但要结合实际调整)
        max_keepalive_connections=20, # Keep-Alive 连接池大小(默认 20)
        keepalive_expiry=30,        # Keep-Alive 连接的最大空闲时间(秒)
    ),
    timeout=httpx.Timeout(
        connect=5.0,    # TCP 握手超时
        read=120.0,     # 等待响应数据超时(非流式要设长)
        write=10.0,     # 发送请求体超时
        pool=10.0,      # 等待从连接池获取连接的超时(极重要!)
    ),
)

最容易被忽略的是 pool 超时:它控制"排队等待连接池空出一个连接"的最大时间。如果不设,默认可能是 None(无限等待),高并发时请求会无限堆积,表现为 P99 劣化但没有报错。

Node.js undici 配置

import { Pool } from 'undici';
import OpenAI from 'openai'; // 示例:替换为你实际使用的 SDK

// undici Pool 是大模型 SDK Node.js 版本的底层实现
const pool = new Pool('https://api.your-llm-provider.com', {
  connections: 50,              // 最大连接数
  pipelining: 1,                // LLM 场景建议为 1(不用 pipeline)
  keepAliveTimeout: 30_000,     // Keep-Alive 超时(ms)
  keepAliveMaxTimeout: 600_000, // Keep-Alive 最长存活(ms,流式场景设长)
  headersTimeout: 10_000,       // 等待响应头超时(ms)
  bodyTimeout: 120_000,         // 等待响应体超时(非流式,ms)
  connectTimeout: 5_000,        // TCP 连接超时(ms)
});

// 将 Pool 作为 fetch 的底层传给 SDK
const client = new OpenAI({
  baseURL: 'https://api.your-llm-provider.com/v1',
  fetch: (url, options) => pool.fetch(url, options),
});

Go net/http 配置

import (
    "net/http"
    "time"
    "github.com/anthropics/anthropic-sdk-go"
)

transport := &http.Transport{
    MaxIdleConns:        100,          // 全局最大 Keep-Alive 连接
    MaxIdleConnsPerHost: 50,           // 每个 Host 的最大 Keep-Alive 连接
    MaxConnsPerHost:     100,          // 每个 Host 的最大总连接(含活跃)
    IdleConnTimeout:     90 * time.Second, // Keep-Alive 空闲超时
    TLSHandshakeTimeout: 5 * time.Second,
    ResponseHeaderTimeout: 30 * time.Second, // 等响应头超时
    // DisableKeepAlives: false,  // 默认 false,不要改成 true
}

httpClient := &http.Client{
    Transport: transport,
    Timeout:   0, // 流式场景设 0 = 无总超时,靠上层 context 控制
}

client := anthropic.NewClient(
    option.WithHTTPClient(httpClient),
)

Keep-Alive 的三个常见陷阱

陷阱 1:服务端先关闭连接,客户端不知道

这是 CLOSE_WAIT 积压的直接原因。

LLM API 服务端通常有自己的 Keep-Alive 超时(比如 60 秒不活动就关连接)。客户端的连接池以为连接还活着,把它放在池里复用,但实际上服务端已经关闭了。下一次用这个连接发请求时:

客户端发 TCP 段 → 服务端返回 RST(连接已关)→ 客户端收到 ConnectionResetError

解决方案:客户端的 keepalive_expiry 要比服务端的超时短 10-20 秒。

如果你不知道服务端的 Keep-Alive 超时是多少,用保守值:

# httpx:设置 20 秒(通常比 API 服务端的 30-60s 超时短)
limits=httpx.Limits(keepalive_expiry=20)
// undici:设置 20 秒
keepAliveTimeout: 20_000

陷阱 2:流式响应期间连接被误判为空闲

某些连接池实现会把"正在等待下一个 SSE chunk"的连接误判为"空闲超时",提前关闭它,导致流式响应中断。

复现方式:

async with client.stream("POST", url, json=payload) as response:
    async for chunk in response.aiter_bytes():
        # 模拟慢消费(客户端处理 chunk 耗时)
        await asyncio.sleep(2)  # 如果 keepalive_expiry < 2,连接可能被杀掉
        process(chunk)

解决方案:流式请求的 keepalive_expiry 要比预期的 chunk 间隔长,或者为流式请求单独维护一个 client 实例:

# 非流式 client:激进配置,快速回收
sync_client = httpx.AsyncClient(
    limits=httpx.Limits(
        max_keepalive_connections=20,
        keepalive_expiry=20,
    ),
    timeout=httpx.Timeout(read=60.0, pool=5.0),
)

# 流式 client:保守配置,允许长连接
stream_client = httpx.AsyncClient(
    limits=httpx.Limits(
        max_keepalive_connections=10,
        keepalive_expiry=300,  # 5 分钟,覆盖长流
    ),
    timeout=httpx.Timeout(read=None, pool=10.0),  # read=None 表示不超时
)

陷阱 3:连接泄漏——用了但没还

流式响应最容易发生连接泄漏。原因是读取流时抛了异常,但没有正确关闭连接:

# ❌ 危险写法:异常时连接可能泄漏
async with client.stream("POST", url, json=payload) as response:
    async for chunk in response.aiter_bytes():
        result = json.loads(chunk)  # 如果这里抛 JSONDecodeError,连接泄漏!
        process(result)

# ✅ 正确写法:确保异常时也关闭连接
try:
    async with client.stream("POST", url, json=payload) as response:
        async for chunk in response.aiter_bytes():
            try:
                result = json.loads(chunk)
                process(result)
            except json.JSONDecodeError:
                logger.warning("Invalid JSON chunk, skipping")
                continue
except httpx.ReadTimeout:
    logger.error("Stream read timeout")
    raise

httpx 的 async with client.stream(...) 实际上会在上下文管理器退出时调用 response.aclose(),但如果你在 async for 外层 break 了,要手动关:

response = await client.send(request, stream=True)
try:
    async for chunk in response.aiter_bytes():
        if should_stop:
            break  # break 不会自动关闭!
        process(chunk)
finally:
    await response.aclose()  # 必须显式关闭

连接泄漏的检测方法

方法 1:暴露连接池状态指标

# httpx 提供了连接池状态查询
import httpx
import asyncio

client = httpx.AsyncClient(
    limits=httpx.Limits(max_connections=50, max_keepalive_connections=20)
)

async def get_pool_metrics():
    pool = client._transport._pool
    return {
        "active": len(pool._requests),         # 正在使用的连接数
        "keepalive": len(pool._keepalive_connections),  # Keep-Alive 池里的连接数
        "connecting": len([c for c in pool._connections if c._connect_failed is False]),
    }

# 定期采集,写入 Prometheus metrics
async def collect_pool_metrics():
    while True:
        metrics = await get_pool_metrics()
        gauge_active_connections.set(metrics["active"])
        gauge_keepalive_connections.set(metrics["keepalive"])
        await asyncio.sleep(10)

方法 2:操作系统级监控

# 查看进程的连接状态分布
PID=$(pgrep -f "python app.py")
ss -p -n | grep "pid=$PID" | awk '{print $2}' | sort | uniq -c | sort -rn

# 输出示例:
# 847 CLOSE_WAIT   ← 这个数量在增长就是泄漏
#  23 ESTABLISHED
#  12 TIME_WAIT

# 实时监控
watch -n 5 "ss -p -n | grep 'pid=$PID' | awk '{print \$2}' | sort | uniq -c"

方法 3:周期性连接数告警

import asyncio
import logging

async def connection_leak_detector(client: httpx.AsyncClient, threshold: int = 40):
    """当 CLOSE_WAIT 连接超过阈值时告警"""
    while True:
        await asyncio.sleep(30)
        try:
            pool = client._transport._pool
            # httpx 内部 API,版本升级可能变化,做好异常处理
            active = len(getattr(pool, '_requests', []))
            if active > threshold:
                logging.warning(
                    f"Connection pool pressure: {active} active connections "
                    f"(threshold={threshold}). Possible leak or overload."
                )
        except Exception:
            pass  # 内部 API 访问失败不影响主逻辑

高并发场景的连接池调优

基准:按 QPS 和平均持续时间计算所需连接数

最小连接数 = QPS × 平均响应时间(秒)
安全系数 × 1.5(应对峰值)

示例:

  • QPS = 50
  • 平均响应时间 = 3s(非流式)
  • 最小连接数 = 50 × 3 = 150
  • 加安全系数:150 × 1.5 = 225

如果你的连接池 max_connections=100,50 QPS 下就会出现排队等待,P99 飙升。

流式场景调整:

  • 流式平均持续时间 = 20s(输出 1000 tokens @ 50 token/s)
  • QPS = 20(流式并发通常低于非流式)
  • 最小连接数 = 20 × 20 = 400

这意味着流式场景需要显著更大的连接池,或者对流式请求的并发数做独立限流:

# 对流式请求独立限流,避免它们耗尽连接池
stream_semaphore = asyncio.Semaphore(30)  # 最多 30 个并发流式请求

async def streaming_llm_call(prompt: str):
    async with stream_semaphore:
        async with stream_client.stream("POST", url, json={...}) as response:
            async for chunk in response.aiter_bytes():
                yield chunk

实际调优的参数清单

参数默认值非流式推荐流式推荐说明
max_connections100QPS×响应时间×1.5QPS×响应时间×2按计算公式设
max_keepalive_connections2050~10010~30流式长连接占资源,少设
keepalive_expiry5s20~30s120~300s短于服务端超时 10s
pool 超时None5~10s10~20s必须设,防止无限排队
read 超时5s60~120sNone流式设 None,靠 context 控

连接池监控 Dashboard 关键指标

# 建议暴露的 Prometheus metrics
from prometheus_client import Gauge, Histogram

llm_pool_active = Gauge('llm_pool_active_connections', 'Active LLM connections')
llm_pool_keepalive = Gauge('llm_pool_keepalive_connections', 'Keepalive LLM connections')
llm_pool_wait_time = Histogram(
    'llm_pool_wait_seconds',
    'Time waiting for a connection from the pool',
    buckets=[0.01, 0.05, 0.1, 0.5, 1.0, 5.0, 10.0]
)

# 在请求包装器里采集
async def llm_request_with_metrics(client, *args, **kwargs):
    wait_start = time.monotonic()
    # pool 超时会在这里抛 PoolTimeout,记录为等待时间异常
    response = await client.request(*args, **kwargs)
    llm_pool_wait_time.observe(time.monotonic() - wait_start)
    return response

一个完整的生产级 LLM Client 封装

将上述实践整合成可直接使用的封装:

import asyncio
import httpx
import logging
import time
from contextlib import asynccontextmanager
from typing import AsyncIterator
from prometheus_client import Gauge, Histogram, Counter

logger = logging.getLogger(__name__)

# Prometheus metrics
pool_active = Gauge('llm_http_pool_active', 'Active connections')
pool_wait = Histogram('llm_http_pool_wait_seconds', 'Pool wait time',
                      buckets=[.01, .05, .1, .5, 1., 5., 10.])
pool_timeout_total = Counter('llm_http_pool_timeout_total', 'Pool timeout count')
stream_leak_total = Counter('llm_http_stream_leak_total', 'Stream connection leak events')


class LLMHttpClient:
    """生产级 LLM HTTP 客户端,内置连接池管理、泄漏检测与 metrics"""

    def __init__(
        self,
        base_url: str,
        api_key: str,
        max_connections: int = 100,
        max_keepalive: int = 30,
        keepalive_expiry: float = 25.0,
        pool_timeout: float = 8.0,
        read_timeout: float = 90.0,
        stream_max_connections: int = 40,
        stream_keepalive_expiry: float = 300.0,
    ):
        self.base_url = base_url
        self.headers = {
            "Authorization": f"Bearer {api_key}",
            "Content-Type": "application/json",
        }

        # 非流式 client:快速回收,严格超时
        self._sync_client = httpx.AsyncClient(
            base_url=base_url,
            headers=self.headers,
            limits=httpx.Limits(
                max_connections=max_connections,
                max_keepalive_connections=max_keepalive,
                keepalive_expiry=keepalive_expiry,
            ),
            timeout=httpx.Timeout(
                connect=5.0,
                read=read_timeout,
                write=10.0,
                pool=pool_timeout,
            ),
        )

        # 流式 client:宽松 keepalive,read 超时由上层 context 控制
        self._stream_client = httpx.AsyncClient(
            base_url=base_url,
            headers=self.headers,
            limits=httpx.Limits(
                max_connections=stream_max_connections,
                max_keepalive_connections=10,
                keepalive_expiry=stream_keepalive_expiry,
            ),
            timeout=httpx.Timeout(
                connect=5.0,
                read=None,   # 流式不设 read 超时,靠 asyncio.timeout 控制
                write=10.0,
                pool=pool_timeout + 5.0,
            ),
        )

        # 流式并发限制
        self._stream_sem = asyncio.Semaphore(stream_max_connections)

    async def post(self, path: str, json: dict) -> dict:
        """非流式请求"""
        wait_start = time.monotonic()
        try:
            resp = await self._sync_client.post(path, json=json)
            pool_wait.observe(time.monotonic() - wait_start)
            resp.raise_for_status()
            return resp.json()
        except httpx.PoolTimeout:
            pool_timeout_total.inc()
            logger.error("LLM pool timeout on non-stream request")
            raise

    @asynccontextmanager
    async def stream(
        self, path: str, json: dict, total_timeout: float = 120.0
    ) -> AsyncIterator[httpx.Response]:
        """流式请求,内置并发限制和泄漏防护"""
        async with self._stream_sem:
            try:
                async with asyncio.timeout(total_timeout):
                    async with self._stream_client.stream(
                        "POST", path, json=json
                    ) as response:
                        response.raise_for_status()
                        yield response
            except httpx.PoolTimeout:
                pool_timeout_total.inc()
                logger.error("LLM pool timeout on stream request")
                raise
            except Exception:
                stream_leak_total.inc()  # 非正常退出计数
                raise

    async def aclose(self):
        await self._sync_client.aclose()
        await self._stream_client.aclose()

    async def health(self) -> dict:
        """连接池健康状态,用于 /healthz 接口"""
        def _pool_info(client):
            try:
                pool = client._transport._pool
                return {
                    "active": len(getattr(pool, '_requests', [])),
                    "keepalive": len(getattr(pool, '_keepalive_connections', [])),
                }
            except Exception:
                return {"active": -1, "keepalive": -1}

        return {
            "sync_pool": _pool_info(self._sync_client),
            "stream_pool": _pool_info(self._stream_client),
        }


# 单例,作为应用级共享 client
_llm_client: LLMHttpClient | None = None

def get_llm_client() -> LLMHttpClient:
    global _llm_client
    if _llm_client is None:
        raise RuntimeError("LLMHttpClient not initialized. Call init_llm_client() first.")
    return _llm_client

def init_llm_client(api_key: str, **kwargs) -> LLMHttpClient:
    global _llm_client
    _llm_client = LLMHttpClient(
        base_url="https://api.your-llm-provider.com",
        api_key=api_key,
        **kwargs,
    )
    return _llm_client

事故复盘:参数调整前后对比

回到文章开头的案例,这是他们调整前后的连接池配置和效果:

调整前(httpx 默认值):

client = httpx.AsyncClient()
# max_connections=100, max_keepalive_connections=20
# keepalive_expiry=5s, pool_timeout=None(无限等待)
# read_timeout=5s(LLM 经常超过 5s!)

问题:

  1. read_timeout=5s 导致大量超时报错(但被业务层重试掩盖了)
  2. pool_timeout=None 导致高并发时请求无限排队,P99 飙升
  3. keepalive_expiry=5s 比服务端短太多,大量连接在池里就已经失效

调整后:

client = httpx.AsyncClient(
    limits=httpx.Limits(
        max_connections=200,
        max_keepalive_connections=50,
        keepalive_expiry=25,  # 短于服务端 ~30s
    ),
    timeout=httpx.Timeout(
        connect=5.0,
        read=90.0,   # 覆盖非流式最长响应时间
        write=10.0,
        pool=8.0,    # 超过 8s 等不到连接就报错,不再无限排队
    ),
)

效果:

指标调整前调整后
P99 延迟(高峰)18s2.1s
CLOSE_WAIT 连接数600+<10
连接池等待超时报错0(无限等待变成堆积)偶发,有告警
平均延迟1.2s1.1s

P99 从 18s 降到 2.1s,平均延迟基本不变——这是连接池问题的典型特征:平均值正常,尾部极差。


总结

LLM HTTP 连接池和普通服务的连接池有几个关键差异需要特别对待:

  1. 超时全家桶都要设:connect、read、write、pool 四个维度,pool 超时最容易被漏掉
  2. 流式和非流式分开配置:它们对 keepalive_expiry 和 read_timeout 的要求完全相反
  3. keepalive_expiry 要短于服务端超时:否则复用失效连接,引发大量 ConnectionReset
  4. 按 QPS×响应时间 计算连接数上限:LLM 响应时间长,需要的连接数远超普通 API
  5. 流式请求必须显式关闭连接:break、异常都可能造成泄漏,用 finally + aclose() 保护

连接池配置错误不会立刻崩溃,它会以 P99 劣化的形式慢慢侵蚀用户体验,直到某次流量高峰才彻底暴露。提前建立监控、按场景调优,比事后排查要划算得多。


本文配套代码已整合到 LLMHttpClient 封装,可直接用于生产。