第5篇 Nginx 监控:流量、连接与性能指标

0 阅读9分钟

Nginx 在现代架构中往往扮演着双重角色——它既是承载静态资源的 Web 服务器,也是转发动态请求的反向代理服务器。这种双重身份意味着,任何一次性能波动都可能来自两个截然不同的层面:静态文件读取受阻,抑或是上游服务响应迟滞。

正因如此,监控 Nginx 从来不是简单地看一眼 QPS 曲线。我们需要回答三个问题:

  • 流量是否健康——请求速率、状态码分布、带宽消耗是否在预期区间内;
  • 连接是否可控——活跃连接数、等待连接、长连接复用率是否逼近瓶颈;
  • 性能是否达标——请求处理延迟、上游响应时间、缓存命中率是否出现了劣化趋势。

本文将以 nginx 1.31 与 nginx-prometheus-exporter 1.5.0 为基础环境,先厘清 Nginx 的指标体系与采集链路,再逐一拆解流量、连接与性能三条监控主线,最终给出可落地的告警规则与可视化方案。目标是:让你在故障发生之前,就从数据中读出征兆。

本文环境:

  • nginx 1.31(官方 Docker 镜像)
  • nginx-prometheus-exporter 1.5.0
  • prom/prometheus:v3

本篇只讲开源版 nginx。--nginx.plus 模式下的 nginx_plus_* 指标(upstream 耗时、缓存命中、状态码分布)不在讨论范围内。

一、验证环境搭建

Prometheus 的文本协议只认 metric_name{label="value"} 数值 这一种形态。stub_status 输出的是给人看的对齐文本,两者格式不兼容,因此中间需要搭建翻译层 nginx-prometheus-exporter。

1、搭建 Nginx 并开启 stub_status

1)开启 stub_status

server {
       listen 8011;
       server_name localhost;
       location /nginx_status {
           stub_status on;
           access_log off;
#           allow 127.0.0.1;
#           deny all;
       }
}

2)启动 Nginx

docker run -d --name nginx \
  --restart=no \
  --network my-bridge \
  -p 80:80 \
  -p 8011:8011 \
  -e TZ=Asia/Shanghai \
  -v /data/volumes/nginx/nginx.conf:/etc/nginx/nginx.conf \
  -v /data/volumes/nginx/conf.d:/etc/nginx/conf.d \
  nginx:1.31

3)验证

/nginx_status 地址返回示例:

$ curl localhost:8011/nginx_status
Active connections: 2
server accepts handled requests
 2 2 3
Reading: 0 Writing: 1 Waiting: 1

stub_status 一共就 7 个数字,全是 /proc 级别的累计计数器和瞬时值。它是 nginx 监控唯一的零成本数据源,但边界也很明确——状态码、upstream、耗时,一个都没有。

stub_status 字段含义对应 Prometheus 指标类型
Active connections当前已建立的全部连接数nginx_connections_activeGauge
accepts累计接受的连接数nginx_connections_acceptedCounter(无 _total 后缀)
handled累计成功处理的连接数nginx_connections_handledCounter(无 _total 后缀)
requests累计客户端请求数(是请求,不是连接)nginx_http_requests_totalCounter
Reading正在读取请求头的连接数nginx_connections_readingGauge
Writing正在向客户端写响应的连接数nginx_connections_writingGauge
Waiting空闲的 keepalive 连接数nginx_connections_waitingGauge

三条恒等关系:

  • Active = Reading + Writing + Waiting:任何时候都成立,对不上说明抓取瞬间前后不一致。
  • accepts - handled:累计被丢弃的连接数。nginx 的 listen 队列满了才会出现。
  • requests / handled:平均每个连接跑了几个请求。这个比值是 keepalive 有没有生效的直接证据。

2、搭建 nginx-prometheus-exporter

docker run -d --name nginx-exporter \
  --rm \
  --restart no \
  --network my-bridge \
  -p 9113:9113 \
  -e TZ=Asia/Shanghai \
  nginx/nginx-prometheus-exporter:1.5.0 \
    --nginx.scrape-uri=http://nginx:8011/nginx_status

Web 服务相关(Exporter 自身服务)

参数默认值说明
--web.listen-address:9113Exporter 监听地址端口,可多组:--web.listen-address=0.0.0.0:9113
--web.telemetry-path/metricsprometheus 拉取指标路径
--web.config.file""web 配置文件,配置 TLS、basicAuth 认证,exporter‑toolkit 格式
--[no‑]web.systemd‑socketfalse使用 systemd socket 监听(Linux)

Nginx 采集核心参数

参数默认值说明
--nginx.scrape‑urihttp://127.0.0.1:8080/stub_statusNginx 状态接口地址:
开源Nginx:http://ip/stub_status
Plus:http://ip/api
unix socket:unix:/var/run/nginx.sock:/stub_status
--[no‑]nginx.plusfalse启用 Nginx‑Plus 模式,读取 api 接口;普通 Nginx 务必加 --no‑nginx.plus
--nginx.timeout5s请求 nginx status 接口超时时间

3、接入 prometheus

- targets:
  - 10.0.2.15:9113
  labels:
    svc: nginx
    type: nginx
    nginx_port: "80"

验证:

$ docker exec -it prometheus promtool check config /etc/prometheus/prometheus.yml
Checking /etc/prometheus/prometheus.yml
 SUCCESS: /etc/prometheus/prometheus.yml is valid prometheus config file syntax

打开 Prometheus 的 Status → Targets,svc="nginx" 显示 UP,就完成了。

二、核心指标解析

OSS 版 nginx 原生指标只有 8 个,加上 exporter 自带的运行时指标,一共四类。

#分类指标前缀监控内容
1可用性nginx_up、nginx_exporter_build_infostub_status 是否抓得到、exporter 版本
2连接nginx_connections_*活跃/读取/写入/等待连接、累计接受与处理
3流量nginx_http_requests_total累计客户端请求数,唯一能算 QPS 的指标
4exporter 自身process_*、go_*、promhttp_*是 exporter 进程的,不是 nginx 的
5nginx 进程资源原生无,需 node_exporter 或 cAdvisor 补nginx 自身 CPU、内存、文件描述符

1、可用性

主要指标:

Prometheus 指标名类型含义说明
nginx_upGauge上次抓取 stub_status 是否成功。1=成功,0=失败。这是最该配的第一条告警
nginx_exporter_build_infoGauge恒为 1,版本信息在标签里(version、revision、goversion)

nginx_up 为 0 的含义比想象中宽:nginx 挂了是一种,stub_status 返回 403 是另一种,DNS 解析失败、连接超时、TLS 握手失败也都算。它只能告诉你"这一环断了",具体断在哪要配合 up{job="nginx"} 一起看:

  • up == 0:Prometheus 到 exporter 之间的问题。
  • up == 1 且 nginx_up == 0:exporter 到 nginx 之间的问题。

nginx_exporter_build_info 指标示例:

# HELP nginx_exporter_build_info A metric with a constant '1' value labeled by version, revision, branch, goversion from which nginx_exporter was built, and the goos and goarch for the build.
# TYPE nginx_exporter_build_info gauge
nginx_exporter_build_info{branch="HEAD",goarch="amd64",goos="linux",goversion="go1.25.1",revision="b14979c9f3634dcd5a2b158874e713beb3aca3d7",tags="unknown",version="1.5.0"} 1

常用 PromQL:

# 查看 Nginx 状态
nginx_up{job="cim"}

2、连接

主要指标:

Prometheus 指标名类型含义说明
nginx_connections_activeGauge当前已建立的全部连接数,恒等于 reading + writing + waiting
nginx_connections_readingGauge正在读取请求头的连接数,反映请求进入的速度
nginx_connections_writingGauge正在向客户端写响应的连接数,持续偏高说明响应慢、连接被占住
nginx_connections_waitingGauge空闲的 keepalive 长连接数,不消耗 CPU,偏高属正常现象
nginx_connections_acceptedCounter(无 _total 后缀)累计接受的连接数,计算速率必须用 rate(),不能直接画原始值
nginx_connections_handledCounter(无 _total 后缀)累计成功处理的连接数;与 accepted 之差是被丢弃的连接

常用 PromQL:

# 当前活跃连接数
nginx_connections_active

# 现有连接构成(三条曲线相加应等于 active)
nginx_connections_reading
nginx_connections_writing
nginx_connections_waiting

# 连接水位:与理论上限的比值
nginx_connections_active
  / (count(count by (instance) (nginx_connections_active)) * 512)   # 512 为 worker_connections 默认值,按实际改

# 每秒新建连接数
rate(nginx_connections_accepted[5m])

# 每秒被丢弃的连接数(> 0 就要看 listen backlog)
rate(nginx_connections_accepted[5m]) - rate(nginx_connections_handled[5m])

# 平均每连接请求数(keepalive 生效程度,正常情况下 > 1)
rate(nginx_http_requests_total[5m]) / rate(nginx_connections_accepted[5m])

# 正在写响应的连接占比(偏高说明响应慢,连接被占住)
nginx_connections_writing / nginx_connections_active

分析要点:

  • active 高不等于有问题,active 高且 writing 占比高才是问题。前者说明连接多,后者说明连接被慢响应占住了。
  • waiting 高是正常的。keepalive 空闲连接本来就应该停在 waiting 状态,它不用 CPU。看到 waiting 涨就告警,是新手最常见的误判。
  • accepts - handled 增速大于 0,说明有连接在下沉。这个值不会回落,只能看增速。

连接数上限的账要算两遍:worker_processes × worker_connections 是理论值,nginx 默认 worker_connections 512。但 nginx 做反向代理时,一个客户端连接会同时占掉两个槽位(客户端侧一个、upstream 侧一个)。所以真正能承载的客户端连接数大约只有理论值的一半。60 个活跃连接配 512 的上限看着很空,如果 worker_processes 是 2、又在代理 5 个后端,实际余量远比数字上看起来小。

3、流量

主要指标:

Prometheus 指标名类型含义说明
nginx_http_requests_totalCounter累计客户端请求数,OSS 版 nginx 唯一能算 QPS 的指标;按请求计不按连接计,一个 keepalive 连接跑 100 个请求会累加 100

nginx_http_requests_total 是 OSS 版 nginx 唯一的原生流量指标。

常用 PromQL:

# 总 QPS
sum(rate(nginx_http_requests_total[5m])) by (job,instance)

# QPS 环比波动(识别流量异动)
sum(rate(nginx_http_requests_total[5m]))
  / sum(rate(nginx_http_requests_total[5m] offset 1h)) - 1

分析要点:

  • 抓取间隔 15s、rate 窗口 5m,是分辨率和平滑度比较平衡的一组值。窗口用 1m 会抖,用 15m 会盖掉 5 分钟以内的流量尖峰。
  • QPS 本身没有绝对阈值,它取决于容量规划。有意义的判断是环比和同比波动。
  • nginx_http_requests_total 是 Counter,nginx 重启会归零。rate() 会自动处理重置,但 Grafana 里如果直接画原始值会出现一个向下的尖峰,别当成异常。

4、exporter 自身

exporter 自身指标无实际意义,不扩展介绍。

5、nginx 进程资源

/metrics 里的 process_* 描述的是 exporter 进程,不是 nginx。

nginx 自身工作进程已经拆成多个 worker,用 exporter 的 process_* 去告警 nginx 的 CPU 和内存,结论一定是错的。要拿 nginx 真实的进程资源,有两条路:

方案关键指标优点缺点适用场景
node_exporter 进程采集器namedprocess_namegroup_cpu_seconds_total、namedprocess_namegroup_memory_bytes、namedprocess_namegroup_open_filedesc、namedprocess_namegroup_num_procs无侵入、按进程名聚合、能拿到 worker 总资源指标量大,--collector.processes 默认关闭需显式打开物理机 / 虚拟机裸部署
cAdvisor(容器指标)container_cpu_usage_seconds_total、container_memory_working_set_bytes直接对应 cgroup、容器场景天然可用只对容器有效Docker / K8s 部署

实际场景中几乎不会使用,且默认的 /metrics 里也没有以上指标,因此,此处也不扩展介绍。

三、告警规则参考

#分类告警名称 (alert)级别for触发条件 / 阈值说明
1可用性NginxExporterDowncritical1mup{job="nginx"} == 0Prometheus 到 exporter 链路断
2可用性NginxStubStatusUnavailablecritical1mnginx_up == 0exporter 到 nginx 链路断,多见于 403
3连接NginxActiveConnectionsHighwarning5mactive / 上限 > 70%连接水位预警(建议值)
4连接NginxActiveConnectionsCriticalcritical5mactive / 上限 > 90%接近上限,新连接开始排队
5连接NginxConnectionDroppedwarning5m丢弃速率 > 0.5/saccepts - handled 增速转正,listen backlog 溢出
6连接NginxWritingRatioHighwarning5mwriting/active > 0.5 且 active > 50写响应连接占比过高,响应变慢
7连接NginxKeepaliveIneffectiveinfo15mrequests/accepted < 1.2keepalive 未生效,连接复用率低(建议值 1.2)
8流量NginxQpsAnomalywarning10mQPS 环比波动 > 50% 且 QPS > 10流量暴涨或暴跌(建议值)

四、小结

本文围绕 Nginx 监控梳理了三条主线:流量、连接与性能。流量层面,QPS、带宽与状态码分布反映请求的规模与质量;连接层面,活跃连接数、等待队列与握手开销揭示服务端的承压状态;性能层面,响应延迟、上游耗时与缓存命中率则刻画请求处理的效率边界。三者相互关联,单独观察任一维度都难以定位问题全貌——流量陡增未必是故障,但若伴随连接堆积与延迟抬升,则往往指向真实的容量瓶颈。

监控体系的有效性,不取决于采集指标的数量,而取决于是否形成**“采集—展示—告警”的闭环**:采集保证数据可追溯,展示保证异常可感知,告警保证响应可执行。任一环节缺位,监控便退化为事后翻查的日志工具。

实践中,建议遵循三条原则:指标精简优先于大而全,告警阈值需结合业务基线动态校准,避免静态阈值引发的告警疲劳;同时,将监控项与容量规划、故障复盘联动,让数据真正参与决策,而非停留在仪表盘之上。