第6篇 依赖服务探活:Blackbox Exporter 核心指标

10 阅读15分钟

依赖服务的可用性直接决定业务可用性,而“接口通不通、证书有没有过期、域名解析是否正常”这类问题,靠服务自身的指标发现不了。业务埋点、运行时指标本就生于服务内部,一旦进程假死或整机宕机,指标链路同步中断,故障反而“隐身”。因此需要一个站在外部视角的探活机制,模拟真实客户端发起请求来验证服务可用。

Blackbox Exporter 是 Prometheus 官方的黑盒探测组件,以独立进程从外部对目标发起 HTTP、TCP、ICMP、DNS、gRPC 探测,把“拨测结果”转化为可采集、可告警的指标,与业务埋点形成互补。

本文从零搭建 Blackbox Exporter、读懂 probe_* 指标、给出可直接落地的告警规则,并延伸到其他 Prober 类型与替代方案。

一、搭建 Blackbox Exporter

1、 blackbox.yml 配置示例

module 是"探测剧本":prober 决定用什么协议,其余字段决定"什么算成功"。

modules:

  http_2xx:
    prober: http
    http:
      valid_http_versions: ["HTTP/1.1", "HTTP/2"]
      method: GET
      preferred_ip_protocol: "ip4"
      tls_config:
        insecure_skip_verify: true  # 忽略证书验证(测试环境可用)
  tcp_connect:
    prober: tcp
  icmp:
    prober: icmp
    timeout: 5s
    icmp:
      preferred_ip_protocol: "ip4"

2、启动 Blackbox Exporter

docker run -d --name blackbox_exporter \
  --rm \
  --restart=no \
  --network my-bridge \
  -p 9115:9115 \
  -v /data/volumes/blackbox:/blackbox \
  quay.io/prometheus/blackbox-exporter:latest \
    --config.file=/blackbox/blackbox.yml

如果需要通过代理访问,有两种方式:

一是在模块中配置 proxy_url(仅 HTTP 探测生效),可以为不同的探测模块指定不同的代理,参考如下:

modules:
  http_2xx_via_proxy:
    prober: http
    timeout: 10s
    http:
      proxy_url: "http://proxy-server:3128"  # 替换为实际代理地址
      # 如果代理需要认证,可嵌入用户名密码:
      # proxy_url: "http://username:password@proxy-server:3128"
      preferred_ip_protocol: "ip4"
      # 建议开启,避免 Exporter 自身解析目标域名导致代理拒绝
      skip_resolve_phase_with_proxy: true

二是通过环境变量设置全局代理,所有 HTTP 探测都默认走同一个代理

docker run -d \
  --name blackbox_exporter \
  -p 9115:9115 \
  -e HTTP_PROXY="http://proxy-server:3128" \
  -e HTTPS_PROXY="http://proxy-server:3128" \
  -e NO_PROXY="localhost,127.0.0.1" \
  -v ~/docker/blackbox/config:/config \
  prom/blackbox-exporter:latest \
    --config.file=/config/blackbox.yml

查看版本:

$ docker exec -it blackbox_exporter blackbox_exporter --version

blackbox_exporter, version 0.28.0 (branch: HEAD, revision: 5a059bee8d8ffa4e75947c5055fb0abeefc582e6)
  build user:       root@967d444a1ba3
  build date:       20251206-13:23:49
  go version:       go1.25.5
  platform:         linux/amd64
  tags:             unknown

3、接入 prometheus

# 每个系统、每个 Prober 类型一个 job
# 这因为 Prometheus 的 params.module 是 job 级参数,无法按 target 区分。
- job_name: 'cim_blackbox_http'
  metrics_path: /probe
  relabel_configs:
  # 当 source_labels 未指定时,replacement 的值会直接赋给 target_label,从而为所有 target 统一添加标签。
  - target_label: 'job_alias'
    replacement: '测试系统'
  - source_labels: [__address__]
    target_label: __param_target
  - source_labels: [__param_target]
    target_label: instance
  - target_label: __address__
    replacement: blackbox_exporter:9115   # 若在同一 Docker 网络,可用容器名
  params:
    module: [http_2xx]          # 对应 blackbox.yml 中的模块名
  static_configs:
  - targets:
    - https://www.baidu.com/
    - https://www.csdn.net/
    labels:
      svc: blackbox
      type: http

4、验证

$ docker exec -it prometheus promtool check config /etc/prometheus/prometheus.yml
Checking /etc/prometheus/prometheus.yml
  SUCCESS: 1 rule files found
 SUCCESS: /etc/prometheus/prometheus.yml is valid prometheus config file syntax

Checking /etc/prometheus/rules/rules.yml
  SUCCESS: 13 rules found

打开 Prometheus 的 Status → Targets,job=~"blackbox.*" 的两个 Target 显示 UP,就完成了。

二、核心指标解析

注意:

  • probe_* 并不是 /metrics 端点的全局指标
  • Blackbox Exporter 的 /metrics 端点主要暴露的是 Exporter 自身的运行状态指标(如配置重载、构建信息等),与具体的探测目标无关
  • 若需验证,则直接请求 /probe 端点,如 curl "http://localhost:9115/probe?target=https://www.baidu.com&module=http_2xx"

1、通用探测指标(所有 Prober 共有)

1)主要指标

每次探测(无论 HTTP、TCP、ICMP、DNS 还是 gRPC)都会产生以下标准指标:

指标名类型含义
probe_successGauge唯一的成败判据。1 = 通过 module 的全部断言;0 = 任一条断言失败
probe_duration_secondsGauge本次探测总耗时(秒)
probe_timeout_secondsGauge本次探测被分配的超时上限(秒),等于 min(module.timeout, scrape_timeout - 0.5),但搭建的验证环境无该指标
probe_dns_lookup_time_secondsGauge探针本地 DNS 解析耗时(秒)。除 UNIX socket 外都有
probe_ip_protocolGauge实际使用的 IP 协议:4 或 6
probe_ip_addr_hashGauge解析到的 IP 的 FNV-32 哈希,用于检测 IP 变化
probe_failed_due_to_regexGauge因正则匹配/不匹配导致的探测失败,1=是,0=否

probe_success 是最核心的指标,配合 probe_duration_seconds 可计算探测成功率与响应延迟。

2)常用 PromQL

# 探测是否成功
probe_success{type="http"}

# 5 分钟成功率
avg_over_time(probe_success[5m]) * 100

# 周期内可用率(1h)
avg_over_time(probe_success[1h]) * 100

# 探测耗时,单位是「秒」
probe_duration_seconds

# 距超时还有多远(> 0.8 说明频繁逼近超时,随时会变成误报)
# 需注意,并不一定有 probe_timeout_seconds 指标
probe_duration_seconds / probe_timeout_seconds

# 探针本地 DNS 解析耗时
probe_dns_lookup_time_seconds

# IP 是否发生变化
changes(probe_ip_addr_hash[1h]) > 0

3)分析要点

  • up 与 probe_success 的组合是定位第一个岔路口: up == 0 → Prometheus 到 exporter 的问题(exporter 挂了、网络不通、/probe 超时); up == 1 && probe_success == 0 → exporter 到目标的问题(目标侧或网络侧); up == 1 && probe_success == 1 → 拨测通过。
  • "注册即暴露"决定了一个指标是"值为 0"还是"不存在"。 HTTP 的 probe_http_* 在函数入口就注册,探测失败时仍会以 0 暴露; 而 TLS 指标只在 resp.TLS != nil(HTTPS 且握手成功)时才注册,探测失败时该指标直接不存在。 所以规则要区分 == 0 与 absent(...),否则会出现"目标挂了,证书告警却不响"的盲区。
  • probe_timeout_seconds 是探针实际拿到的超时,但不是一定有该指标,默认等于 scrape_timeout - 0.5s。用 probe_duration_seconds / probe_timeout_seconds 做"濒临超时"预警,比直接给耗时定绝对阈值更通用。
  • probe_ip_addr_hash 是 FNV-32 哈希值,只用于比较是否变化,不要试图反解或直接展示。
  • 每个探针实例的并发能力有限,指标数量随 target 数线性增长:N 个 target × 每 target 约 10~20 条序列。

2、HTTP Prober 专属指标

1)主要指标

HTTP 探测器提供最丰富的指标集,覆盖请求各阶段耗时、响应详情及校验结果:

指标名类型含义
probe_http_status_codeGauge最后一次响应的 HTTP 状态码;探测未拿到响应时为 0
probe_http_duration_secondsGaugeVec按阶段拆分的耗时(含所有重定向的累加),标签 phase
probe_http_content_lengthGauge响应 Content-Length(字节)
probe_http_uncompressed_body_lengthGauge解压后的响应体大小(字节)
probe_http_redirectsGauge重定向次数
probe_http_sslGauge最后一次请求是否使用 SSL。1 = HTTPS
probe_http_versionGauge响应使用的 HTTP 版本,数值形式(1.1 / 2 / 3);失败时为 0
probe_http_last_modified_timestamp_secondsGauge响应头 Last-Modified 的 Unix 时间戳(仅当服务端返回该头时才注册)
probe_failed_due_to_regexGauge是否因正文正则断言失败
probe_failed_due_to_celGauge是否因正文 JSON 的 CEL 断言失败(0.27.0+),验证环境无该指标
probe_ssl_* / probe_tls_*HTTPS 且握手成功时才注册

probe_http_duration_seconds 的 phase 标签将一次 HTTP 请求拆解为五个阶段,便于定位性能瓶颈:

phase含义偏高说明什么
resolveDNS 解析耗时本地 DNS 慢(仅无代理时产生)
connectTCP 建连耗时网络/防火墙/目标 accept 队列
tlsTLS 握手耗时证书链过大、加密套件协商、CPU
processing服务端处理耗时后端业务慢(最该关注的一段)
transfer响应体传输耗时带宽不足或响应体过大

2)常用 PromQL

# ---------- 状态码 ----------
probe_http_status_code{type="http"}
probe_http_status_code{type="http"} >= 500          # 服务端错误
probe_http_status_code{type="http"} == 0            # 没拿到响应(失败原因之一)

# ---------- 内容与协议 ----------
probe_failed_due_to_regex{type="http"} == 1         # 内容断言失败
probe_failed_due_to_cel{type="http"}   == 1
probe_http_redirects{type="http"} > 3               # 跳转链过长
probe_http_ssl{type="http"} == 0                    # 期望 HTTPS 却走到了 HTTP
probe_http_version{type="http"} < 2                 # HTTP/2 未生效(协议降级)

# ---------- TLS / 证书 ----------
probe_ssl_earliest_cert_expiry{type="http"}
(probe_ssl_earliest_cert_expiry - time()) / 86400                 # 剩余天数
probe_ssl_earliest_cert_expiry - time() < 0                       # 已过期
(probe_ssl_last_chain_expiry_timestamp_seconds - time()) / 86400  # 整条链的剩余天数
probe_tls_version_info{version="TLS 1.3"}                         # 协商到的版本
probe_tls_cipher_info                                             # 协商到的套件

3)分析要点

  • 排查顺序固定为五步:up → probe_success → probe_http_status_code → probe_failed_due_to_* → 五个 phase。 能连上但 probe_success == 0,通常是状态码不在 valid_status_codes、或内容断言失败,而不是网络问题。
  • probe_http_status_code == 0 是"没有响应",不是"状态码 0"。要区分 0、4xx、5xx 三种失败语义。
  • phase="processing" 才是业务耗时;resolve/connect/tls 偏高都属于基础设施问题。把 processing / probe_duration_seconds 画出来,能一眼区分"后端慢"与"网络慢"。
  • probe_duration_seconds 是整个探测的耗时(含 DNS、TCP、TLS、请求、读取、重定向),必然 ≥ 各阶段之和。
  • probe_ssl_last_chain_expiry_timestamp_seconds 有个大坑:当 TLS 校验未产生 VerifiedChains(例如 insecure_skip_verify: true,或自签未被信任)时,它取到的是 Go 的零值时间,Unix 时间戳为 -62135596800。若直接算"剩余天数",会得到"证书 190 万年前就过期了"的荒谬结论。用它做告警必须加 > 0 过滤,或以 probe_ssl_earliest_cert_expiry 为准。
  • probe_tls_cipher_info 只有 HTTP 有,别给 TCP/gRPC 写套件相关的规则。
  • fail_if_not_ssl: true 是比告警更前置的防线:一旦证书失效或跳转降级,probe_success 立刻为 0,再由 probe_http_ssl 区分原因。

三、告警规则参考

注意,以下规则存在重叠,实际使用过程中要有取舍,或者通过Alertmanager 抑制规则,避免告警风暴。

1、通用层 probe_common

规则名表达式for级别触发含义
ProbeExporterDownup{job=~"blackbox.*"} == 01mcriticalPrometheus 无法抓取 Blackbox Exporter,所有拨测数据失真
ProbeSuccessRateLowavg_over_time(probe_success[5m]) * 100 < 995mwarning5 分钟成功率 < 99%
ProbeSuccessRateCriticalavg_over_time(probe_success[5m]) * 100 < 955mcritical5 分钟成功率 < 95%
ProbeTargetDownprobe_success == 02mcritical连续 2 分钟探测未通过断言
ProbeDurationP95Highquantile_over_time(0.95, probe_duration_seconds[10m]) > 210mwarning10 分钟 P95 探测耗时 > 2s

2、HTTP 层 probe_http

告警名表达式for级别触发含义
HttpStatus5xxprobe_http_status_code{type="http"} >= 5002mcriticalHTTP 返回 5xx
HttpStatus4xxprobe_http_status_code{type="http"} >= 400 and probe_http_status_code{type="http"} < 5005mwarningHTTP 返回 4xx
HttpNoResponseprobe_http_status_code{type="http"} == 0 and probe_success{type="http"} == 02mcritical未获得任何 HTTP 响应(连接/TLS 阶段失败)
HttpContentCheckFailedprobe_failed_due_to_regex{type="http"} == 1 or probe_failed_due_to_cel{type="http"} == 12mcritical响应内容与预期不符(正则/CEL 断言失败)
HttpDurationHighprobe_duration_seconds{type="http"} > 25mwarningHTTP 单次探测总耗时 > 2s

四、Blackbox Exporter 扩展

本节内容没有实际验证,仅供参考。

1、支持的 Prober 类型

类型协议/层级适用场景典型模块示例关键注意事项
httpHTTP / HTTPS网站、API、Web 服务可用性监控;状态码校验;重定向检查;响应内容匹配;认证接口探测;TLS 证书检查http_2xx、http_post_2xx、http_3xx、http_401可配置 method、headers、body、valid_status_codes、fail_if_not_matches_regexp、tls_config 等
tcpTCP端口连通性检查;SSH、MySQL、Redis、SMTP 等 TCP 服务存活探测;简单 Banner/协议交互验证tcp_connect、ssh_banner、irc_banner可通过 query_response 发送内容并匹配响应;不适用于完整应用层协议解析
icmpICMP / 网络层Ping 探测;主机或网络层连通性;丢包率与延迟监控icmp容器需添加 --cap-add=NET_RAW;部分云环境或网络策略可能禁用 ICMP
dnsDNS / 应用层DNS 服务器可用性;域名解析正确性;特定记录类型查询;解析延迟监控dns_lookup、dns_soa可配置 query_name、query_type、valid_rcodes;可验证响应内容
grpcgRPCgRPC 服务健康检查;微服务存活探测;标准健康检查接口监控grpc_health基于标准 grpc.health.v1.Health/Check;注意 TLS 或明文连接配置

注意:

较早的文档可能只列出前四种,grpc 是后续版本加入的,请确保你使用的镜像版本较新。 websocket Prober 已存在于 master 分支但尚未随任何版本发布,本文不纳入介绍。

2、HTTP 类

modules:

  http_2xx:          # 基础 HTTP 探测,期望 2xx 状态码
    prober: http
    http:
      valid_status_codes: []   # 默认 2xx
      method: GET

  http_post_2xx:     # 使用 POST 方法探测
    prober: http
    http:
      method: POST

  http_3xx:          # 期望 3xx 重定向
    prober: http
    http:
      valid_status_codes: [301, 302, 303, 307, 308]

  http_401:          # 期望 401 未授权(用于验证需认证的端点存在)
    prober: http
    http:
      valid_status_codes: [401]

3、TCP 类

TCP 是"最小"的 Prober:没有任何专属指标。配置参考:

  tcp_connect:       # 基础 TCP 端口连通性检查

    prober: tcp

  ssh_banner:        # 检查 SSH 服务 Banner
    prober: tcp
    tcp:
      query_response:
        - expect: "^SSH-2.0-"

  irc_banner:        # 检查 IRC 服务交互
    prober: tcp
    tcp:
      query_response:
        - send: "NICK prober"
        - send: "USER prober prober prober :prober"
        - expect: "PING :([^ ]+)"
          send: "PONG ${1}"
        - expect: "^:[^ ]+ 001"

4、ICMP 类

  icmp:              # Ping 探测
    prober: icmp
    icmp:
      preferred_ip_protocol: "ip4"

指标清单

指标名类型含义
probe_icmp_duration_secondsGaugeVec按阶段拆分,标签 phase ∈ {resolve、setup、rtt}
probe_icmp_reply_hop_limitGauge回包跳数(IPv4 下即 TTL)
probe_success、probe_duration_seconds、probe_dns_lookup_time_seconds、probe_ip_protocol、probe_ip_addr_hash通用—

phase 语义:

phase含义偏高说明什么
resolve目标域名解析本地 DNS 慢
setup本机 ICMP socket 建立(含原始套接字权限检查)本机问题:缺 CAP_NET_RAW、内核限速、探针负载高
rtt往返时延真正的网络质量

常用 PromQL

# 丢包率(5m / 1h)
(1 - avg_over_time(probe_success{type="icmp"}[5m])) * 100
(1 - avg_over_time(probe_success{type="icmp"}[1h])) * 100

# 完全不可达
probe_success{type="icmp"} == 0

# RTT(秒 → 毫秒)
probe_icmp_duration_seconds{type="icmp", phase="rtt"} * 1000

# 本机 ICMP 栈是否异常
probe_icmp_duration_seconds{type="icmp", phase="setup"}

# 路由是否变化(跳数/TTL 变化)
changes(probe_icmp_reply_hop_limit{type="icmp"}[1h]) > 0

# 时延抖动
stddev_over_time(probe_duration_seconds{type="icmp"}[10m]) * 1000

5、DNS 类

  dns_lookup:        # DNS 解析探测
    prober: dns
    dns:
      query_name: "example.com"
      query_type: "A"
      valid_rcodes: [0]   # NOERROR
  • DNS 的 target 是 DNS 服务器地址,被查询的域名写在 module 的 query_name 里,不是 target,这点和别的 Prober 相反。

指标清单

指标名类型含义
probe_dns_duration_secondsGaugeVec按阶段拆分,标签 phase ∈ {resolve、connect、request}
probe_dns_query_succeededGauge查询本身是否成功(rcode 在 valid_rcodes 内)
probe_dns_answer_rrsGaugeAnswer 段记录条数
probe_dns_authority_rrsGaugeAuthority 段记录条数
probe_dns_additional_rrsGaugeAdditional 段记录条数
probe_dns_serialGaugezone 的 SOA serial(仅 query_type: SOA 时注册)
probe_dns_lookup_time_seconds通用注意:这是探针为了连到 DNS 服务器而做的"寻址"解析,不是被查询域名的解析耗时
probe_ip_protocol / probe_ip_addr_hash通用解析到的 DNS 服务器 IP

phase 语义:

phase含义
resolve解析 DNS 服务器地址(即上表中的 probe_dns_lookup_time_seconds)
connect与 DNS 服务器建立连接(UDP 下近似为 0;TCP / DoT 下有意义)
request真正的查询往返耗时(最该关注的一段)

常用 PromQL

# 查询是否成功(rcode 是否在 valid_rcodes 内)
probe_dns_query_succeeded{type="dns"} == 0

# 成功但没解析出记录(空 Answer 段)
probe_dns_answer_rrs{type="dns"} == 0 and probe_dns_query_succeeded{type="dns"} == 1

# 查询往返耗时
probe_dns_duration_seconds{type="dns", phase="request"}
quantile_over_time(0.95, probe_dns_duration_seconds{type="dns", phase="request"}[10m])

# DoT / TCP 传输的建连耗时
probe_dns_duration_seconds{type="dns", phase="connect"}

# 整体可用率
avg_over_time(probe_success{type="dns"}[5m]) * 100

# 多个 NS 的 zone serial 是否一致(> 1 表示出现了不一致的 serial)
count by (job) (count_values("serial", probe_dns_serial) by (job)) > 1

# SOA serial 原始值
probe_dns_serial

6、gRPC 类

  grpc_health:       # gRPC 健康检查
    prober: grpc
    grpc:
      # 使用标准 grpc.health.v1.Health/Check

gRPC 探测基于标准健康检查协议,支持 TLS 和明文连接。

指标清单

指标名类型含义
probe_grpc_duration_secondsGaugeVec按阶段拆分,标签 phase ∈ {resolve、check}
probe_grpc_sslGauge是否使用 TLS
probe_grpc_status_codeGaugegRPC 状态码,0 = OK,非 0 即失败
probe_grpc_healthcheck_responseGaugeVec健康检查响应,标签 serving_status
probe_ssl_earliest_cert_expiry、probe_ssl_last_chain_expiry_timestamp_seconds、probe_ssl_last_chain_info、probe_tls_version_infoTLS握手成功后注册

phase 只有两个:resolve(解析目标地址)、check(发起健康检查并等待响应)。

gRPC 探测是通过标准的 gRPC Health Checking Protocol 实现的,因此目标服务必须实现 grpc.health.v1.Health。

常用 PromQL

# 成功与否
probe_success{type="grpc"}
avg_over_time(probe_success{type="grpc"}[5m]) * 100

# gRPC 状态码(0 = OK)
probe_grpc_status_code{type="grpc"} != 0

# 健康检查响应的具体状态
probe_grpc_healthcheck_response{type="grpc"}
probe_grpc_healthcheck_response{type="grpc", serving_status!="SERVING"} == 1

# 健康检查耗时
probe_grpc_duration_seconds{type="grpc", phase="check"}

# 是否走了 TLS
probe_grpc_ssl{type="grpc"} == 0

# 证书
probe_ssl_earliest_cert_expiry{type="grpc"}
probe_tls_version_info{type="grpc"}

五、替代方案

由于日常工作中使用的依赖服务以 HTTP 类为主,因此还可以有一种替代方案,比如增加 Gateway 服务,用于路由依赖服务地址,prometheus 只需监控 Gateway 服务即可,但这会增加请求链路。

六、小结

  • 配置侧:一个 Prober 类型对应一个 job,module 中的 prober 字段决定协议,其余字段决定「什么算成功」。
  • 排障侧:先用 up 与 probe_success 判断问题出在抓取侧还是目标侧,再结合状态码与 phase 定位具体环节。
  • 告警侧:按通用层与 HTTP 层分层配置,重叠规则需取舍或依赖抑制,避免告警风暴。