依赖服务的可用性直接决定业务可用性,而“接口通不通、证书有没有过期、域名解析是否正常”这类问题,靠服务自身的指标发现不了。业务埋点、运行时指标本就生于服务内部,一旦进程假死或整机宕机,指标链路同步中断,故障反而“隐身”。因此需要一个站在外部视角的探活机制,模拟真实客户端发起请求来验证服务可用。
Blackbox Exporter 是 Prometheus 官方的黑盒探测组件,以独立进程从外部对目标发起 HTTP、TCP、ICMP、DNS、gRPC 探测,把“拨测结果”转化为可采集、可告警的指标,与业务埋点形成互补。
本文从零搭建 Blackbox Exporter、读懂 probe_* 指标、给出可直接落地的告警规则,并延伸到其他 Prober 类型与替代方案。
一、搭建 Blackbox Exporter
1、 blackbox.yml 配置示例
module 是"探测剧本":prober 决定用什么协议,其余字段决定"什么算成功"。
modules:
http_2xx:
prober: http
http:
valid_http_versions: ["HTTP/1.1", "HTTP/2"]
method: GET
preferred_ip_protocol: "ip4"
tls_config:
insecure_skip_verify: true # 忽略证书验证(测试环境可用)
tcp_connect:
prober: tcp
icmp:
prober: icmp
timeout: 5s
icmp:
preferred_ip_protocol: "ip4"
2、启动 Blackbox Exporter
docker run -d --name blackbox_exporter \
--rm \
--restart=no \
--network my-bridge \
-p 9115:9115 \
-v /data/volumes/blackbox:/blackbox \
quay.io/prometheus/blackbox-exporter:latest \
--config.file=/blackbox/blackbox.yml
如果需要通过代理访问,有两种方式:
一是在模块中配置 proxy_url(仅 HTTP 探测生效),可以为不同的探测模块指定不同的代理,参考如下:
modules:
http_2xx_via_proxy:
prober: http
timeout: 10s
http:
proxy_url: "http://proxy-server:3128" # 替换为实际代理地址
# 如果代理需要认证,可嵌入用户名密码:
# proxy_url: "http://username:password@proxy-server:3128"
preferred_ip_protocol: "ip4"
# 建议开启,避免 Exporter 自身解析目标域名导致代理拒绝
skip_resolve_phase_with_proxy: true
二是通过环境变量设置全局代理,所有 HTTP 探测都默认走同一个代理
docker run -d \
--name blackbox_exporter \
-p 9115:9115 \
-e HTTP_PROXY="http://proxy-server:3128" \
-e HTTPS_PROXY="http://proxy-server:3128" \
-e NO_PROXY="localhost,127.0.0.1" \
-v ~/docker/blackbox/config:/config \
prom/blackbox-exporter:latest \
--config.file=/config/blackbox.yml
查看版本:
$ docker exec -it blackbox_exporter blackbox_exporter --version
blackbox_exporter, version 0.28.0 (branch: HEAD, revision: 5a059bee8d8ffa4e75947c5055fb0abeefc582e6)
build user: root@967d444a1ba3
build date: 20251206-13:23:49
go version: go1.25.5
platform: linux/amd64
tags: unknown
3、接入 prometheus
# 每个系统、每个 Prober 类型一个 job
# 这因为 Prometheus 的 params.module 是 job 级参数,无法按 target 区分。
- job_name: 'cim_blackbox_http'
metrics_path: /probe
relabel_configs:
# 当 source_labels 未指定时,replacement 的值会直接赋给 target_label,从而为所有 target 统一添加标签。
- target_label: 'job_alias'
replacement: '测试系统'
- source_labels: [__address__]
target_label: __param_target
- source_labels: [__param_target]
target_label: instance
- target_label: __address__
replacement: blackbox_exporter:9115 # 若在同一 Docker 网络,可用容器名
params:
module: [http_2xx] # 对应 blackbox.yml 中的模块名
static_configs:
- targets:
- https://www.baidu.com/
- https://www.csdn.net/
labels:
svc: blackbox
type: http
4、验证
$ docker exec -it prometheus promtool check config /etc/prometheus/prometheus.yml
Checking /etc/prometheus/prometheus.yml
SUCCESS: 1 rule files found
SUCCESS: /etc/prometheus/prometheus.yml is valid prometheus config file syntax
Checking /etc/prometheus/rules/rules.yml
SUCCESS: 13 rules found
打开 Prometheus 的 Status → Targets,job=~"blackbox.*" 的两个 Target 显示 UP,就完成了。
二、核心指标解析
注意:
probe_*并不是/metrics端点的全局指标- Blackbox Exporter 的
/metrics端点主要暴露的是 Exporter 自身的运行状态指标(如配置重载、构建信息等),与具体的探测目标无关 - 若需验证,则直接请求
/probe端点,如curl "http://localhost:9115/probe?target=https://www.baidu.com&module=http_2xx"
1、通用探测指标(所有 Prober 共有)
1)主要指标
每次探测(无论 HTTP、TCP、ICMP、DNS 还是 gRPC)都会产生以下标准指标:
| 指标名 | 类型 | 含义 |
|---|---|---|
probe_success | Gauge | 唯一的成败判据。1 = 通过 module 的全部断言;0 = 任一条断言失败 |
probe_duration_seconds | Gauge | 本次探测总耗时(秒) |
probe_timeout_seconds | Gauge | 本次探测被分配的超时上限(秒),等于 min(module.timeout, scrape_timeout - 0.5),但搭建的验证环境无该指标 |
probe_dns_lookup_time_seconds | Gauge | 探针本地 DNS 解析耗时(秒)。除 UNIX socket 外都有 |
probe_ip_protocol | Gauge | 实际使用的 IP 协议:4 或 6 |
probe_ip_addr_hash | Gauge | 解析到的 IP 的 FNV-32 哈希,用于检测 IP 变化 |
probe_failed_due_to_regex | Gauge | 因正则匹配/不匹配导致的探测失败,1=是,0=否 |
probe_success是最核心的指标,配合probe_duration_seconds可计算探测成功率与响应延迟。
2)常用 PromQL
# 探测是否成功
probe_success{type="http"}
# 5 分钟成功率
avg_over_time(probe_success[5m]) * 100
# 周期内可用率(1h)
avg_over_time(probe_success[1h]) * 100
# 探测耗时,单位是「秒」
probe_duration_seconds
# 距超时还有多远(> 0.8 说明频繁逼近超时,随时会变成误报)
# 需注意,并不一定有 probe_timeout_seconds 指标
probe_duration_seconds / probe_timeout_seconds
# 探针本地 DNS 解析耗时
probe_dns_lookup_time_seconds
# IP 是否发生变化
changes(probe_ip_addr_hash[1h]) > 0
3)分析要点
up与probe_success的组合是定位第一个岔路口:up == 0→ Prometheus 到 exporter 的问题(exporter 挂了、网络不通、/probe超时);up == 1 && probe_success == 0→ exporter 到目标的问题(目标侧或网络侧);up == 1 && probe_success == 1→ 拨测通过。- "注册即暴露"决定了一个指标是"值为 0"还是"不存在"。 HTTP 的
probe_http_*在函数入口就注册,探测失败时仍会以 0 暴露; 而 TLS 指标只在resp.TLS != nil(HTTPS 且握手成功)时才注册,探测失败时该指标直接不存在。 所以规则要区分== 0与absent(...),否则会出现"目标挂了,证书告警却不响"的盲区。 probe_timeout_seconds是探针实际拿到的超时,但不是一定有该指标,默认等于scrape_timeout - 0.5s。用probe_duration_seconds / probe_timeout_seconds做"濒临超时"预警,比直接给耗时定绝对阈值更通用。probe_ip_addr_hash是 FNV-32 哈希值,只用于比较是否变化,不要试图反解或直接展示。- 每个探针实例的并发能力有限,指标数量随 target 数线性增长:N 个 target × 每 target 约 10~20 条序列。
2、HTTP Prober 专属指标
1)主要指标
HTTP 探测器提供最丰富的指标集,覆盖请求各阶段耗时、响应详情及校验结果:
| 指标名 | 类型 | 含义 |
|---|---|---|
probe_http_status_code | Gauge | 最后一次响应的 HTTP 状态码;探测未拿到响应时为 0 |
probe_http_duration_seconds | GaugeVec | 按阶段拆分的耗时(含所有重定向的累加),标签 phase |
probe_http_content_length | Gauge | 响应 Content-Length(字节) |
probe_http_uncompressed_body_length | Gauge | 解压后的响应体大小(字节) |
probe_http_redirects | Gauge | 重定向次数 |
probe_http_ssl | Gauge | 最后一次请求是否使用 SSL。1 = HTTPS |
probe_http_version | Gauge | 响应使用的 HTTP 版本,数值形式(1.1 / 2 / 3);失败时为 0 |
probe_http_last_modified_timestamp_seconds | Gauge | 响应头 Last-Modified 的 Unix 时间戳(仅当服务端返回该头时才注册) |
probe_failed_due_to_regex | Gauge | 是否因正文正则断言失败 |
probe_failed_due_to_cel | Gauge | 是否因正文 JSON 的 CEL 断言失败(0.27.0+),验证环境无该指标 |
probe_ssl_* / probe_tls_* | HTTPS 且握手成功时才注册 |
probe_http_duration_seconds 的 phase 标签将一次 HTTP 请求拆解为五个阶段,便于定位性能瓶颈:
| phase | 含义 | 偏高说明什么 |
|---|---|---|
resolve | DNS 解析耗时 | 本地 DNS 慢(仅无代理时产生) |
connect | TCP 建连耗时 | 网络/防火墙/目标 accept 队列 |
tls | TLS 握手耗时 | 证书链过大、加密套件协商、CPU |
processing | 服务端处理耗时 | 后端业务慢(最该关注的一段) |
transfer | 响应体传输耗时 | 带宽不足或响应体过大 |
2)常用 PromQL
# ---------- 状态码 ----------
probe_http_status_code{type="http"}
probe_http_status_code{type="http"} >= 500 # 服务端错误
probe_http_status_code{type="http"} == 0 # 没拿到响应(失败原因之一)
# ---------- 内容与协议 ----------
probe_failed_due_to_regex{type="http"} == 1 # 内容断言失败
probe_failed_due_to_cel{type="http"} == 1
probe_http_redirects{type="http"} > 3 # 跳转链过长
probe_http_ssl{type="http"} == 0 # 期望 HTTPS 却走到了 HTTP
probe_http_version{type="http"} < 2 # HTTP/2 未生效(协议降级)
# ---------- TLS / 证书 ----------
probe_ssl_earliest_cert_expiry{type="http"}
(probe_ssl_earliest_cert_expiry - time()) / 86400 # 剩余天数
probe_ssl_earliest_cert_expiry - time() < 0 # 已过期
(probe_ssl_last_chain_expiry_timestamp_seconds - time()) / 86400 # 整条链的剩余天数
probe_tls_version_info{version="TLS 1.3"} # 协商到的版本
probe_tls_cipher_info # 协商到的套件
3)分析要点
- 排查顺序固定为五步:
up→probe_success→probe_http_status_code→probe_failed_due_to_*→ 五个phase。 能连上但probe_success == 0,通常是状态码不在valid_status_codes、或内容断言失败,而不是网络问题。 probe_http_status_code == 0是"没有响应",不是"状态码 0"。要区分 0、4xx、5xx 三种失败语义。phase="processing"才是业务耗时;resolve/connect/tls偏高都属于基础设施问题。把processing / probe_duration_seconds画出来,能一眼区分"后端慢"与"网络慢"。probe_duration_seconds是整个探测的耗时(含 DNS、TCP、TLS、请求、读取、重定向),必然 ≥ 各阶段之和。probe_ssl_last_chain_expiry_timestamp_seconds有个大坑:当 TLS 校验未产生VerifiedChains(例如insecure_skip_verify: true,或自签未被信任)时,它取到的是 Go 的零值时间,Unix 时间戳为-62135596800。若直接算"剩余天数",会得到"证书 190 万年前就过期了"的荒谬结论。用它做告警必须加> 0过滤,或以probe_ssl_earliest_cert_expiry为准。probe_tls_cipher_info只有 HTTP 有,别给 TCP/gRPC 写套件相关的规则。fail_if_not_ssl: true是比告警更前置的防线:一旦证书失效或跳转降级,probe_success立刻为 0,再由probe_http_ssl区分原因。
三、告警规则参考
注意,以下规则存在重叠,实际使用过程中要有取舍,或者通过Alertmanager 抑制规则,避免告警风暴。
1、通用层 probe_common
| 规则名 | 表达式 | for | 级别 | 触发含义 |
|---|---|---|---|---|
ProbeExporterDown | up{job=~"blackbox.*"} == 0 | 1m | critical | Prometheus 无法抓取 Blackbox Exporter,所有拨测数据失真 |
ProbeSuccessRateLow | avg_over_time(probe_success[5m]) * 100 < 99 | 5m | warning | 5 分钟成功率 < 99% |
ProbeSuccessRateCritical | avg_over_time(probe_success[5m]) * 100 < 95 | 5m | critical | 5 分钟成功率 < 95% |
ProbeTargetDown | probe_success == 0 | 2m | critical | 连续 2 分钟探测未通过断言 |
ProbeDurationP95High | quantile_over_time(0.95, probe_duration_seconds[10m]) > 2 | 10m | warning | 10 分钟 P95 探测耗时 > 2s |
2、HTTP 层 probe_http
| 告警名 | 表达式 | for | 级别 | 触发含义 |
|---|---|---|---|---|
HttpStatus5xx | probe_http_status_code{type="http"} >= 500 | 2m | critical | HTTP 返回 5xx |
HttpStatus4xx | probe_http_status_code{type="http"} >= 400 and probe_http_status_code{type="http"} < 500 | 5m | warning | HTTP 返回 4xx |
HttpNoResponse | probe_http_status_code{type="http"} == 0 and probe_success{type="http"} == 0 | 2m | critical | 未获得任何 HTTP 响应(连接/TLS 阶段失败) |
HttpContentCheckFailed | probe_failed_due_to_regex{type="http"} == 1 or probe_failed_due_to_cel{type="http"} == 1 | 2m | critical | 响应内容与预期不符(正则/CEL 断言失败) |
HttpDurationHigh | probe_duration_seconds{type="http"} > 2 | 5m | warning | HTTP 单次探测总耗时 > 2s |
四、Blackbox Exporter 扩展
本节内容没有实际验证,仅供参考。
1、支持的 Prober 类型
| 类型 | 协议/层级 | 适用场景 | 典型模块示例 | 关键注意事项 |
|---|---|---|---|---|
| http | HTTP / HTTPS | 网站、API、Web 服务可用性监控;状态码校验;重定向检查;响应内容匹配;认证接口探测;TLS 证书检查 | http_2xx、http_post_2xx、http_3xx、http_401 | 可配置 method、headers、body、valid_status_codes、fail_if_not_matches_regexp、tls_config 等 |
| tcp | TCP | 端口连通性检查;SSH、MySQL、Redis、SMTP 等 TCP 服务存活探测;简单 Banner/协议交互验证 | tcp_connect、ssh_banner、irc_banner | 可通过 query_response 发送内容并匹配响应;不适用于完整应用层协议解析 |
| icmp | ICMP / 网络层 | Ping 探测;主机或网络层连通性;丢包率与延迟监控 | icmp | 容器需添加 --cap-add=NET_RAW;部分云环境或网络策略可能禁用 ICMP |
| dns | DNS / 应用层 | DNS 服务器可用性;域名解析正确性;特定记录类型查询;解析延迟监控 | dns_lookup、dns_soa | 可配置 query_name、query_type、valid_rcodes;可验证响应内容 |
| grpc | gRPC | gRPC 服务健康检查;微服务存活探测;标准健康检查接口监控 | grpc_health | 基于标准 grpc.health.v1.Health/Check;注意 TLS 或明文连接配置 |
注意:
较早的文档可能只列出前四种,
grpc是后续版本加入的,请确保你使用的镜像版本较新。 websocket Prober 已存在于 master 分支但尚未随任何版本发布,本文不纳入介绍。
2、HTTP 类
modules:
http_2xx: # 基础 HTTP 探测,期望 2xx 状态码
prober: http
http:
valid_status_codes: [] # 默认 2xx
method: GET
http_post_2xx: # 使用 POST 方法探测
prober: http
http:
method: POST
http_3xx: # 期望 3xx 重定向
prober: http
http:
valid_status_codes: [301, 302, 303, 307, 308]
http_401: # 期望 401 未授权(用于验证需认证的端点存在)
prober: http
http:
valid_status_codes: [401]
3、TCP 类
TCP 是"最小"的 Prober:没有任何专属指标。配置参考:
tcp_connect: # 基础 TCP 端口连通性检查
prober: tcp
ssh_banner: # 检查 SSH 服务 Banner
prober: tcp
tcp:
query_response:
- expect: "^SSH-2.0-"
irc_banner: # 检查 IRC 服务交互
prober: tcp
tcp:
query_response:
- send: "NICK prober"
- send: "USER prober prober prober :prober"
- expect: "PING :([^ ]+)"
send: "PONG ${1}"
- expect: "^:[^ ]+ 001"
4、ICMP 类
icmp: # Ping 探测
prober: icmp
icmp:
preferred_ip_protocol: "ip4"
指标清单
| 指标名 | 类型 | 含义 |
|---|---|---|
probe_icmp_duration_seconds | GaugeVec | 按阶段拆分,标签 phase ∈ {resolve、setup、rtt} |
probe_icmp_reply_hop_limit | Gauge | 回包跳数(IPv4 下即 TTL) |
probe_success、probe_duration_seconds、probe_dns_lookup_time_seconds、probe_ip_protocol、probe_ip_addr_hash | 通用 | — |
phase 语义:
| phase | 含义 | 偏高说明什么 |
|---|---|---|
resolve | 目标域名解析 | 本地 DNS 慢 |
setup | 本机 ICMP socket 建立(含原始套接字权限检查) | 本机问题:缺 CAP_NET_RAW、内核限速、探针负载高 |
rtt | 往返时延 | 真正的网络质量 |
常用 PromQL
# 丢包率(5m / 1h)
(1 - avg_over_time(probe_success{type="icmp"}[5m])) * 100
(1 - avg_over_time(probe_success{type="icmp"}[1h])) * 100
# 完全不可达
probe_success{type="icmp"} == 0
# RTT(秒 → 毫秒)
probe_icmp_duration_seconds{type="icmp", phase="rtt"} * 1000
# 本机 ICMP 栈是否异常
probe_icmp_duration_seconds{type="icmp", phase="setup"}
# 路由是否变化(跳数/TTL 变化)
changes(probe_icmp_reply_hop_limit{type="icmp"}[1h]) > 0
# 时延抖动
stddev_over_time(probe_duration_seconds{type="icmp"}[10m]) * 1000
5、DNS 类
dns_lookup: # DNS 解析探测
prober: dns
dns:
query_name: "example.com"
query_type: "A"
valid_rcodes: [0] # NOERROR
- DNS 的 target 是 DNS 服务器地址,被查询的域名写在 module 的
query_name里,不是 target,这点和别的 Prober 相反。
指标清单
| 指标名 | 类型 | 含义 |
|---|---|---|
probe_dns_duration_seconds | GaugeVec | 按阶段拆分,标签 phase ∈ {resolve、connect、request} |
probe_dns_query_succeeded | Gauge | 查询本身是否成功(rcode 在 valid_rcodes 内) |
probe_dns_answer_rrs | Gauge | Answer 段记录条数 |
probe_dns_authority_rrs | Gauge | Authority 段记录条数 |
probe_dns_additional_rrs | Gauge | Additional 段记录条数 |
probe_dns_serial | Gauge | zone 的 SOA serial(仅 query_type: SOA 时注册) |
probe_dns_lookup_time_seconds | 通用 | 注意:这是探针为了连到 DNS 服务器而做的"寻址"解析,不是被查询域名的解析耗时 |
probe_ip_protocol / probe_ip_addr_hash | 通用 | 解析到的 DNS 服务器 IP |
phase 语义:
| phase | 含义 |
|---|---|
resolve | 解析 DNS 服务器地址(即上表中的 probe_dns_lookup_time_seconds) |
connect | 与 DNS 服务器建立连接(UDP 下近似为 0;TCP / DoT 下有意义) |
request | 真正的查询往返耗时(最该关注的一段) |
常用 PromQL
# 查询是否成功(rcode 是否在 valid_rcodes 内)
probe_dns_query_succeeded{type="dns"} == 0
# 成功但没解析出记录(空 Answer 段)
probe_dns_answer_rrs{type="dns"} == 0 and probe_dns_query_succeeded{type="dns"} == 1
# 查询往返耗时
probe_dns_duration_seconds{type="dns", phase="request"}
quantile_over_time(0.95, probe_dns_duration_seconds{type="dns", phase="request"}[10m])
# DoT / TCP 传输的建连耗时
probe_dns_duration_seconds{type="dns", phase="connect"}
# 整体可用率
avg_over_time(probe_success{type="dns"}[5m]) * 100
# 多个 NS 的 zone serial 是否一致(> 1 表示出现了不一致的 serial)
count by (job) (count_values("serial", probe_dns_serial) by (job)) > 1
# SOA serial 原始值
probe_dns_serial
6、gRPC 类
grpc_health: # gRPC 健康检查
prober: grpc
grpc:
# 使用标准 grpc.health.v1.Health/Check
gRPC 探测基于标准健康检查协议,支持 TLS 和明文连接。
指标清单
| 指标名 | 类型 | 含义 |
|---|---|---|
probe_grpc_duration_seconds | GaugeVec | 按阶段拆分,标签 phase ∈ {resolve、check} |
probe_grpc_ssl | Gauge | 是否使用 TLS |
probe_grpc_status_code | Gauge | gRPC 状态码,0 = OK,非 0 即失败 |
probe_grpc_healthcheck_response | GaugeVec | 健康检查响应,标签 serving_status |
probe_ssl_earliest_cert_expiry、probe_ssl_last_chain_expiry_timestamp_seconds、probe_ssl_last_chain_info、probe_tls_version_info | TLS | 握手成功后注册 |
phase 只有两个:resolve(解析目标地址)、check(发起健康检查并等待响应)。
gRPC 探测是通过标准的 gRPC Health Checking Protocol 实现的,因此目标服务必须实现 grpc.health.v1.Health。
常用 PromQL
# 成功与否
probe_success{type="grpc"}
avg_over_time(probe_success{type="grpc"}[5m]) * 100
# gRPC 状态码(0 = OK)
probe_grpc_status_code{type="grpc"} != 0
# 健康检查响应的具体状态
probe_grpc_healthcheck_response{type="grpc"}
probe_grpc_healthcheck_response{type="grpc", serving_status!="SERVING"} == 1
# 健康检查耗时
probe_grpc_duration_seconds{type="grpc", phase="check"}
# 是否走了 TLS
probe_grpc_ssl{type="grpc"} == 0
# 证书
probe_ssl_earliest_cert_expiry{type="grpc"}
probe_tls_version_info{type="grpc"}
五、替代方案
由于日常工作中使用的依赖服务以 HTTP 类为主,因此还可以有一种替代方案,比如增加 Gateway 服务,用于路由依赖服务地址,prometheus 只需监控 Gateway 服务即可,但这会增加请求链路。
六、小结
- 配置侧:一个 Prober 类型对应一个 job,module 中的 prober 字段决定协议,其余字段决定「什么算成功」。
- 排障侧:先用 up 与 probe_success 判断问题出在抓取侧还是目标侧,再结合状态码与 phase 定位具体环节。
- 告警侧:按通用层与 HTTP 层分层配置,重叠规则需取舍或依赖抑制,避免告警风暴。