Open table of contents
Who observes this Prometheus when it stops?
Running Prometheus, Alertmanager and Grafana on Kubernetes with multiple replicas, anti-affinity and disruption budgets can improve resilience to individual pod or node failures.
But when monitoring shares networking, power, DNS or cloud-account dependencies with its targets, a common failure can remove observation and notification together. Preserving replicas through HA and preserving an observation point outside a failure domain are separate design problems.
A control-plane outage does not necessarily stop every existing pod immediately. Conversely, healthy control-plane components do not guarantee user access through ingress or public DNS. Replace the single label “cluster down” with explicit unavailable functions and observation paths.
Following From Monitoring to Observability, this article examines shared failures of the observation platform itself.
The configuration and review bundle describes an imaginary production cluster and external VM. Prometheus, Blackbox Exporter and Alertmanager configuration validation, probing, notifications, heartbeat reception and failure injection were not executed. Everything remains NOT_RUN. The bundle neither provisions a complete system nor reports measured detection performance.
1. Separate HA from failure-domain independence
Prometheus instances on separate nodes can survive some node failures. Alertmanager HA shares silences and notification state and favors duplicate notifications over missed notifications during a partition. Prometheus should send alerts to every Alertmanager instance rather than choose one through a load balancer. [1]
That does not guarantee delivery when all outbound connectivity is lost. PodDisruptionBudgets primarily govern voluntary disruption; they do not prevent site-wide power failure.
| Common failure to detect | Separation to review |
|---|---|
| Pod or node loss | Placement, host, volume and node networking |
| Cluster infrastructure failure | Another cluster or VM with independent startup and management |
| VPC or site connectivity loss | Network, upstream, VPN gateway and site |
| Account or administrative failure | Account, IAM, authentication, billing and operational authority |
| Notification-provider failure | Provider, credentials, delivery and escalation |
Another VM or provider is not automatically independent. DNS, identity, secret retrieval and deployment automation may still be shared. Work backward from the faults to detect. Different placement alone also does not justify treating failure probabilities as independent and multiplying them.
2. Keep internal diagnosis and add external observation
Retain monitoring inside the cluster. Service discovery and internal metrics help explain pod, storage, application and dependency behavior. Add a small external sentinel alongside that diagnostic system.
Scroll horizontally to view the complete diagram.
Read the diagram as text
Production provides diagnosis; an external sentinel probes service and monitoring endpoints and uses an independent pager. Each monitoring system sends a distinct heartbeat to an external receiver. Placement alone does not prove independence.
The sentinel checks important endpoints, internal monitoring readiness and a notification route that bypasses production. It need not duplicate Loki, Thanos and Grafana. A VM running Prometheus with short retention, Blackbox Exporter and Alertmanager is one candidate.
A single sentinel remains a single point of failure. Send its own heartbeat to an external receiver using a stream distinct from production’s. Record the failure domain delegated to that receiver and its notification path. A managed service does not eliminate residual provider or human-delivery risk.
Sharing an IaC repository differs from sharing runtime dependencies. If rebuilding through NixOS or OpenTofu, verify access to source, artifacts and credentials while production is unavailable.
3. Define what each probe path establishes
Public probes follow a user-facing route through DNS, TLS, CDN, load balancer and ingress. Private probes observe management endpoints such as Prometheus and Alertmanager. Kubernetes APIs need not become internet-facing solely for monitoring.
A private VPN adds a dependency: its failure can make every private probe fail. Record that interpretation. Inventory the path between observer and target, including shared gateways and resolvers.
| Observation | Meaning of success |
|---|---|
| Process liveness | The process responds |
Prometheus / Alertmanager /-/ready | That component reports readiness |
| Public read-only synthetic | The specified path and response contract work |
| External Watchdog receipt | A heartbeat arrives through a particular alert route |
Readiness does not validate notification credentials or rule content. A health endpoint returning 200 does not establish every business function. Keep synthetic health separate from Kubernetes liveness so dependency failure does not unnecessarily trigger widespread restarts. [9][10]
4. Make blackbox probes representative of user access
Blackbox Exporter supports HTTP, DNS, TCP, ICMP, gRPC and other probes, exposing results through probe_success. Prometheus scrapes the exporter’s /probe endpoint rather than the target directly. [2]
The illustrative public HTTP module is:
modules:
http_public_ipv4:
prober: http
timeout: 5s
http:
method: GET
valid_status_codes: [200]
follow_redirects: false
fail_if_not_ssl: true
preferred_ip_protocol: ip4
ip_protocol_fallback: false
fail_if_body_not_matches_regexp: ['^ready\n?$']
Its contract is IPv4 only, no redirects, HTTPS, status 200 and a body containing only ready with an optional newline. The dedicated endpoint itself is not implemented. IPv4 success leaves IPv6 unverified; add a separate module and series where necessary. A login redirect or cached CDN response must not be mistaken for origin health. [3]
The corresponding scrape fragment passes the target as __param_target and keeps target identity separate from the exporter address: [4]
scrape_configs:
- job_name: external-blackbox
scrape_interval: 15s
scrape_timeout: 10s
metrics_path: /probe
params:
module: [http_public_ipv4]
static_configs:
- targets: [https://app.example.com/synthetic/ready]
labels:
target_role: public-application
relabel_configs:
- source_labels: [__address__]
target_label: __param_target
- source_labels: [__param_target]
target_label: instance
- target_label: __address__
replacement: 127.0.0.1:9115
DNS, connection, TLS and HTTP timings provide clues rather than a conclusive root cause. An HTTP probe using a hostname observes its resolver and cache path; it does not establish authoritative DNS correctness. For that requirement, design a separate DNS probe with an explicit resolver and question.
The exporter’s /probe?target=... can initiate requests to supplied destinations. The example assumes loopback binding, avoiding a publicly usable probe proxy. Private-module certificates and keys must be supplied separately at the documented paths; no secret values are included.
5. Separate failed probes, failed scrapes and missing series
probe_success == 0 means a probe result was obtained and reported failure. If the exporter stops and no result is returned, that expression alone does not alert.
Scroll horizontally to view the complete diagram.
Read the diagram as text
A successful scrape can report probe success or failure. up=0 means scrape failure; up=1 without a result means missing output; absent expected up requires inventory review. Evaluator failure needs external heartbeat detection.
| State | Observation | Investigation starting point |
|---|---|---|
up=1, probe_success=1 | Successful probe result | Contract met from this path at this time |
up=1, probe_success=0 | Failed target probe | Target or intervening path |
up=0 | Failed scrape | Exporter, scrape path, timeout or configuration |
up=1, no probe series | Missing result | Module or exporter-output contract |
Expected up series absent | Missing target or job | Discovery, removal or missing evaluation input |
The supplied rules distinguish these states. A subset follows:
groups:
- name: external-observation
rules:
- alert: EndpointProbeFailed
expr: probe_success{job=~"external-blackbox|management-blackbox"} == 0
for: 2m
labels:
severity: critical
annotations:
summary: 'Probe contract failed from the external sentinel.'
- alert: ProbeScrapeFailed
expr: up{job=~"external-blackbox|management-blackbox"} == 0
for: 1m
labels:
severity: critical
annotations:
summary: 'Probe result could not be scraped; target state is unknown.'
- alert: ProbeResultMissing
expr: |
(up{job=~"external-blackbox|management-blackbox"} == 1)
unless on (job, instance)
probe_success{job=~"external-blackbox|management-blackbox"}
for: 1m
labels:
severity: critical
annotations:
summary: 'Scrape succeeded but the expected probe result is absent.'
The full rules also contain absent(up{...}) for each known target. Job-wide absence alone misses a single removed target. Deleting an expected target and its absence rule together defeats the check, so inventory changes themselves require review. [6]
for: 2m measures how long an expression remains active at evaluation times; it is not a guaranteed counter of eight consecutive failures. Missing series and interrupted evaluation affect it. If Prometheus itself stops, these rules stop running too, requiring an external dead-man mechanism. [5]
6. Separate external alert delivery from production
Sending external probe alerts only to an Alertmanager inside production loses notification during the failure being observed. Run Alertmanager with the sentinel and use a receiver path that bypasses production.
Independence extends beyond process placement. Follow DNS, proxies, NAT, credentials, paging provider, on-call configuration and the authentication or devices used by responders. Email and SMS through one provider may still share its outage.
Choose which notifications need multiple paths. Duplicating every alert indiscriminately adds noise. Escalating monitoring or site failures through another provider can be a more focused design.
For HA Alertmanager, review direct delivery to each instance, duplicates during partitions and how replica labels affect deduplication. The supplied example describes one external VM; it does not provision an HA cluster. [1]
7. Use Watchdog to detect silence on a particular alert route
Watchdog remains firing and sends repeated notifications to a receiver. An external system detects their absence. The Prometheus Operator runbook explicitly requires external notification when this alerting system stops working. [7]
groups:
- name: internal-watchdog
rules:
- alert: Watchdog
expr: vector(1)
labels:
severity: none
monitoring_system: production-cluster
annotations:
summary: 'Production alert-route heartbeat.'
Add a dedicated route and receiver to internal Alertmanager without replacing ordinary paging routes:
# Merge into internal Alertmanager; preserve its normal receivers/routes.
route:
routes:
- matchers:
- alertname="Watchdog"
- monitoring_system="production-cluster"
receiver: production-heartbeat
group_by: [alertname, monitoring_system]
group_wait: 0s
group_interval: 1m
repeat_interval: 1m
receivers:
- name: production-heartbeat
webhook_configs:
- url_file: /run/secrets/production-heartbeat-url
send_resolved: false
http_config:
authorization:
credentials_file: /run/secrets/production-heartbeat-token
This fragment must be merged into configuration. Runtime secret files supply an HTTPS destination and credential. The receiver contract accepts Alertmanager webhook payloads; arbitrary heartbeat services are not necessarily compatible. Any required adapter is not implemented in the bundle. [8]
Use send_resolved: false, and require the receiver to validate firing status, alert name, stream identity and authentication. Resolved alerts, unrelated alerts and arbitrary POST requests must not count as healthy heartbeats. Decide whether silencing or inhibiting Watchdog creates a missing-heartbeat incident or an explicitly time-limited maintenance interval.
One logical Watchdog in an HA deployment establishes that at least one route is alive. Separate component checks are needed to inspect every Prometheus and Alertmanager replica.
8. Do not turn Watchdog into proof of all notification health
Watchdog covers its rule and route to a particular receiver. It does not validate every business rule, another receiver’s credentials, paging delivery or human awareness.
Heartbeat credentials may work while ordinary paging credentials expire. A dedicated Watchdog cannot be claimed to detect that fault. Separately exercise a synthetic alert through the actual critical route and record human acknowledgement.
Scroll horizontally to view the complete diagram.
Read the diagram as text
Watchdog observes its rule, dedicated route and receiver arrival. Ordinary paging credentials and human acknowledgement need a separate synthetic exercise. Receipt time does not guarantee fresh rule evaluation.
Heartbeat delivery also need not stop immediately when Prometheus stops. Alertmanager’s handling of existing firing alerts, retries and delayed delivery can postpone detection. Distinguish an alert’s validity from receiver arrival time. resolve_timeout alone does not uniformly determine expiry for Prometheus-originated alerts. [8]
Ordinary webhook arrival cannot always distinguish a fresh rule evaluation from delayed delivery or replay. startsAt is the start of a continuously firing alert, not the generation time of every repeat. Rejecting identical payloads can also reject legitimate repeats. Strong freshness requirements need an additional timestamp or sequence protocol, with its implementation and validation tracked separately.
9. Estimate detection time across the whole chain
A 15-second probe or one-minute heartbeat does not establish a human-notification deadline. Endpoint detection includes:
Probe scheduling + probe timeout + rule evaluation alignment
+ alert for duration + Alertmanager grouping
+ notification delivery + human acknowledgement
Heartbeat detection adds the delay between upstream silence and the last delivered heartbeat, receiver grace, missing-heartbeat evaluation, delivery and escalation. repeat_interval interacts with group_interval; align them on the dedicated route. The example’s one-minute repeats and proposed three-minute grace are not measured SLOs. [8]
Monitor streams that never produce their first heartbeat. A receiver armed only after first arrival misses deployments broken from the start. Manage expected-stream registration, startup grace, maintenance expiry and authorized deletion separately.
For an SLO, define whether timing ends at condition detection, rule firing, receiver receipt, paging delivery or acknowledgement. This design alone cannot guarantee an unconditional detection deadline while the destination itself is unavailable.
10. Read combinations as hypotheses
Public endpoints, private management endpoints and Watchdog provide complementary evidence, not a definitive cause lookup table.
| Watchdog receipt | Public probe | Investigation starting point |
|---|---|---|
| Continuing | Success | These two observation contracts are currently met |
| Missing | Success | Heartbeat route, receiver or internal monitoring |
| Continuing | Failure | Application, edge, DNS or observer-side path |
| Missing | Failure | Common failure; investigate through independent paths |
| Unknown | Missing data | Establish observer health first |
Failures from several regions can still share a resolver or faulty configuration. Ignoring one failed region can hide impact to users there.
A two-of-three policy requires expected observer count, series freshness, treatment of absence and availability of the aggregation system. Data from separate Prometheus instances does not automatically join one query. Hosting the aggregator inside production reintroduces a common dependency. Missing data must not silently become a healthy vote.
11. Make failure and recovery exercises acceptance obligations
Plan separate exercises in an isolated environment or approved maintenance window. None was executed for this article.
| Fault | Expected observation | Evidence still required |
|---|---|---|
| Synthetic endpoint unavailable | Probe-failure alert | Pager receipt and responder confirmation |
| Blackbox Exporter stopped | Scrape-failure alert | No incorrect target-failure classification |
| Target removed from configuration | Expected-series absence alert | Inventory expectation was not removed too |
| Internal Prometheus stopped | Readiness failure, then missing heartbeat | Measured delay from last repeat to notification |
| Heartbeat credential invalidated | Internal firing continues; external heartbeat missing | Receiver does not count invalid traffic as health |
| Ordinary pager credential invalidated | Critical-route synthetic fails to arrive | Failure is noticed even while Watchdog continues |
| Sentinel stopped | Sentinel’s own stream becomes missing | Receiver does not depend on production |
Distinguish stopping one replica from stopping all replicas. Watchdog continuing after one replica stops may be the intended HA behavior. Add cluster, DNS, VPN, site and provider faults with explicit scope and recovery procedures.
During recovery, retain external observations and verify the return of API, ingress, applications, internal monitoring and notifications. Readiness recovery and actual rule/notification recovery are separate results.
Scroll horizontally to view the complete diagram.
Read the diagram as text
For each fault, record shared dependencies and observation/notification ownership. Exercise isolated failures, measure last-good observation through delivery and acknowledgement, then verify both readiness and notification recovery.
12. Model loss of observation as a failure mode
For each monitoring component, record who observes it and who delivers its alert. The bundle separates failure-domain inventory, expected targets and heartbeat streams, receiver obligations, unexecuted exercises and a blank evidence record.
Owners, notification providers, VM placement, credential supply and SLOs remain unassigned where no adoption has occurred. An architecture diagram is not evidence of independence in the deployed environment.
Internal observability provides diagnostic detail. The external sentinel checks whether specified paths respond and whether observation and notification routes have fallen silent. Both are useful, and their claims differ.
Place observation and notification outside the common failures they must detect, then exercise that independence. Being able to answer “Who notices this Prometheus stopping?” belongs in the completion criteria for monitoring architecture.
The related observability-platform modernization case supplies design and responsibility context. This sentinel, configuration and unexecuted drill plan are not claimed as adopted or validated evidence from that engagement.
References
Official documentation source revisions are pinned below. These are editorial baselines, not installed binaries or validated deployment versions.