Skip to content
CoRISE

Who Notices Monitoring Failure? Observation Outside Kubernetes

Combine external probes, Watchdog heartbeats and independent notification, with explicit failure domains, missing-data handling, detection timing and recovery evidence.

T. AsanoPublished Updated 17 min read
  • Observability
  • Reliability
  • Operations
Open table of contents

Who observes this Prometheus when it stops?

Running Prometheus, Alertmanager and Grafana on Kubernetes with multiple replicas, anti-affinity and disruption budgets can improve resilience to individual pod or node failures.

But when monitoring shares networking, power, DNS or cloud-account dependencies with its targets, a common failure can remove observation and notification together. Preserving replicas through HA and preserving an observation point outside a failure domain are separate design problems.

A control-plane outage does not necessarily stop every existing pod immediately. Conversely, healthy control-plane components do not guarantee user access through ingress or public DNS. Replace the single label “cluster down” with explicit unavailable functions and observation paths.

Following From Monitoring to Observability, this article examines shared failures of the observation platform itself.

The configuration and review bundle describes an imaginary production cluster and external VM. Prometheus, Blackbox Exporter and Alertmanager configuration validation, probing, notifications, heartbeat reception and failure injection were not executed. Everything remains NOT_RUN. The bundle neither provisions a complete system nor reports measured detection performance.

1. Separate HA from failure-domain independence

Prometheus instances on separate nodes can survive some node failures. Alertmanager HA shares silences and notification state and favors duplicate notifications over missed notifications during a partition. Prometheus should send alerts to every Alertmanager instance rather than choose one through a load balancer. [1]

That does not guarantee delivery when all outbound connectivity is lost. PodDisruptionBudgets primarily govern voluntary disruption; they do not prevent site-wide power failure.

Common failure to detectSeparation to review
Pod or node lossPlacement, host, volume and node networking
Cluster infrastructure failureAnother cluster or VM with independent startup and management
VPC or site connectivity lossNetwork, upstream, VPN gateway and site
Account or administrative failureAccount, IAM, authentication, billing and operational authority
Notification-provider failureProvider, credentials, delivery and escalation

Another VM or provider is not automatically independent. DNS, identity, secret retrieval and deployment automation may still be shared. Work backward from the faults to detect. Different placement alone also does not justify treating failure probabilities as independent and multiplying them.

2. Keep internal diagnosis and add external observation

Retain monitoring inside the cluster. Service discovery and internal metrics help explain pod, storage, application and dependency behavior. Add a small external sentinel alongside that diagnostic system.

Scroll horizontally to view the complete diagram.

Separate diagnosis, endpoint observation and heartbeat paths Production provides diagnosis; an external sentinel probes service and monitoring endpoints and uses an independent pager. Each monitoring system sends a distinct heartbeat to an external receiver. Placement alone does not prove independence.
Fig. 01 — Separate diagnosis, endpoint observation and heartbeat paths
Read the diagram as text

Production provides diagnosis; an external sentinel probes service and monitoring endpoints and uses an independent pager. Each monitoring system sends a distinct heartbeat to an external receiver. Placement alone does not prove independence.

The sentinel checks important endpoints, internal monitoring readiness and a notification route that bypasses production. It need not duplicate Loki, Thanos and Grafana. A VM running Prometheus with short retention, Blackbox Exporter and Alertmanager is one candidate.

A single sentinel remains a single point of failure. Send its own heartbeat to an external receiver using a stream distinct from production’s. Record the failure domain delegated to that receiver and its notification path. A managed service does not eliminate residual provider or human-delivery risk.

Sharing an IaC repository differs from sharing runtime dependencies. If rebuilding through NixOS or OpenTofu, verify access to source, artifacts and credentials while production is unavailable.

3. Define what each probe path establishes

Public probes follow a user-facing route through DNS, TLS, CDN, load balancer and ingress. Private probes observe management endpoints such as Prometheus and Alertmanager. Kubernetes APIs need not become internet-facing solely for monitoring.

A private VPN adds a dependency: its failure can make every private probe fail. Record that interpretation. Inventory the path between observer and target, including shared gateways and resolvers.

ObservationMeaning of success
Process livenessThe process responds
Prometheus / Alertmanager /-/readyThat component reports readiness
Public read-only syntheticThe specified path and response contract work
External Watchdog receiptA heartbeat arrives through a particular alert route

Readiness does not validate notification credentials or rule content. A health endpoint returning 200 does not establish every business function. Keep synthetic health separate from Kubernetes liveness so dependency failure does not unnecessarily trigger widespread restarts. [9][10]

4. Make blackbox probes representative of user access

Blackbox Exporter supports HTTP, DNS, TCP, ICMP, gRPC and other probes, exposing results through probe_success. Prometheus scrapes the exporter’s /probe endpoint rather than the target directly. [2]

The illustrative public HTTP module is:

modules:
  http_public_ipv4:
    prober: http
    timeout: 5s
    http:
      method: GET
      valid_status_codes: [200]
      follow_redirects: false
      fail_if_not_ssl: true
      preferred_ip_protocol: ip4
      ip_protocol_fallback: false
      fail_if_body_not_matches_regexp: ['^ready\n?$']

Its contract is IPv4 only, no redirects, HTTPS, status 200 and a body containing only ready with an optional newline. The dedicated endpoint itself is not implemented. IPv4 success leaves IPv6 unverified; add a separate module and series where necessary. A login redirect or cached CDN response must not be mistaken for origin health. [3]

The corresponding scrape fragment passes the target as __param_target and keeps target identity separate from the exporter address: [4]

scrape_configs:
  - job_name: external-blackbox
    scrape_interval: 15s
    scrape_timeout: 10s
    metrics_path: /probe
    params:
      module: [http_public_ipv4]
    static_configs:
      - targets: [https://app.example.com/synthetic/ready]
        labels:
          target_role: public-application
    relabel_configs:
      - source_labels: [__address__]
        target_label: __param_target
      - source_labels: [__param_target]
        target_label: instance
      - target_label: __address__
        replacement: 127.0.0.1:9115

DNS, connection, TLS and HTTP timings provide clues rather than a conclusive root cause. An HTTP probe using a hostname observes its resolver and cache path; it does not establish authoritative DNS correctness. For that requirement, design a separate DNS probe with an explicit resolver and question.

The exporter’s /probe?target=... can initiate requests to supplied destinations. The example assumes loopback binding, avoiding a publicly usable probe proxy. Private-module certificates and keys must be supplied separately at the documented paths; no secret values are included.

5. Separate failed probes, failed scrapes and missing series

probe_success == 0 means a probe result was obtained and reported failure. If the exporter stops and no result is returned, that expression alone does not alert.

Scroll horizontally to view the complete diagram.

Distinguish target failure from observation failure A successful scrape can report probe success or failure. up=0 means scrape failure; up=1 without a result means missing output; absent expected up requires inventory review. Evaluator failure needs external heartbeat detection.
Fig. 02 — Distinguish target failure from observation failure
Read the diagram as text

A successful scrape can report probe success or failure. up=0 means scrape failure; up=1 without a result means missing output; absent expected up requires inventory review. Evaluator failure needs external heartbeat detection.

StateObservationInvestigation starting point
up=1, probe_success=1Successful probe resultContract met from this path at this time
up=1, probe_success=0Failed target probeTarget or intervening path
up=0Failed scrapeExporter, scrape path, timeout or configuration
up=1, no probe seriesMissing resultModule or exporter-output contract
Expected up series absentMissing target or jobDiscovery, removal or missing evaluation input

The supplied rules distinguish these states. A subset follows:

groups:
  - name: external-observation
    rules:
      - alert: EndpointProbeFailed
        expr: probe_success{job=~"external-blackbox|management-blackbox"} == 0
        for: 2m
        labels:
          severity: critical
        annotations:
          summary: 'Probe contract failed from the external sentinel.'
      - alert: ProbeScrapeFailed
        expr: up{job=~"external-blackbox|management-blackbox"} == 0
        for: 1m
        labels:
          severity: critical
        annotations:
          summary: 'Probe result could not be scraped; target state is unknown.'
      - alert: ProbeResultMissing
        expr: |
          (up{job=~"external-blackbox|management-blackbox"} == 1)
          unless on (job, instance)
          probe_success{job=~"external-blackbox|management-blackbox"}
        for: 1m
        labels:
          severity: critical
        annotations:
          summary: 'Scrape succeeded but the expected probe result is absent.'

The full rules also contain absent(up{...}) for each known target. Job-wide absence alone misses a single removed target. Deleting an expected target and its absence rule together defeats the check, so inventory changes themselves require review. [6]

for: 2m measures how long an expression remains active at evaluation times; it is not a guaranteed counter of eight consecutive failures. Missing series and interrupted evaluation affect it. If Prometheus itself stops, these rules stop running too, requiring an external dead-man mechanism. [5]

6. Separate external alert delivery from production

Sending external probe alerts only to an Alertmanager inside production loses notification during the failure being observed. Run Alertmanager with the sentinel and use a receiver path that bypasses production.

Independence extends beyond process placement. Follow DNS, proxies, NAT, credentials, paging provider, on-call configuration and the authentication or devices used by responders. Email and SMS through one provider may still share its outage.

Choose which notifications need multiple paths. Duplicating every alert indiscriminately adds noise. Escalating monitoring or site failures through another provider can be a more focused design.

For HA Alertmanager, review direct delivery to each instance, duplicates during partitions and how replica labels affect deduplication. The supplied example describes one external VM; it does not provision an HA cluster. [1]

7. Use Watchdog to detect silence on a particular alert route

Watchdog remains firing and sends repeated notifications to a receiver. An external system detects their absence. The Prometheus Operator runbook explicitly requires external notification when this alerting system stops working. [7]

groups:
  - name: internal-watchdog
    rules:
      - alert: Watchdog
        expr: vector(1)
        labels:
          severity: none
          monitoring_system: production-cluster
        annotations:
          summary: 'Production alert-route heartbeat.'

Add a dedicated route and receiver to internal Alertmanager without replacing ordinary paging routes:

# Merge into internal Alertmanager; preserve its normal receivers/routes.
route:
  routes:
    - matchers:
        - alertname="Watchdog"
        - monitoring_system="production-cluster"
      receiver: production-heartbeat
      group_by: [alertname, monitoring_system]
      group_wait: 0s
      group_interval: 1m
      repeat_interval: 1m
receivers:
  - name: production-heartbeat
    webhook_configs:
      - url_file: /run/secrets/production-heartbeat-url
        send_resolved: false
        http_config:
          authorization:
            credentials_file: /run/secrets/production-heartbeat-token

This fragment must be merged into configuration. Runtime secret files supply an HTTPS destination and credential. The receiver contract accepts Alertmanager webhook payloads; arbitrary heartbeat services are not necessarily compatible. Any required adapter is not implemented in the bundle. [8]

Use send_resolved: false, and require the receiver to validate firing status, alert name, stream identity and authentication. Resolved alerts, unrelated alerts and arbitrary POST requests must not count as healthy heartbeats. Decide whether silencing or inhibiting Watchdog creates a missing-heartbeat incident or an explicitly time-limited maintenance interval.

One logical Watchdog in an HA deployment establishes that at least one route is alive. Separate component checks are needed to inspect every Prometheus and Alertmanager replica.

8. Do not turn Watchdog into proof of all notification health

Watchdog covers its rule and route to a particular receiver. It does not validate every business rule, another receiver’s credentials, paging delivery or human awareness.

Heartbeat credentials may work while ordinary paging credentials expire. A dedicated Watchdog cannot be claimed to detect that fault. Separately exercise a synthetic alert through the actual critical route and record human acknowledgement.

Scroll horizontally to view the complete diagram.

Dedicated heartbeat and ordinary paging are different contracts Watchdog observes its rule, dedicated route and receiver arrival. Ordinary paging credentials and human acknowledgement need a separate synthetic exercise. Receipt time does not guarantee fresh rule evaluation.
Fig. 03 — Dedicated heartbeat and ordinary paging are different contracts
Read the diagram as text

Watchdog observes its rule, dedicated route and receiver arrival. Ordinary paging credentials and human acknowledgement need a separate synthetic exercise. Receipt time does not guarantee fresh rule evaluation.

Heartbeat delivery also need not stop immediately when Prometheus stops. Alertmanager’s handling of existing firing alerts, retries and delayed delivery can postpone detection. Distinguish an alert’s validity from receiver arrival time. resolve_timeout alone does not uniformly determine expiry for Prometheus-originated alerts. [8]

Ordinary webhook arrival cannot always distinguish a fresh rule evaluation from delayed delivery or replay. startsAt is the start of a continuously firing alert, not the generation time of every repeat. Rejecting identical payloads can also reject legitimate repeats. Strong freshness requirements need an additional timestamp or sequence protocol, with its implementation and validation tracked separately.

9. Estimate detection time across the whole chain

A 15-second probe or one-minute heartbeat does not establish a human-notification deadline. Endpoint detection includes:

Probe scheduling + probe timeout + rule evaluation alignment
+ alert for duration + Alertmanager grouping
+ notification delivery + human acknowledgement

Heartbeat detection adds the delay between upstream silence and the last delivered heartbeat, receiver grace, missing-heartbeat evaluation, delivery and escalation. repeat_interval interacts with group_interval; align them on the dedicated route. The example’s one-minute repeats and proposed three-minute grace are not measured SLOs. [8]

Monitor streams that never produce their first heartbeat. A receiver armed only after first arrival misses deployments broken from the start. Manage expected-stream registration, startup grace, maintenance expiry and authorized deletion separately.

For an SLO, define whether timing ends at condition detection, rule firing, receiver receipt, paging delivery or acknowledgement. This design alone cannot guarantee an unconditional detection deadline while the destination itself is unavailable.

10. Read combinations as hypotheses

Public endpoints, private management endpoints and Watchdog provide complementary evidence, not a definitive cause lookup table.

Watchdog receiptPublic probeInvestigation starting point
ContinuingSuccessThese two observation contracts are currently met
MissingSuccessHeartbeat route, receiver or internal monitoring
ContinuingFailureApplication, edge, DNS or observer-side path
MissingFailureCommon failure; investigate through independent paths
UnknownMissing dataEstablish observer health first

Failures from several regions can still share a resolver or faulty configuration. Ignoring one failed region can hide impact to users there.

A two-of-three policy requires expected observer count, series freshness, treatment of absence and availability of the aggregation system. Data from separate Prometheus instances does not automatically join one query. Hosting the aggregator inside production reintroduces a common dependency. Missing data must not silently become a healthy vote.

11. Make failure and recovery exercises acceptance obligations

Plan separate exercises in an isolated environment or approved maintenance window. None was executed for this article.

FaultExpected observationEvidence still required
Synthetic endpoint unavailableProbe-failure alertPager receipt and responder confirmation
Blackbox Exporter stoppedScrape-failure alertNo incorrect target-failure classification
Target removed from configurationExpected-series absence alertInventory expectation was not removed too
Internal Prometheus stoppedReadiness failure, then missing heartbeatMeasured delay from last repeat to notification
Heartbeat credential invalidatedInternal firing continues; external heartbeat missingReceiver does not count invalid traffic as health
Ordinary pager credential invalidatedCritical-route synthetic fails to arriveFailure is noticed even while Watchdog continues
Sentinel stoppedSentinel’s own stream becomes missingReceiver does not depend on production

Distinguish stopping one replica from stopping all replicas. Watchdog continuing after one replica stops may be the intended HA behavior. Add cluster, DNS, VPN, site and provider faults with explicit scope and recovery procedures.

During recovery, retain external observations and verify the return of API, ingress, applications, internal monitoring and notifications. Readiness recovery and actual rule/notification recovery are separate results.

Scroll horizontally to view the complete diagram.

Review independence, then exercise failure and recovery For each fault, record shared dependencies and observation/notification ownership. Exercise isolated failures, measure last-good observation through delivery and acknowledgement, then verify both readiness and notification recovery.
Fig. 04 — Review independence, then exercise failure and recovery
Read the diagram as text

For each fault, record shared dependencies and observation/notification ownership. Exercise isolated failures, measure last-good observation through delivery and acknowledgement, then verify both readiness and notification recovery.

12. Model loss of observation as a failure mode

For each monitoring component, record who observes it and who delivers its alert. The bundle separates failure-domain inventory, expected targets and heartbeat streams, receiver obligations, unexecuted exercises and a blank evidence record.

Owners, notification providers, VM placement, credential supply and SLOs remain unassigned where no adoption has occurred. An architecture diagram is not evidence of independence in the deployed environment.

Internal observability provides diagnostic detail. The external sentinel checks whether specified paths respond and whether observation and notification routes have fallen silent. Both are useful, and their claims differ.

Place observation and notification outside the common failures they must detect, then exercise that independence. Being able to answer “Who notices this Prometheus stopping?” belongs in the completion criteria for monitoring architecture.

The related observability-platform modernization case supplies design and responsibility context. This sentinel, configuration and unexecuted drill plan are not claimed as adopted or validated evidence from that engagement.

References

Official documentation source revisions are pinned below. These are editorial baselines, not installed binaries or validated deployment versions.

  1. Alertmanager High Availability
  2. Blackbox Exporter
  3. Blackbox Exporter Configuration
  4. Prometheus Configuration
  5. Prometheus Alerting Rules
  6. PromQL Functions / absent
  7. Prometheus Operator: Watchdog Runbook
  8. Alertmanager Configuration
  9. Prometheus Management API
  10. Alertmanager Management API

T. Asano

More articles by T. Asano

Contact

Tell us about your engineering challenge.

Talk with CoRISE about the design, implementation and operation of your systems.

Start a Conversation