Skip to content
CoRISE
CASE / 03Platform & Operations

From Legacy Monitoring to an Observability Platform

From individually configured monitoring to a platform for observing the whole system. We progressively modernized a long-running monitoring environment into a Kubernetes-based observability platform.

The previous environment combined Munin, Nagios, Cacti and syslog-ng, with checks and alert conditions configured for individual hosts and services. CoRISE mapped their existing responsibilities and migrated progressively toward Prometheus, Thanos, Grafana, Vector, Loki, OpenTelemetry and Alertmanager.

  • Kubernetes
  • Prometheus
  • Thanos
  • Grafana
  • Vector
  • Loki
  • OpenTelemetry
  • Alertmanager
  • Nix / NixOS

Progressive migration from legacy monitoring

Mapped existing tool responsibilities and operational constraints before migrating in stages.

Declarative observability on Kubernetes

Organized telemetry collection and visualization as a shared platform.

Continued platform lifecycle management

Managed component versions with Nix and monitored the observability platform itself.

From host availability to understanding service behavior.

The previous platform monitored servers, processes, networks and individual metrics, notifying operators when predefined conditions were met. Containers and Kubernetes introduced an environment in which monitoring targets change continuously.

We moved toward systems exposing telemetry for collection and evaluation by a shared platform, extending visibility from individual hosts to services and their users.

Host-centered monitoring with individual configuration.

Munin, Nagios, Cacti and syslog-ng each served important purposes. In this environment, configuration and data were managed separately, creating challenges as the system grew and changed.

↔ Scroll horizontally to see the full diagram.

Monitoring before migrationHosts branch to Munin for metrics, Cacti for metrics and graphs, Nagios for checks and alerts, and syslog-ng for logs. This describes the separately managed paths in this environment, not a limitation on the tools’ integration capabilities.MetricsMetrics / GraphsChecks / AlertsLogsHOSTSMUNINCACTINAGIOSSYSLOG-NG
Fig. 01 — Separately managed monitoring paths in the previous environment.

A declarative observability platform on Kubernetes.

The new platform organizes components around their telemetry responsibilities.

Thanos

Provides long-term retention and queries across multiple Prometheus environments, supporting analysis beyond immediate anomaly detection.

  • Long-term trends
  • Capacity planning
  • Past incident comparisons
  • Seasonality
  • Resource usage changes

Grafana

Provides a shared interface for exploring metrics and logs, understanding system state and moving into investigation.

OpenTelemetry

Provides a standard telemetry entry point for newer applications and distributed systems, separating metrics, logs and traces instrumentation from backend products.

Alerting Rules + Alertmanager

Prometheus evaluates alert rules; Alertmanager handles grouping, routing and silences. Declarative rules encode both abnormal conditions and notification policy.

  • Which conditions are abnormal
  • How long a condition must persist
  • How related alerts are grouped
  • Who receives notifications

Design telemetry paths as one connected system.

↔ Scroll horizontally to see the full diagram.

Telemetry collection, querying and notificationPrometheus collects metrics; Thanos supports long-term retention and cross-environment queries. Vector sends logs to Loki. Grafana queries metrics and logs. OpenTelemetry instrumentation sends telemetry to a collector for export to backends; the collector is not a trace storage or query backend. Prometheus evaluates alert rules and sends alerts through Alertmanager to operators, independently of Grafana. This is a conceptual component and data-flow view.KUBERNETES — DECLARATIVE CONFIGURATIONTELEMETRY BACKENDSQuery metrics / logsEvaluated by PrometheusGrouping / Routing / Silence → NotificationAPPLICATIONSMETRICSLOGSTRACESPROMETHEUSVECTOROPENTELEMETRYTHANOSLOKIOTEL COLLECTORGRAFANAALERTING RULESALERTMANAGEROPERATIONS TEAM
Fig. 02 — Separate telemetry querying, export and alert notification paths.

Make monitoring configuration maintainable software.

Monitoring targets, alert rules, dashboards and Kubernetes resources are managed as YAML and configuration files with a Git history. The platform is designed to be reproducible, changeable and reviewable.

↔ Scroll horizontally to see the full diagram.

Declarative configuration changesConfiguration for targets, rules and dashboards is managed in Git, reviewed and changed, then applied as declarative configuration to Kubernetes.CONFIGURATIONGITREVIEW / CHANGEDECLARATIVECONFIGURATIONKUBERNETES
Fig. 03 — Connect configuration changes to history and review.

Connect detection to the information needed for diagnosis.

Detection, context, investigation, diagnosis and improvement form a connected operating experience. The following sequence illustrates incident investigation.

↔ Scroll horizontally to see the full diagram.

From detection to improvementDetection leads to context, investigation, diagnosis and improvement. A return arrow from improvement to detection represents findings informing monitoring rules.DETECTIONCONTEXTINVESTIGATIONDIAGNOSISIMPROVEMENT
Fig. 04 — Feed incident findings into subsequent detection and investigation.

For example, an incident investigation follows this sequence.

  1. 01Detect an anomaly in Prometheus metrics
  2. 02Inspect related metrics in Grafana
  3. 03Examine Loki logs for the same period
  4. 04Consult traces when needed
  5. 05Inspect Kubernetes state
  6. 06Correct the cause
  7. 07Feed findings into new metrics and alerts

Findings feed back into monitoring rules and dashboards, allowing the platform to improve through operational experience.

Operate the monitoring platform as a production system.

If the observability platform stops working, information needed during an incident may become unavailable. Its own health is therefore part of the monitoring scope.

The monitoring system is treated as production infrastructure.

  • Prometheus
  • Thanos
  • Grafana
  • Loki
  • Vector
  • OTel Collector
  • Alertmanager
  • Kubernetes
  • etcd

Keep the platform updatable throughout its life.

The platform runs on self-hosted Kubernetes on NixOS. CoRISE contributed to lifecycle management of both monitoring applications and the underlying components.

Nix manages Kubernetes and related tool versions, reducing environment differences and supporting continued upgrades.

  • Kubernetes Components
  • etcd
  • Middleware
  • Container Images
  • Configuration
  • Dependencies

Modernize the operating model alongside the tools.

  • Host-centered monitoring
  • Static configuration
  • Individually managed tools
  • Separated metrics and logs
  • Threshold-oriented notifications
  • Manual configuration changes
  • Short-term monitoring focus
  • Service and platform observability
  • Kubernetes-aware discovery
  • Declarative configuration
  • Design for correlating metrics, logs and traces
  • Prometheus alert evaluation
  • Git-managed operational configuration
  • Long-term metrics with Thanos
  • Centralized log analysis with Loki
  • OpenTelemetry instrumentation
  • Continuous platform lifecycle management

An operating platform that can follow change.

The work aimed to keep monitoring aligned with system changes, connect notifications to investigation, feed operational knowledge into configuration and design, and make the platform itself continuously updatable.

The central objective was to make monitoring a changeable engineering system that improves through continued operation.

Our contribution to design, migration and continued operation.

  1. 01Progressive migration from legacy monitoring
  2. 02Prometheus metrics infrastructure
  3. 03Long-term metrics with Thanos
  4. 04Grafana visualization
  5. 05Vector log pipelines
  6. 06Loki log aggregation
  7. 07OpenTelemetry telemetry architecture
  8. 08Alerting rules and Alertmanager notifications
  9. 09Observability deployment on Kubernetes
  10. 10Declarative configuration management
  11. 11Component lifecycle management with Nix
  12. 12Continued Kubernetes and etcd operation
  13. 13Monitoring the observability platform itself
  14. 14Middleware upgrades
  15. 15Continuous operational improvement
  • Kubernetes
  • Nix / NixOS
  • Prometheus
  • Thanos
  • Grafana
  • Vector
  • Loki
  • OpenTelemetry
  • Alertmanager
  • etcd
  • Git
  • YAML

Contact

Build an observability platform you can keep operating.

Talk to us about modernizing Nagios, Cacti or Munin environments, adopting Prometheus, Grafana and OpenTelemetry, or building and operating an observability platform on Kubernetes.

Start a conversation

Back to all work