From Legacy Monitoring to an Observability Platform
From individually configured monitoring to a platform for observing the whole system. We progressively modernized a long-running monitoring environment into a Kubernetes-based observability platform.
The previous environment combined Munin, Nagios, Cacti and syslog-ng, with checks and alert conditions configured for individual hosts and services. CoRISE mapped their existing responsibilities and migrated progressively toward Prometheus, Thanos, Grafana, Vector, Loki, OpenTelemetry and Alertmanager.
- Kubernetes
- Prometheus
- Thanos
- Grafana
- Vector
- Loki
- OpenTelemetry
- Alertmanager
- Nix / NixOS
01 / MIGRATION
Progressive migration from legacy monitoring
Mapped existing tool responsibilities and operational constraints before migrating in stages.
02 / PLATFORM
Declarative observability on Kubernetes
Organized telemetry collection and visualization as a shared platform.
03 / OPERATIONS
Continued platform lifecycle management
Managed component versions with Nix and monitored the observability platform itself.
Background
From host availability to understanding service behavior.
The previous platform monitored servers, processes, networks and individual metrics, notifying operators when predefined conditions were met. Containers and Kubernetes introduced an environment in which monitoring targets change continuously.
We moved toward systems exposing telemetry for collection and evaluation by a shared platform, extending visibility from individual hosts to services and their users.
- Overall service health
- Components showing abnormal behavior
- When symptoms first appeared
- Relationships between metrics, logs and traces
- The effect of incidents on users
Before
Host-centered monitoring with individual configuration.
Munin, Nagios, Cacti and syslog-ng each served important purposes. In this environment, configuration and data were managed separately, creating challenges as the system grew and changed.
↔ Scroll horizontally to see the full diagram.
View diagram as text
Hosts branch to Munin for metrics, Cacti for metrics and graphs, Nagios for checks and alerts, and syslog-ng for logs. This describes the separately managed paths in this environment, not a limitation on the tools’ integration capabilities.
- Manual additions to monitoring targets and configuration
- Separated metrics and logs
- Limited service-level context from host checks
- Difficult long-term retention and cross-system analysis
- Operational effort to reconfigure and upgrade the monitoring platform
After
A declarative observability platform on Kubernetes.
The new platform organizes components around their telemetry responsibilities.
METRICS
Prometheus
Collects metrics from services, infrastructure and Kubernetes, with PromQL for evaluation. Kubernetes service discovery helps the configuration follow changes in the environment.
LONG-TERM METRICS
Thanos
Provides long-term retention and queries across multiple Prometheus environments, supporting analysis beyond immediate anomaly detection.
- Long-term trends
- Capacity planning
- Past incident comparisons
- Seasonality
- Resource usage changes
VISUALIZATION
Grafana
Provides a shared interface for exploring metrics and logs, understanding system state and moving into investigation.
LOGS
Vector + Loki
Vector collects and transforms container and system logs before sending them to Loki. Parsing, filtering, enrichment and routing accommodate differences between applications and environments.
Loki provides log search integrated with Grafana. The operating model aims to connect metric anomalies to logs from the same time window.
TELEMETRY STANDARDIZATION
OpenTelemetry
Provides a standard telemetry entry point for newer applications and distributed systems, separating metrics, logs and traces instrumentation from backend products.
ALERTING
Alerting Rules + Alertmanager
Prometheus evaluates alert rules; Alertmanager handles grouping, routing and silences. Declarative rules encode both abnormal conditions and notification policy.
- Which conditions are abnormal
- How long a condition must persist
- How related alerts are grouped
- Who receives notifications
Architecture
Design telemetry paths as one connected system.
↔ Scroll horizontally to see the full diagram.
View diagram as text
Prometheus collects metrics; Thanos supports long-term retention and cross-environment queries. Vector sends logs to Loki. Grafana queries metrics and logs. OpenTelemetry instrumentation sends telemetry to a collector for export to backends; the collector is not a trace storage or query backend. Prometheus evaluates alert rules and sends alerts through Alertmanager to operators, independently of Grafana. This is a conceptual component and data-flow view.
Declarative Operations
Make monitoring configuration maintainable software.
Monitoring targets, alert rules, dashboards and Kubernetes resources are managed as YAML and configuration files with a Git history. The platform is designed to be reproducible, changeable and reviewable.
- Trace who changed what
- Review changes
- Reproduce configuration
- Automate updates
- Make rollback easier
↔ Scroll horizontally to see the full diagram.
View diagram as text
Configuration for targets, rules and dashboards is managed in Git, reviewed and changed, then applied as declarative configuration to Kubernetes.
From Monitoring to Observability
Connect detection to the information needed for diagnosis.
Detection, context, investigation, diagnosis and improvement form a connected operating experience. The following sequence illustrates incident investigation.
↔ Scroll horizontally to see the full diagram.
View diagram as text
Detection leads to context, investigation, diagnosis and improvement. A return arrow from improvement to detection represents findings informing monitoring rules.
For example, an incident investigation follows this sequence.
- 01Detect an anomaly in Prometheus metrics
- 02Inspect related metrics in Grafana
- 03Examine Loki logs for the same period
- 04Consult traces when needed
- 05Inspect Kubernetes state
- 06Correct the cause
- 07Feed findings into new metrics and alerts
Findings feed back into monitoring rules and dashboards, allowing the platform to improve through operational experience.
Monitoring the Monitoring System
Operate the monitoring platform as a production system.
If the observability platform stops working, information needed during an incident may become unavailable. Its own health is therefore part of the monitoring scope.
The monitoring system is treated as production infrastructure.
- Prometheus
- Thanos
- Grafana
- Loki
- Vector
- OTel Collector
- Alertmanager
- Kubernetes
- etcd
Kubernetes Lifecycle
Keep the platform updatable throughout its life.
The platform runs on self-hosted Kubernetes on NixOS. CoRISE contributed to lifecycle management of both monitoring applications and the underlying components.
Nix manages Kubernetes and related tool versions, reducing environment differences and supporting continued upgrades.
- Kubernetes Components
- etcd
- Middleware
- Container Images
- Configuration
- Dependencies
What Changed
Modernize the operating model alongside the tools.
Before
- Host-centered monitoring
- Static configuration
- Individually managed tools
- Separated metrics and logs
- Threshold-oriented notifications
- Manual configuration changes
- Short-term monitoring focus
After
- Service and platform observability
- Kubernetes-aware discovery
- Declarative configuration
- Design for correlating metrics, logs and traces
- Prometheus alert evaluation
- Git-managed operational configuration
- Long-term metrics with Thanos
- Centralized log analysis with Loki
- OpenTelemetry instrumentation
- Continuous platform lifecycle management
Outcome
An operating platform that can follow change.
The work aimed to keep monitoring aligned with system changes, connect notifications to investigation, feed operational knowledge into configuration and design, and make the platform itself continuously updatable.
The central objective was to make monitoring a changeable engineering system that improves through continued operation.
CoRISE’s Role
Our contribution to design, migration and continued operation.
- 01Progressive migration from legacy monitoring
- 02Prometheus metrics infrastructure
- 03Long-term metrics with Thanos
- 04Grafana visualization
- 05Vector log pipelines
- 06Loki log aggregation
- 07OpenTelemetry telemetry architecture
- 08Alerting rules and Alertmanager notifications
- 09Observability deployment on Kubernetes
- 10Declarative configuration management
- 11Component lifecycle management with Nix
- 12Continued Kubernetes and etcd operation
- 13Monitoring the observability platform itself
- 14Middleware upgrades
- 15Continuous operational improvement
Technologies
- Kubernetes
- Nix / NixOS
- Prometheus
- Thanos
- Grafana
- Vector
- Loki
- OpenTelemetry
- Alertmanager
- etcd
- Git
- YAML