Open table of contents
Conclusion
Connect collection upgrades to signals that explain failures and declarative change management.
Context
Moving from availability checks to explaining user symptoms requires signal correlation and managed configuration changes.
Design and verification scope
Assess the following responsibilities and boundaries when designing and verifying a configuration.
- Munin, Nagios and Cacti
- Prometheus and Thanos
- Vector and Loki
- OpenTelemetry
- Grafana
- Declarative operations
Decision rationale
Beyond adding stores for metrics, logs and traces, align identifiers and queries that connect user symptoms to application and infrastructure behavior. Compare Thanos retention, Vector log ingestion and OpenTelemetry instrumentation as distinct responsibilities.
Trade-offs
Longer retention and more labels provide context but increase storage and query costs. Account for information lost through sampling or aggregation, collection-system failures and the work needed to maintain declarative configuration.
Limitations
Review old and new topology, retention, labels, correlation procedures and operational diffs. Added tools alone do not establish better diagnosis.
Related case context
These cases provide attributed design context. They do not establish that the proposed experiments or configurations were delivered in those engagements.