Open table of contents
What should we check after traces arrive?
OpenTelemetry is installed. Traces reach the backend, and metrics and logs are visible. That establishes an important part of the collection path. The next question is whether those signals help an operator explain a failure.
Suppose checkout becomes slow for only some customers after a deployment. Which version, dependency and request characteristics are involved? What do the logs from the same execution show? Can an operator move beyond prepared dashboards and investigate questions that arise during the incident? OpenTelemetry’s observability primer emphasizes understanding internal state from external outputs, including problems not anticipated in advance. [1]
Building on From Monitoring to Observability and Vector and OpenTelemetry Collector, this article asks how to evaluate the quality of collected signals.
The example bundle contains a question inventory, signal contract, sampling fragment, acceptance plan and incident-drill plan for an imaginary checkout service. Collector validation/startup, SDK instrumentation, backend delivery, failure injection and timing measurements were not executed. All execution statuses remain NOT_RUN. No measured improvement is claimed.
1. Start with a question inventory
Starting with available spans and metrics can constrain questions to the data already collected. First identify what operators need to know to make decisions.
| Question | Required evidence | Investigation path |
|---|---|---|
| Q1: When did it start, and how much impact? | Request volume, latency distribution and final outcome | Service metrics and time window |
| Q2: Which version or cell? | Identity reported by running workloads | Version and cell comparisons |
| Q3: Which dependency consumed time? | Request and dependency-call timing | Trace relationships and critical path |
| Q4: What do affected requests share? | Provider, tier and feature variant | Approved attribute filters |
| Q5: Did retries recover? | Attempt and final-operation outcomes | Spans, events and structured logs |
| Q6: What happened in that execution? | Trace/span IDs and resource context | Trace-to-log and log-to-trace navigation |
| Q7: Why is data missing? | Sampling, delay, refusal and delivery failures | Independent pipeline observation |
Scroll horizontally to view the complete diagram.
Read the diagram as text
Questions about versions, dependencies and affected requests determine context and correlation. Inspect SDK, pipeline and backend behavior; return unanswered questions to a gap. Data presence alone does not close the work.
An attribute’s presence is insufficient to mark a question answerable. Record the environment, time window, query procedure and inspected evidence. Preserve unverified, unanswered and out-of-scope items. We cannot predict every future question, but common identities and comparison dimensions support new combinations of queries.
2. Align service identity and change context
Repeated unknown_service entries make source attribution difficult. Establish resource naming and configuration first. This example makes service.name, service.namespace, service.version and deployment.environment.name organizational requirements; it does not imply that OpenTelemetry universally requires all four. [2][3]
# Illustrative staging identity; SDK support for env configuration varies.
export OTEL_SERVICE_NAME=checkout-api
export OTEL_RESOURCE_ATTRIBUTES='service.namespace=commerce,service.version=1.0.0,deployment.environment.name=staging'
service.instance.id helps distinguish replicas. A pod UID can be useful where it identifies the intended service instance; inspect cases such as multiple services within one pod. Resource attributes need not all become metric labels. [2]
A version change coinciding with latency is evidence for a hypothesis, not proof of causation. The new version may run in another region, receive different tenants or traffic, or coincide with a feature-flag change. Compare those conditions before assigning cause.
Link deployment events, configuration revisions and feature variants when relevant. Check the version reported by the running process, rather than only the intended Git revision. Environment is an important query dimension but is not interchangeable with service identity; include namespace and environment in query scope explicitly.
3. Place spans at operational responsibility boundaries
Use automatic instrumentation for HTTP and database edges, then add necessary domain meaning. Every function need not become a span. Official library guidance recommends operations meaningful to the library’s users. [4]
POST /checkout
├─ inventory.reserve
├─ payment.authorize
│ ├─ payment.attempt [timeout]
│ └─ payment.attempt [success]
└─ db.create_order
Span names describe operation classes, without order or customer IDs. Aggregate HTTP paths using route templates such as /orders/{id} rather than unbounded raw URLs and query strings. Prefix custom domain vocabulary, for example corise.payment.provider, to distinguish it from standard conventions.
A 1.6-second payment span establishes time spent in that call interval. It does not identify remote CPU as the cause: connection-pool waits, DNS, networking and retries may contribute. Parallel span durations cannot simply be added. Inspect overlap and synchronization to understand the critical path.
| Operation | Error interpretation |
|---|---|
| One attempt times out | Record failure on that attempt |
| Authorization succeeds after retry | Do not promote the recovered error to the whole operation |
| Checkout ultimately fails | Record its final failed outcome |
| HTTP 404 | Apply client/server conventions and operation context |
Generic HTTP instrumentation treats client and server 4xx differently. Business rejection is another semantic layer. Recording an exception does not necessarily set span status; check the API. Successful operations generally need no explicit OK override and can remain UNSET. [5][6]
4. Propagate context and verify log navigation
Review queues, tasks and callbacks as well as HTTP. Inject/extract support can be present while asynchronous work loses the active context, producing unrelated spans or logs. [4]
A worker starting another trace is not inherently incorrect. Long jobs, batches and fan-in may be better represented through span links than a single parent. Review messaging conventions, instrumentation and backend link navigation together. Preserving causality is the requirement; using one trace ID everywhere is not. [7]
{
"event": "payment.attempt.failed",
"severity": "WARN",
"trace_id": "0123456789abcdef0123456789abcdef",
"span_id": "0123456789abcdef",
"error.type": "timeout",
"corise.payment.provider": "provider-a",
"corise.retry.attempt": 1
}
This is an illustrative structured-log representation, not OTLP wire format. Actual LogRecords need appropriate resource, timestamp and trace/span context. Configure the logger bridge and backend navigation in both directions. A trace ID in a log does not guarantee that its trace has been retained. [9]
For Rust tracing, check span integration and log export/correlation separately. Adding tracing-opentelemetry does not necessarily export every log as an OTel LogRecord. Record compatible SDK, bridge, runtime and exporter versions and verify their configuration. [14]
5. Connect metrics to traces, with explicit limits
Find a latency change in a histogram, inspect a relevant request trace, and follow its logs for detail. That sequence makes investigation easier.
Scroll horizontally to view the complete diagram.
Read the diagram as text
Use aggregate metrics for impact, traces for timing and logs for detail. Exemplars are optional references, not guaranteed retained traces or p99 representatives. Time/service/version search provides a fallback.
An exemplar records context for a particular observation associated with a metric aggregate and may contain trace/span IDs. It depends on SDK reservoirs and filters, exporters, collectors, storage and UI support. It is not the unique request responsible for p99, nor is it guaranteed to identify a slow request. [10]
Tail sampling can drop the referenced trace; mismatched retention can also break the link. Provide a fallback trace search carrying service, version and time range when exemplars are unavailable.
Measure user-impact request counts, error rates and latency independently of trace selection where possible. A retained set favoring errors and slow requests cannot yield unbiased overall rates or p99 without appropriate treatment. For span-derived metrics, identify which sampling decisions occur before metric generation.
Averaging per-instance p99 values does not produce the global p99. Aggregate compatible histogram buckets or equivalent distributions before computing quantiles. Directly recorded metrics can still be lost during export; unsampled does not mean lossless.
6. Give attributes cardinality, privacy and trust boundaries
Investigating one affected tenant does not require putting every tenant ID on every metric. Use bounded dimensions such as tier and cell for aggregate metrics. Restrict individual identifiers to traces or logs where justified, with appropriate access controls. Trace indexing and retention also have costs.
Opaque or hashed identifiers are not automatically anonymous. Consider reidentification and joins with other data. Define permitted content across attributes, span names, events, log bodies and exceptions; do not casually collect payloads, authorization headers, tokens or email addresses.
Baggage carries context. It does not automatically become span, metric or log attributes without explicit copying or a processor. It also lacks built-in integrity protection and may propagate to third-party services. [8]
Untrusted inbound baggage
→ validate / allowlist / discard at the boundary
→ approved diagnostic context
→ explicit signal attributes where needed
Receiving role=admin or a tenant identifier does not authorize an action. Use authenticated identity and policy, reconstructing diagnostic context from them where needed. Apply destination-specific controls to outbound propagation.
7. Treat sampling as an evidence-retention policy
Keeping errors, slow traces and a normal baseline is useful, but it cannot guarantee retention of every error. Tail sampling cannot reconstruct spans never recorded or exported because of head sampling.
The following unvalidated fragment references Collector Contrib v0.161.0. It supplies no receiver, exporter, memory limit, TLS, authentication or trace-ID routing. Values are illustrative, not capacity recommendations. [11]
# Reference: Collector Contrib v0.161.0. NOT validated or executed.
# Fragment only: no receiver/exporter/TLS/auth/routing/memory limit.
# Illustrative values, not a sizing recommendation.
processors:
tail_sampling:
decision_wait: 10s
num_traces: 50000
policies:
- name: errors
type: status_code
status_code:
status_codes: [ERROR]
- name: slow
type: latency
latency:
threshold_ms: 1000
- name: baseline
type: probabilistic
probabilistic:
sampling_percentage: 5
This simple policy intends to retain traces matching errors or latency, plus a probabilistically selected baseline. Five percent is neither the total retention fraction nor a guaranteed observed percentage. Recheck decision semantics when adding drop policies, changing combinations or enabling feature gates.
Scroll horizontally to view the complete diagram.
Read the diagram as text
Head decisions, SDK export, trace-ID routing, tail decisions and backend storage each impose conditions. Error/slow retention policies cannot recover unrecorded spans, late evidence or loss during failures.
All spans of a trace must reach the same tail-sampler instance. Consider routing changes during scaling, restarts, buffer limits and late spans. decision_wait starts from receipt of the first span; it does not guarantee completion of arbitrary-length traces.
The latency policy uses the earliest observed start and latest observed end. That need not equal root request latency, especially when long asynchronous jobs share the trace. A failed-attempt span and a retry event inside an otherwise successful operation also match status policies differently. [11]
8. Review collector ordering and observe loss
Order processors according to their dependencies. Put memory limiting early. Processors such as k8sattributes that depend on original connection context must precede tail sampling, which reassembles spans into new batches. Enrichment supplying sampling attributes must also come first; batching commonly follows. “Always sample before enrichment” is not a valid universal rule. [11]
Unconditionally upserting deployment.environment.name=production on a shared gateway can relabel staging data as production. Establish trusted source identity, separated paths and rejection or quarantine for conflicting context. Enrichment is more than rewriting strings.
Filtering health-check spans can remove parents while leaving children, or affect another operation using the same path. Record precisely which metrics, logs and traces the filter removes.
Observe reception/refusal, sampling decisions, queues, export failures and backend queryability. Received minus exported is not automatically a drop rate. Batching, fan-out, retries, intentional sampling and different observation windows affect the comparison. A failure counter may count an attempt later retried successfully. Verify metric names, units and suffixes for the deployed version and exporter. [12]
Use an observation path outside the same failure domain to distinguish quiet applications from a broken collection pipeline. An exporter acknowledgement also differs from immediate searchability in the backend UI.
9. Turn the signal contract into acceptance obligations
Use critical user journeys as the contract boundary. This example records checkout entry, dependencies, final outcome, correlation and approved context.
{
"id": "OBS-CHECKOUT-001",
"status": "NOT_RUN",
"journey": "checkout",
"resource": {
"required_by_this_contract": [
"service.namespace",
"service.name",
"service.version",
"deployment.environment.name"
],
"conditional": ["service.instance.id", "corise.deployment.cell"],
"authority": "runtime deployment identity; reject/quarantine conflicts before trusted enrichment"
},
"operations": {
"entry": "HTTP POST /checkout using route-template naming",
"dependencies": ["inventory.reserve", "payment.authorize", "db.create_order"],
"retry": "failed attempt has ERROR and error.type; recovered parent uses final successful outcome",
"error_conventions": "apply HTTP client/server and operation-specific semantics; exception recording is not automatic status assignment"
}
}
If a standard HTTP instrument already answers the question, another custom instrument with the same meaning may add little. Add domain-specific meaning and record units and permitted label values.
Acceptance has three layers:
| Layer | What to inspect | What it does not establish alone |
|---|---|---|
| Instrumentation | In-memory exporter output: spans, outcomes, correlation and absence of sensitive content | Delivery beyond the SDK |
| Pipeline | Fixed fixtures: resource preservation, policy decisions, transformations and filtering | Backend queries and UI navigation |
| End-to-end | Synthetic journey through backend retrieval | Every possible failure or continuous losslessness |
SDK testing APIs vary by language and release. Include asynchronous completion and flushing in the design. Conceptual test code is not execution evidence. The supplied CSV is an acceptance plan, not a test suite. [4]
10. Exercise investigation through synthetics and incident drills
In staging, introduce delay only into a payment stub and ask an operator to investigate slow checkout. Define an isolated path without actual charges or notifications, scope, stopping conditions and recovery procedure first. This bundle executes no fault injection.
Use metrics for impact, exemplars or scoped search for traces, and correlate dependency, version, provider and logs. Record both supported conclusions and remaining hypotheses rather than stopping at “payment is slow.”
A synthetic header alone should not allow any public client to force expensive retention. Use an authenticated mechanism with a budget. Preserve a run’s correlation across signals without creating a new metric-label value for every run ID. A few requests cannot validate a probabilistic sampling percentage reliably.
If measuring outcomes, distinguish injection time, detection, evidence-supported understanding and the start of mitigation. Record clock differences and familiarity with the exercise. Do not publish “45 minutes to seven minutes” before measurement. Unmeasured fields in the supplied evidence record are null.
11. Return unanswered questions to the development backlog
If operators cannot compare providers, record the missing dimension rather than merely requesting more spans.
OBS-124
Question: Is only provider-a slow?
Current: payment.authorize is visible.
Gap: approved provider identity is absent.
Change: add corise.payment.provider from a bounded allowlist.
Acceptance: compare providers in retained traces and relevant metrics.
Status: proposed / NOT_RUN
Missing instrumentation is only one possible cause. Investigate inconsistent names, unindexed fields, retention, sampling, access permissions and noise. Removing spans or improving query navigation can be the appropriate change.
Scroll horizontally to view the complete diagram.
Read the diagram as text
Move from incident/drill questions through gaps and contract/instrumentation/query changes to three-layer acceptance. Track owners and versions; only observed results close evidence obligations.
SDK, collector and semantic-convention upgrades can change data contracts. Review names and stability at the chosen baseline, including migration from deployment.environment to deployment.environment.name and database attributes such as db.system.name. Conventions do not all share one maturity level. Plan queries, dashboards, alerts and coexistence with older producers. [3][13]
Platform teams own transport and common schema, application teams own domain meaning, and operations contribute questions and investigation paths. An unassigned owner does not constitute an adopted operational contract.
12. Make explanation part of completion
At release review, ask what evidence will explain a slow new feature. Critical dependencies should be visible, version and outcome identifiable, signals correlated, and retention limits understood.
Detection and explanation are useful design roles, not a strict division between monitoring and observability. Metrics can explain causes and traces can reveal anomalies. Separating user-impact paging from diagnostic evidence is often useful in practice.
Completion means an operator can answer important inventory questions through the real collection path and backend, recording evidence and limitations. It does not mean predicting every unknown. Preserve the ability to combine new queries and return missing evidence to instrumentation work.
Move from signal volume to evidence that supports explanation. The development loop after OpenTelemetry adoption keeps resource identity, context, meaning, correlation and retention policy evolving together.
The related observability-platform modernization case provides attributed design context. This checkout scenario, fragment, acceptance plan and any future drill results are not presented as executed evidence from that engagement.
For shared monitoring failures and independent notification, see Who Notices Monitoring Failure?.
References
Official source revisions and the Collector Contrib reference release are pinned below. They are editorial baselines, not a validated running configuration or a compatibility claim.
- Observability Primer
- Service Resource
- Deployment Environment (Deprecated deployment attributes)
- Instrumenting Libraries
- Recording Errors
- HTTP Spans
- Messaging Spans
- Baggage
- Logs Data Model
- Metrics Data Model / Exemplars
- Tail Sampling Processor v0.161.0
- Collector Internal Telemetry
- Database Spans
- tracing-opentelemetry 0.33.0 README