1. What It Is
You can't debug a distributed system from logs alone. We instrument metrics (what), logs (why), and traces (where in the call graph) β and tie them together with a shared trace ID.
What:
Observability measures a distributed system's internal health based on its external outputs: structured Logs, real-time numeric Metrics, and distributed Traces.
Primary purpose:
Diagnosing root causes of failures, tracing latencies across microservice networks, and keeping system downtime minimal.
Usually used for:
Microservices health checking, cloud cluster monitoring, API gateway profiling, and alert automation.
2. Core Mental Model
Metrics tell you something is wrong; logs tell you what happened locally; traces tell you where in the call chain. You need all three, but at different cost β concept #33 Time-Series Storage covers metrics backends; problem #33 Distributed Tracing is the tracing HLD.
πͺ΅ The Log Microscope
Logs detail exactly what happened inside a single node process at a specific timestamp. Always structure logs as searchable JSON key-values.
π The Metric Radar
Low-cost time-series counters. Monitor the four golden signals: Latency, Traffic (QPS), Errors (5xx), and Saturation (CPU/RAM).
πΈοΈ Distributed Request Tracing
Trace request flows across servers. Gateways generate a `trace_id` header propagated to all downstream RPC calls to group child `spans` together.
In the room
Candidates say "we'll add monitoring" without the three pillars. Mention sampling for high-QPS traces, RED/USE metrics for services, and structured logging with trace_id propagation. OpenTelemetry is the safe name-drop.
3. Why It Matters in HLD
You cannot fix what you cannot see β metrics, logs, and traces are the production debugging stack. Three lenses:
Needed When:
Deploying multi-tier distributed microservices where identifying whether API latency originates from database locks or network lag is impossible without traces.
Avoids:
Blind debug sessions during outages, slow response bottlenecks, bloated storage bills, and missing critical error warnings.
Optimizes For:
Mean Time To Resolution (MTTR), tail latency profiles (P99), infrastructure capacity planning, and alert signal accuracy.
4. Architecture & Data Flow
Walk a traced request as interview steps. Step 1 β Ingress: gateway generates trace-id, spans root span. Step 2 β Propagate: inject trace context headers on every RPC hop. Step 3 β Span per service: each microservice records start/end, tags, errors. Step 4 β Export: spans batch to collector (Jaeger, Tempo). Step 5 β Correlate: link metrics (RED) and logs via trace-id for incident drill-down.
In the room
Mention RED (rate, errors, duration) for each service and trace-id propagation across gRPC β that shows ops maturity beyond drawing boxes.
5. Key Characteristics
Metrics vs logs vs traces β three pillars with different cardinality and cost β we compare:
- The three pillars β complementary views of system behavior:
| Observability Pillar | Data Model & Mechanic | Architectural Role |
|---|---|---|
| Logs (Event Records) | Structured, timestamped text lines (JSON) recording discrete local events. |
|
| Metrics (Numeric Counters) | Pre-aggregated time-series numeric values (CPU%, QPS, memory profiles). |
|
| Traces (Request Pathways) | Correlated spans tracking a single request flow across multiple microservices nodes. |
|
- SLI / SLO / error budget β an SLI is a measured signal (e.g., p99 latency < 300ms); an SLO is the target (99.9% of requests meet it over 30 days); the error budget is how much failure you can afford before feature freeze. Burn-rate alerts fire when budget consumption accelerates.
- Cost order of magnitude β metrics: ~2 bytes/sample compressed, cheap at millions/sec. Logs: ~500 bytesβ2 KB per line, 100β1000x more expensive than metrics at same QPS. Traces: ~1β5 KB per span, most expensive β sample aggressively on success paths, retain all errors.
6. Strategic Tradeoffs
Full visibility trades storage cost and instrumentation overhead β we state both:
| Benefit | Cost |
|---|---|
| Faster incident diagnosis (correlating traces with logs shortens mean time to resolution) | Storage cost (retaining all logs and traces at high QPS is expensive without sampling) |
| Precise Tail Latency Maps (distributed tracing exposes hidden down-stream queuing latencies (P99) immediately) | Application CPU Overhead (injecting span context and exporting trace payloads consumes compute cycles) |
7. Failure / Bottleneck Awareness
Cardinality explosion, missing context propagation, and sampling bias β we name mitigations:
Problem: At 50k requests/sec, storing every span and debug log fills Elasticsearch quickly and can cost more than the services being monitored.
Mitigation: Sample traces (e.g., 1% of success paths) but always retain failures and slow requests (5xx, latency above SLO). Use tail-based sampling where the collector decides after the request completes.
Problem: When transitioning from HTTP API calls to asynchronous message queues (e.g., publishing to Kafka), developers forget to serialize and pass trace context headers. The trace breaks, appearing as two isolated request fragments, blinding root-cause analyses.
Mitigation: Propagate W3C Trace Context headers into Kafka record metadata (and other async handoffs) so spans stay linked across sync and async boundaries.
8. Common HLD Usage
Microservice latency debugging and SLO dashboards rely on observability stacks:
| Instrumented System | Observability Engine | Architectural Rationale |
|---|---|---|
| E-Commerce Checkout Pipeline | OpenTelemetry + Jaeger Tracing | Traces propagate unique context headers across payment, cart, and inventory nodes, pinpointing exactly which downstream service triggers checkout timeouts. |
| API Gateway Metrics | Prometheus + Grafana | Collects low-overhead QPS, 500 error counts, and P99 latency charts dynamically, triggering PagerDuty alerts on failures. |
9. Decision Signals
Add tracing when requests cross more than three services or tail latency is an SLA:
- You are designing distributed microservice graphs where diagnosing single-node latency spikes from standard metrics is impossible.
- You need to configure automated pager alerting systems that monitor P99 latency bounds.
- You want to establish structured, searchable audit trails mapping transaction execution steps.
11. Deep Dive (Optional)
OpenTelemetry Context Propagation under the Hood
To trace a request end-to-end, OpenTelemetry propagates a runtime context across network boundaries using inject/extract on each hop:
- Context Object: An in-memory key-value map holding metadata, primarily:
trace_id: a7e39f88b022e11a\nparent_span_id: c08f921b332d\nspan_id: e39b2b8c2d1c\ntrace_flags: 01 (Sampled)
- Injection (Client Side): When Service A calls Service B via HTTP/gRPC, the tracing library serializes this Context map into standardized headers (W3C Trace Context format):HTTP
traceparent: 00-a7e39f88b022e11a-e39b2b8c2d1c-01 - Extraction (Server Side): Upon receiving the request, Service B's tracing middleware extracts the `traceparent` header, instantiates B's parent context, and creates a child span linked to the caller.
This trace-correlation flow runs transparently inside application middlewares, enabling real-time trace mapping with zero manual developer coding.
Review
How helpful was this walkthrough?
Click a star to rate. We actively use this feedback to refine and update our system design content.
Discussion
Share your thoughts, ask questions, or help others.