Tools / Monitoring and Observability Interview Questions
What is a service mesh and how does it enhance observability?
A service mesh is an infrastructure layer — deployed alongside your application services — that manages service-to-service communication. It intercepts network traffic using sidecar proxies (Envoy is the most common) injected into every pod, handling load balancing, mutual TLS, retries, circuit breaking, and observability without any application code changes.
From an observability perspective, a service mesh provides L7 telemetry automatically for every service-to-service call in the mesh. Because Envoy intercepts all HTTP/gRPC traffic, it can emit:
- Metrics: Request rate, error rate, and latency (p50/p95/p99) per source-destination service pair — exactly the RED method signals, automatically, for every microservice.
- Traces: Envoy can propagate trace context headers and generate spans for every hop, contributing to distributed traces without application-level instrumentation.
- Access logs: Structured per-request logs with HTTP method, path, status, upstream cluster, and duration.
Istio (using Envoy) and Linkerd are the two dominant service meshes. Istio integrates with Prometheus (via native Envoy metrics scraping), Jaeger/Zipkin (for tracing), and Kiali (a service mesh topology visualization tool). Linkerd has its own lightweight Rust-based proxy with built-in Prometheus metrics.
The trade-off is operational complexity: managing a service mesh's control plane (istiod, Linkerd control plane) adds significant overhead, and sidecar injection adds latency and resource consumption per pod.
More Related questions...