Production Container Observability: Monitor Docker Containers, Logs, Metrics, and Performance
Production container observability is the discipline of collecting metrics, logs, and traces from every layer of a containerized system — the container runtime, the orchestrator, and the application inside — so that engineers can understand why something is happening, not just that something happened. Monitoring tells you a service is down; observability is what lets you figure out why, three layers deep, without guessing. That distinction is the difference between a two-minute diagnosis and a two-hour war room. 🔎
The stakes rise sharply with containers and Kubernetes, because the thing you're debugging rarely stays still. A pod that OOM-killed five minutes ago may already be gone, rescheduled to a different node with a new IP and a new set of logs starting from zero. Without metrics, logs, and traces captured and correlated continuously — not just checked reactively — that failure is effectively unrecoverable evidence. This post walks through how the three pillars of observability actually work in a container and Kubernetes environment, grounded in official Kubernetes and Docker documentation. ⚙️
📑 In This Post
- 1. Foundations: the three pillars, and why containers make them harder
- 2. Metrics: how Kubernetes and containers expose system state
- 3. Logs: from container stdout to a searchable pipeline
- 4. Traces: following one request across many containers
- 5. Real example: probes as the feedback loop for orchestration
- 6. Implementation: wiring metrics, logs, and traces together
- 7. From signals to action: dashboards, alerts, and on-call
- 8. Enterprise rollout: governance for observability data
- 9. Common mistakes
- 10. FAQ
- 11. References & further reading
- 12. Summary
🔀 Quick Comparison: The Three Observability Pillars
| Pillar | Answers | Typical container-world source |
|---|---|---|
| Metrics | Is the system healthy right now, and what's the trend? | kubelet /metrics endpoints, cAdvisor, application /metrics endpoint |
| Logs | What exactly happened, in what order, with what detail? | Container stdout/stderr captured by the runtime and kubelet |
| Traces | Where did time go as one request crossed multiple containers? | Application instrumentation exporting spans to a trace collector |
1. Foundations: The Three Pillars, and Why Containers Make Them Harder
Kid-friendly analogy: imagine trying to understand what happened at a school fair using only a single photo taken at the end of the day. You'd see the aftermath but miss every decision that led there. Now imagine the fair tents get torn down and rebuilt in a different spot every few minutes — that's a container getting rescheduled.
Kubernetes' own documentation frames this precisely: observability is the process of collecting and analyzing metrics, logs, and traces — often called the three pillars of observability — to understand the internal state, performance, and health of a cluster, with control plane components and add-ons generating and emitting these signals continuously. What makes this harder in a containerized environment than on a static server is exactly the property that makes containers useful in the first place: they're disposable. A pod can be killed, rescheduled to a different node, and replaced with a fresh instance in seconds, and every piece of local state — including anything written only to that container's own filesystem — disappears with it unless something captured it first.
What it does: the three pillars together give a layered picture — metrics for trend and alerting, logs for detailed forensic record, traces for cross-service request flow.
Why it's needed: in a distributed, ephemeral system, no single signal is sufficient. A metric tells you latency spiked; a log tells you which request failed and why; a trace tells you which downstream container the delay came from.
What fails without it: teams that rely on logging in alone, or metrics alone, consistently hit a wall the first time an incident spans more than one container — they can see that something is wrong, but not which of a dozen ephemeral pods caused it or where in the request path the time went.
🎯 Use this when you're deciding what a "minimum viable observability" setup needs to cover before going to production.
2. Metrics: How Kubernetes and Containers Expose System State
Kid-friendly analogy: a car's dashboard doesn't tell you the life story of the engine — it gives you a handful of numbers (speed, fuel, temperature) updated continuously, so you can react before something breaks. Metrics are that dashboard for a cluster.
What it does: Kubernetes components emit metrics in the Prometheus text-based exposition format, structured plain text designed to be readable by both people and machines. Most components expose these on a /metrics HTTP endpoint. The kubelet specifically exposes several distinct endpoints, including /metrics/cadvisor for per-container resource usage, /metrics/resource, and /metrics/probes, each with its own lifecycle separate from the kubelet's own metrics.
Why it's needed: metrics are lightweight, numeric, and cheap to store at high resolution over long time ranges, which makes them the right tool for dashboards, trend analysis, and threshold-based alerting — something logs and traces are both too heavy and too unstructured to do efficiently at scale.
How it works step by step:
- Each component (kubelet, API server, controllers, and instrumented applications) exposes current metric values on its own
/metricsendpoint. - A metrics scraper, typically Prometheus Server, periodically polls those endpoints and stores the values in a time series database.
- Access to the metrics endpoint can be restricted through RBAC, requiring a ClusterRole that explicitly permits reading the non-resource
/metricsURL. - A visualization layer, commonly Grafana, queries that time series database to render dashboards and evaluate alerting rules.
- For multi-cluster or multi-cloud visibility, a distributed time series store can sit behind Prometheus to aggregate metrics across many clusters.
Kubernetes also assigns metrics a formal stability lifecycle — alpha, beta (or stable), deprecated, hidden, then deleted — so that dashboards and alerts built on stable metrics aren't broken by a routine Kubernetes upgrade, while alpha metrics carry no such guarantee and can change at any time.
3. Logs: From Container Stdout to a Searchable Pipeline
Kid-friendly analogy: a flight recorder doesn't summarize a flight — it writes down everything, in order, so that if something goes wrong, investigators can reconstruct exactly what happened rather than relying on memory.
What it does: logs provide a chronological record of events inside applications, system components, and security-relevant activity. Kubernetes documentation describes how container runtimes capture a containerized application's output from its standard output and standard error streams, and while runtimes implement this differently, the integration with the kubelet is standardized through the CRI logging format — which is what lets kubectl logs work consistently regardless of which container runtime a node uses.
Why it's needed: when a container is rescheduled or deleted, anything it only wrote to its own local disk is gone. Capturing stdout/stderr at the runtime level, outside the container's own lifecycle, is what preserves that record past the container's own existence.
What fails without it: a container that crashes and is immediately replaced leaves no trace of what it was doing at the moment of failure if logs weren't being captured and shipped somewhere durable — the most useful diagnostic evidence for the incident disappears along with the container.
At the Docker engine level, this same capture happens through logging drivers. Docker documentation confirms the default logging driver is json-file, which caches container logs as JSON on disk, and explicitly warns that this default performs no log rotation, which can exhaust disk space for containers that produce heavy output; Docker's documentation recommends the local logging driver instead for production use, since it rotates logs by default and uses a more space-efficient file format. Changing the daemon's default logging driver, notably, only affects containers created after the change — existing containers keep whatever driver they started with.
🎯 Use this when setting the daemon or node-level logging defaults for any environment that will run unattended.
4. Traces: Following One Request Across Many Containers
Kid-friendly analogy: think of a relay race where each runner logs their own lap time separately. Individually, those times don't tell you where the team lost the race. A trace is the stopwatch that follows the baton the whole way around, showing exactly which leg was slow.
What it does: a trace follows a single request as it moves through multiple services and containers, recording how long each hop took and where it went next.
Why it's needed: once an application is split across several containers — an API gateway, an auth service, a database proxy, a backend — a slow response could originate in any one of them, or in the network hops between them. Metrics tell you the overall request got slow; only a trace tells you which specific hop in which specific container caused it.
What fails without it: debugging a latency regression across a multi-container request path without tracing usually means manually correlating timestamps across several different containers' logs by hand — slow, error-prone, and often impossible once containers have already been rescheduled and their logs scattered.
Kubernetes' own observability documentation names traces as one of the three pillars precisely because it recognizes that in a distributed, containerized architecture, request flow across service boundaries can't be reconstructed from logs or metrics alone.
🎯 Use this when your architecture involves more than two or three containers handling a single user-facing request.
5. Real Example: Probes as the Feedback Loop for Orchestration
Kid-friendly analogy: a lifeguard doesn't wait for someone to shout for help — they periodically scan the pool, checking on every swimmer, and act the moment something looks wrong. A probe is that periodic check, run by Kubernetes on every container.
This is a concrete, verifiable example of observability feeding directly back into orchestration decisions rather than just sitting on a dashboard. Kubernetes documentation defines three distinct probe types, each answering a different question:
- Startup probe: verifies whether the application inside a container has finished starting. While configured, it disables both liveness and readiness checks until it succeeds, giving slow-starting applications room to initialize without being killed prematurely by the kubelet.
- Liveness probe: determines when to restart a container — for example, catching a deadlock where the process is running but no longer making progress. If a container fails its liveness probe repeatedly, the kubelet restarts it. Liveness probes do not wait for readiness probes to succeed.
- Readiness probe: determines when a container is ready to accept traffic, useful while an application performs time-consuming initialization such as establishing connections or warming a cache. If a readiness probe fails, Kubernetes removes the pod's IP address from the endpoints of every Service that matches it, and readiness probes continue running for the container's entire lifecycle, not just at startup.
Kubernetes documentation is explicit that liveness probes must be configured carefully to indicate genuinely unrecoverable failure — an incorrectly implemented liveness probe can lead to cascading failures: containers restarting under load, failed client requests as the application becomes less scalable, and increased load on the remaining healthy pods as failed ones are taken out of rotation.
6. Implementation: Wiring Metrics, Logs, and Traces Together
A workable observability stack layers together several purpose-built pieces rather than one tool doing everything:
- Metrics collection: a scraper (typically Prometheus) polls
/metricsendpoints across the cluster, including per-container resource metrics surfaced through/metrics/cadvisoron each kubelet. - Log shipping: a node-level agent tails each container's captured stdout/stderr output and forwards it to centralized log storage, since relying on
kubectl logsalone loses history the moment a pod is deleted. - Trace instrumentation: application code is instrumented to emit spans for each unit of work, tagged with identifiers that let a single request be reassembled across every container it touched.
- Correlation: metrics, logs, and traces are tagged with shared identifiers — pod name, namespace, node — so an engineer can pivot from a metric spike, to the relevant logs, to the specific trace, without starting each search from scratch.
- Retention policy: high-cardinality, high-volume signals like raw logs and traces are typically retained for a shorter window than aggregated metrics, balancing storage cost against how far back an investigation might need to reach.
🎯 Use this when moving from "we have a dashboard" to an observability stack an on-call engineer can actually debug an incident with.
7. From Signals to Action: Dashboards, Alerts, and On-Call
Observability data only has operational value once it triggers a human or automated response at the right moment — not too late, and not so often that alerts get ignored.
- Dashboards should answer "is the system healthy right now" at a glance, built from the same stable metrics that won't silently change meaning across a Kubernetes upgrade.
- Alerting rules belong on symptoms that matter to users — elevated error rate, rising latency, restart loops — rather than on every possible internal metric, to avoid alert fatigue that trains engineers to ignore notifications.
- Runbooks tied to each alert turn "something's wrong" into a concrete first action, cutting the time between an alert firing and someone starting the right diagnostic step.
- Escalation paths should be defined before an incident, not improvised during one, especially for alerts that can fire outside business hours.
8. Enterprise Rollout: Governance for Observability Data
Observability pipelines themselves are production systems with real governance needs, not a side project:
- Access controls on metrics and logs. Kubernetes' own guidance shows that reading the
/metricsendpoint under RBAC requires an explicit ClusterRole grant — treat that same discipline as the baseline, and extend it to log storage and trace backends, since both routinely contain sensitive request data. - Ownership of the observability stack. Assign a clear owner for the Prometheus/Grafana stack, log pipeline, and trace backend themselves — they're production dependencies, and an outage in the observability stack during an incident is its own kind of disaster.
- Privacy of production-derived data. Logs and traces can capture user data incidentally through request payloads; define what gets redacted or excluded before it's shipped, not after a privacy review finds it.
- Budget controls on retention and cardinality. High-cardinality metric labels (per-request IDs, for instance) can silently balloon storage costs; review label design and retention windows as a recurring cost-governance task, not a one-time setup decision.
- Dashboards and alerts as reviewed artifacts. Treat alerting rule changes with the same review process as application code changes, since a broken alert is a silent gap in coverage that isn't discovered until the next incident.
- Incident response integration. Make sure the observability stack's own alerts route to the same on-call system as everything else, so an observability outage doesn't become a blind spot precisely when it's needed most.
🎯 Use this when observability tooling moves from a personal debugging habit to infrastructure other teams depend on.
9. Common Mistakes
- Leaving the default json-file logging driver unrotated in production. Because Docker's own documentation confirms this default performs no log rotation, a single chatty container can silently fill a node's disk over days or weeks, eventually causing failures unrelated to whatever the container itself is doing.
- Configuring only a liveness probe with no startup probe for a slow-starting application. Since liveness probes don't wait for readiness and can begin checking before initialization finishes, this reliably produces restart loops for any container whose startup genuinely takes longer than the liveness probe's timing allows.
- Alerting on every available metric instead of user-facing symptoms. This produces alert fatigue, where engineers start ignoring notifications because most of them aren't actionable — exactly the failure mode that causes a genuinely important alert to get missed.
- Treating kubectl logs as durable log storage. Once a pod is deleted, its logs accessible through kubectl are gone unless a separate log-shipping pipeline captured them first — relying on kubectl alone means losing the most relevant evidence right after the incident that made it relevant.
- Building metrics dashboards on alpha-stability metrics. Kubernetes' own metric lifecycle explicitly warns that alpha metrics carry no stability guarantee and can change or disappear at any time; a dashboard built on one can silently break on a routine version upgrade.
- No correlation identifiers linking metrics, logs, and traces. Without shared tags like pod name or request ID across all three signal types, an engineer has to manually bridge a metric spike to the relevant logs and trace by timestamp guesswork, which is slow exactly when speed matters most.
❓ FAQ
What's the actual difference between monitoring and observability?
Monitoring typically means watching a predefined set of signals for known failure conditions. Observability is broader: it's having enough correlated metrics, logs, and traces available that you can investigate a failure mode nobody predicted in advance, not just the ones a dashboard was built to catch.
Why does my container's disk fill up even though the application isn't writing much data itself?
This is commonly the default json-file logging driver, which Docker's documentation confirms performs no log rotation by default. A container that logs moderately but runs for a long time can still accumulate a large unrotated log file over time; switching to a rotating driver, such as the local logging driver, addresses this directly.
Do I need a startup probe if I already have a liveness probe?
If your container starts quickly, often not. If it has any meaningfully slow initialization step, yes — because liveness probes don't wait for readiness to succeed, a container that's still starting can fail its liveness probe and get killed before it ever finishes, which a properly timed startup probe prevents.
Are logs, metrics, and traces stored the same way?
No, and that's intentional. Metrics are compact numeric time series suited to long retention at low cost. Logs and traces are far higher volume and higher cardinality, so they're typically retained for a shorter window and often sampled, trading some completeness for manageable storage and query cost.
Is it safe to build permanent dashboards on any Kubernetes metric I find?
Not on alpha-stability metrics. Kubernetes documents a formal metric lifecycle, and alpha metrics have no stability guarantee — they can be modified or removed at any time. Stable metrics are the safer foundation for dashboards and alerts you expect to survive a Kubernetes version upgrade.
🔗 References & Further Reading
- Kubernetes Docs — Observability (official concepts documentation, kubernetes.io)
- Kubernetes Docs — Metrics For Kubernetes System Components (official documentation on /metrics endpoints and metric lifecycle)
- Kubernetes Docs — Liveness, Readiness, and Startup Probes (official documentation)
- Docker Docs — Configure Logging Drivers (official documentation, docs.docker.com)
Kubernetes, Docker, Prometheus, and Grafana are trademarks of their respective owners; these names are referenced for identification only.
📝 Summary
- Observability rests on three pillars — metrics, logs, and traces — that together answer questions no single signal can answer alone.
- Containers make observability harder specifically because they're ephemeral: local state, including unshipped logs, disappears when a container is rescheduled.
- Kubernetes components expose metrics in Prometheus format on /metrics endpoints, with a formal stability lifecycle from alpha to stable to deprecated.
- Container logs are captured from stdout/stderr by the runtime, standardized through the CRI logging format, and depend on a properly configured, rotating logging driver in production.
- Traces reconstruct a single request's path across multiple containers, filling a gap metrics and logs can't close alone.
- Liveness, readiness, and startup probes are a concrete, verifiable example of observability signals feeding directly into orchestration decisions.
- A working stack correlates metrics, logs, and traces with shared identifiers, and deliberately balances retention and sampling against cost.
- Enterprise governance covers access controls on observability data itself, ownership of the pipeline, data privacy, and treating alert rules as reviewed artifacts.
- Most production observability failures trace back to unrotated logs, missing startup probes, alert fatigue from over-broad alerting, or metrics dashboards built on unstable foundations.
That's the full path from a container's raw stdout to a governed, correlated production observability stack — happy shipping! 🚀
Comments
Post a Comment