Observability

Observability for Small Distributed Systems

Build useful logs, metrics, and traces for a few services without creating an expensive monitoring project.

On this page
  1. Begin with questions, not dashboards
  2. Metrics show shape over time
  3. Logs explain individual events
  4. Traces show a request across boundaries
  5. Correlation needs common fields
  6. Instrument the failure path
  7. Keep the system proportional

Begin with questions, not dashboards

Observability is the ability to infer what a system is doing from the signals it emits. For a small distributed system, the first challenge is not a lack of dashboards. It is not knowing which questions the signals should answer. When a request is slow, an operator may need to find the affected operation, identify the service that introduced delay, and distinguish a dependency problem from local saturation.

Write down a handful of questions before choosing tools. Which operations fail most often? Where does latency increase? Are workers keeping up with incoming work? Did a deployment change the pattern? These questions suggest what to instrument and prevent a monitoring stack from becoming a collection of charts with no operational purpose.

Metrics show shape over time

Metrics are numeric measurements collected across time. They are useful for rates, ratios, and distributions: request count, error count, queue depth, process memory, and request duration. A counter can show whether errors increased after a release; a histogram can show whether tail latency changed even when the average remains steady.

Keep labels bounded. A metric label that contains a user ID, request ID, raw URL, or unbounded exception text can create a new time series for every value. Cardinality grows quickly, increasing memory and query costs while making the result harder to interpret. Put unique identifiers in logs or traces instead, and normalize paths into a finite set of route names.

Start with a small service-level set: request rate, error ratio, latency distribution, saturation of a constrained resource, and queue age or depth for asynchronous work. Add metrics when they answer a question or support an action. Instrumentation that does not change diagnosis or operation is a candidate for removal.

Logs explain individual events

Structured logs capture details about a specific event. A JSON record with timestamp, severity, service, operation, outcome, and trace identifier can be searched consistently across services. A plain sentence may be readable, but stable fields make aggregation and correlation much easier.

Include enough context to explain an event without copying entire request payloads. Personal data, credentials, session tokens, and customer content should not enter logs by default. Redaction is safer when sensitive fields are never emitted, rather than filtered later by a collector. Retention and access controls are part of logging design too, not just storage settings.

Avoid logging every successful event at high volume if a metric already captures its aggregate behavior. Reserve detailed records for state transitions, failures, and selected sampled events. A flood of duplicate messages can hide the one record that explains why a request failed.

Traces show a request across boundaries

Distributed traces connect work across services. A trace contains spans representing operations, with parent-child relationships and timing. When a request crosses an API, a queue, and a worker, trace context can show where time was spent and which dependency call failed.

Tracing requires context propagation. If a service creates a new trace identifier at every hop, the result becomes several unrelated fragments. Propagate a standard context through HTTP headers and, where appropriate, message metadata. Background jobs may execute much later than their initiating request; decide whether to continue the original trace or start a new trace linked to the scheduling event.

Sampling is necessary at scale, but a small system can still sample selectively. Keep representative normal traces and increase the chance of retaining error or slow traces. A trace is not a replacement for metrics: it explains an individual request, while metrics show how often a pattern occurs.

Correlation needs common fields

Signals become more useful when they share service names, environment names, deployment versions, and trace identifiers. A dashboard may show error rate by version; a log search can filter the same version; a trace can reveal the affected dependency. Agree on a small naming convention early, because inconsistent names make cross-service questions unnecessarily difficult.

Do not overload one field with several meanings. A deployment ID is not an environment, a region is not a host, and a customer key is not a trace ID. Clear dimensions help people compare like with like and keep metric cardinality under control.

Instrument the failure path

Many systems emit good signals on successful requests and almost none during failure. Record timeout outcomes, retries, circuit-breaker transitions, worker lease changes, rejected messages, and dead-letter counts. For queues, message age is often more useful than depth alone: a small queue with old messages can be more concerning than a large queue moving quickly.

Signals should survive partial failure when possible. A service that cannot reach its central log collector should still write locally or to a bounded buffer. An observability pipeline should not block user work indefinitely. Decide what happens when telemetry cannot be exported, and cap local buffering so the monitoring path cannot exhaust the application.

Keep the system proportional

A practical first version may need a metrics endpoint, structured application logs, and trace propagation for a few critical flows. It does not necessarily need a complex topology, every available dashboard, or full retention of all events. Choose retention from actual diagnostic needs and available storage. Set alerts on symptoms that require action, not every metric that crosses an arbitrary line.

Review observability in incident exercises. Ask someone unfamiliar with a service to use the available signals to locate an injected timeout or stuck worker. If they cannot answer what failed, where it happened, and how broad the impact is, improve the instrumentation or its labels. Useful observability is not measured by the number of charts. It is measured by how much uncertainty it removes when the system behaves unexpectedly.