Infrastructure

Understanding Service Health Checks

Separate liveness, readiness, and dependency health so checks help operators instead of amplifying failures.

On this page
  1. A health check is a decision input
  2. Liveness asks whether the process can continue
  3. Readiness asks whether to send work here
  4. Dependency health is an observation
  5. Timeouts and thresholds shape behavior
  6. Rollouts need more than one green response
  7. Test failure behavior deliberately

A health check is a decision input

Health checks look simple because their output is small: success or failure. The difficult part is deciding what that result should cause. A scheduler may restart a process, stop routing traffic to an instance, delay a rollout, or alert an operator. A check that combines unrelated conditions can trigger the wrong action at the wrong time.

Start by naming the decision. Is the question whether a process is alive, whether it can serve a request now, or whether a dependency is functioning? Those are different observations, and they belong at different layers. One endpoint called healthy may hide all three until a failure makes the ambiguity costly.

Liveness asks whether the process can continue

A liveness check is intended to detect a process that is irrecoverably stuck or unable to make progress. It should be cheap and should test something the process can know locally. A minimal HTTP handler can confirm that the event loop is responding; a worker may report whether its main loop has advanced recently.

Liveness should not fail just because a database or remote API is temporarily unavailable. If every instance restarts when the same dependency has an outage, the check converts a partial dependency failure into a fleet-wide restart event. Recovery can become slower because connections, caches, and in-flight work are discarded at the same time.

There is a limit in the other direction. A process that answers a trivial liveness endpoint while its critical worker thread is dead may not be useful. The check needs to represent the component whose failure the restart mechanism can actually repair. If no restart would help, report the condition through metrics or logs instead of a liveness signal.

Readiness asks whether to send work here

Readiness is about accepting new traffic. During startup, an instance may be alive but still loading configuration, warming a required cache, or opening its listener. It should not receive production requests until it can handle them safely. During shutdown, it should become unready before it exits so routers have time to stop sending new work.

A readiness response can include required dependencies, but that choice must match the service's behavior. If every request needs a database connection, database availability may be part of readiness. If the service can serve cached or degraded responses, failing readiness on any database timeout may remove useful capacity. Prefer a bounded check that reflects the actual contract.

The control plane matters too. A readiness check that only changes a local HTTP response has no routing effect by itself. The load balancer or orchestrator must consume it, and its update delay must be understood. Readiness is most valuable when the whole chain from probe to traffic removal is tested.

Dependency health is an observation

Dependency checks are often better exposed as metrics or diagnostic endpoints than as a global healthy flag. Record connection failures, latency, queue depth, and circuit-breaker state. An operator can then tell whether one service is degraded because its own process is sick or because a dependency is slow.

Avoid running an expensive deep test on every probe interval. A check that creates a database transaction or calls several third-party APIs can add load exactly when the system is under stress. Use a short timeout, avoid writes where possible, and decide whether the probe should use a separate connection pool. A check that blocks on the same saturated resource it is meant to diagnose can worsen the saturation.

Timeouts and thresholds shape behavior

Probe settings are part of the reliability design. The interval controls how quickly a failure may be noticed. The timeout limits how long a probe occupies resources. A failure threshold filters transient noise, while a success threshold can prevent rapid flapping back to ready.

Choose values from measured startup and response behavior, not from copied defaults alone. If a service normally takes forty seconds to initialize, a probe that begins after five seconds and allows only two failures will restart it repeatedly before startup can finish. If a network timeout is longer than the upstream request budget, the probe may lag behind the actual user experience.

Every active check should have a clear budget. Its work should be bounded, predictable, and safe to repeat. Keep the endpoint independent of authentication only where the network boundary makes that safe; never expose sensitive diagnostics merely because a monitoring system needs a status code.

Rollouts need more than one green response

During a rollout, a single successful readiness result may be insufficient evidence that a new version is stable. Observe the service over a period that includes representative traffic and background work. Track error rate, latency, restarts, and queue behavior. A readiness probe answers whether traffic can be sent now; it does not prove that the version behaves correctly under load.

For a small service, a useful arrangement is often a local liveness check, an application readiness check, and separate dependency metrics. A deployment controller can use readiness to place traffic, while alerts use sustained error or saturation signals. This separation gives each mechanism a narrower and more explainable responsibility.

Test failure behavior deliberately

Test the check with the process starting slowly, the dependency timing out, the dependency returning an error, and the application shutting down. Verify that each state produces the intended routing or restart behavior. Observe the timeline from failure to detection to action to recovery.

Document which dependency failures make the service unready and which create degraded responses. That short list is more useful than a vague claim that the service is healthy. Good checks reduce uncertainty for automation and for people. They make the system's response to failure explicit, and they avoid turning one weak signal into a second failure.