Designing Reliable Background Workers
Build worker systems that tolerate retries, crashes, slow dependencies, and queue pressure without hiding failure.
On this page
A queue changes when work happens
A background worker separates accepting a task from completing it. That can keep user-facing requests short and absorb uneven workloads, but the queue does not make the work reliable on its own. A worker can crash after performing an external side effect but before acknowledging the message. A dependency can time out after completing the request. A message can be delivered again.
Design around those ordinary failure cases. A worker system should make it possible to determine what work is pending, what is in progress, what has failed, and what has completed. The delivery model may be at-most-once, at-least-once, or closer to exactly-once within a narrow transactional boundary. State the actual guarantee instead of assuming that the broker supplies it everywhere.
Make handlers safe to retry
At-least-once delivery is common because a broker can redeliver work when acknowledgements are lost or a worker exits. The handler should therefore be idempotent where possible: processing the same task twice produces the same intended state as processing it once.
Use a stable task identifier and record the result of applying it. For database changes, a unique key or transaction can prevent duplicate writes. For external APIs, pass an idempotency key if the provider supports one. When the external action cannot be made idempotent, use a state machine that records intent and outcome, then define how an uncertain result is reconciled.
Do not acknowledge a message before the durable work is complete. Acknowledging too early can lose a task if the process crashes immediately afterward. Acknowledging too late can cause repeat delivery, which is safe only when duplicates are expected. Choose the acknowledgement point to match the side-effect boundary.
Bound retries and preserve useful delay
Retries help with transient failures but can amplify an outage. Use a bounded attempt count or deadline, exponential backoff, and jitter so many workers do not retry in lockstep. Classify errors: a temporary connection reset may be retried, while invalid input will not become valid by waiting.
Keep retry state visible. Record attempt count, last error class, and the next scheduled attempt. If a message is repeatedly failing, move it to a dead-letter or quarantine path with enough context for investigation. A dead-letter queue is not a repair process by itself; someone or some system must decide whether to fix, discard, or replay the item.
Retries should respect the dependency's own limits. If a provider returns a retry-after instruction, honor it where appropriate. Apply concurrency limits so recovery does not overload a dependency that is already struggling. A circuit breaker or admission limit can protect the system while still keeping queued work available for later.
Control queue pressure
Queue depth is useful, but message age often shows whether users are waiting too long. Monitor incoming rate, completion rate, oldest-message age, processing duration, retry rate, and worker saturation. A queue that grows continuously means the system is accepting work faster than it can complete it.
Set capacity and back-pressure policies. The service can reject new work, defer acceptance, reduce optional workload, or expand workers up to a safe limit. Unlimited queue growth converts a short overload into delayed failure, memory pressure, or storage exhaustion. Make overload behavior part of the product contract rather than letting it emerge accidentally.
Bound the time and resources each task can consume. A task timeout should be compatible with the broker's visibility or lease timeout. If a task can exceed its lease, a second worker may begin the same work while the first is still running. Renew leases carefully or split long operations into smaller resumable steps.
Treat shutdown and deployment as normal events
A worker should stop taking new work during shutdown, finish or safely release its current task, and exit within the orchestrator's termination window. If the process disappears without acknowledgement, the broker should eventually redeliver the message. Test that lifecycle rather than relying on a graceful path that has never been interrupted.
Rolling deployments can run two versions at once. Messages produced by the new version may be read by an older worker during the rollout. Use version-tolerant schemas, additive changes, and staged migrations. Keep message contracts explicit and avoid coupling every task to the exact current database representation.
Observe work from acceptance to outcome
Assign a task ID at acceptance and carry it through logs, metrics, and traces. Record accepted time, start time, completion time, result, and retry number. Avoid putting sensitive task payloads into logs; log identifiers and safe summaries. This creates a path to answer which work is delayed and whether a specific task completed.
A dashboard can show arrival and completion rates, queue age, active workers, failures by category, and dead-letter volume. Alerts should trigger on sustained conditions that need action, such as oldest-message age exceeding a service objective or a rising permanent-failure rate. A transient depth spike may be normal if the queue drains quickly.
Test the boundaries deliberately
Inject a worker crash before processing, after a side effect, and before acknowledgement. Simulate a slow dependency, a repeated delivery, a malformed message, a full queue, and a deployment that overlaps worker versions. Verify the database or external system ends in the intended state and that operators can find the cause.
Reliable background work comes from explicit guarantees and controlled failure behavior. Idempotency limits the damage from duplicates; bounded retries protect dependencies; back-pressure protects the system; and useful telemetry shows whether the queue is making progress. The design is complete only when a task can fail visibly and still have a deliberate path to recovery.