SaaS development
Design background jobs to survive a retry
Why a SaaS job needs an operation identity, durable state and a plan for partial failure.
By Thomas Watson · · 3 min read
A worker can finish an operation and crash before acknowledging its queue message. When the message returns, the next worker cannot infer from the missing acknowledgement whether the work happened. Design for a repeated attempt, even if duplicates are uncommon.
Give the operation a stable identity
Imagine a SaaS product generating an export. The export request should have an identifier that stays the same across retries. A fresh random identifier on each attempt describes the attempt, not the operation. Store the request and its status durably, with a uniqueness constraint that prevents two records for the same operation.
Keep local changes atomic
If the entire effect is inside one database, a transaction can put the effect and the completion record in the same commit. Concurrent workers still need coordination through the database: a uniqueness constraint or conditional update is stronger than checking a flag and then acting. Two workers can both pass a separate read before either writes.
- Load the operation by its stable identifier.
- If it is already complete, return the stored result.
- Claim the work using a database operation that prevents competing claims.
- Perform the work and record its result.
- Acknowledge the message only after the durable state is updated.
External effects need their own plan
A database transaction cannot roll back an email already sent or an external payment already accepted. Where a provider supports idempotency keys, use a stable operation key and follow that provider’s documented retention and retry behaviour. Where it does not, you need reconciliation or an explicit decision about the risk of duplicates. A local processed flag alone cannot close the gap between the external action and recording its success.
For an export, a deterministic storage key can make repeated uploads replace the same output, provided the input and generation are stable. If the input can change, capture a version or snapshot when creating the export request. Otherwise a retry can quietly produce a different result.
Exercise the failure windows
- Stop a worker immediately before and after its external effect.
- Deliver the same operation to two workers at once.
- Retry after a timeout where the remote result is unknown.
- Check how an expired worker claim becomes recoverable.
- Separate invalid input from temporary failures so permanent errors do not retry forever.
Track the operation identifier in logs alongside an attempt identifier. That gives you a way to reconstruct what happened without confusing several attempts with several user requests.
