OneRuby.devAN ENGINEERING NOTEBOOK

ruby · 5 min read

A unique key doesn't make a Sidekiq job retry-safe

A Ruby experiment shows why a unique payment record can hide unfinished work, and how retries can resume it with one stable operation key.

A payment job creates a pending record, calls a gateway, then marks the record completed. There is a unique index on its operation key. The job rescues a duplicate-record error and returns successfully.

Now interrupt it just after creating the record. On retry, the unique index rejects the insert. The rescue returns. Nobody has charged anything, yet the job has stopped trying.

Two places to stop the job

The accompanying Ruby lab has a store, a fake gateway and two job implementations. Each retry gets a new job object. The store and gateway remain in memory, representing state that survives the failed attempt. Everything runs sequentially in one process.

The broken implementation preserves the original control flow:

Ruby
def perform(operation_id, amount_cents, crash_at: nil)
row = @store.create!(operation_id, amount_cents)
charge_and_complete(row, crash_at)
rescue DuplicateOperation
:ignored
end

charge_and_complete can raise an exception before the gateway call or after the gateway returns but before local completion. These exceptions give us two distinct states:

InterruptionLocal recordGateway effectBroken retry
Before the gateway callPendingNoneReturns :ignored
After the gateway callPendingOne chargeReturns :ignored

In both cases, the local row exists. In neither case does its existence mean the workflow finished. A uniqueness constraint can prevent another row with that key while leaving the first row unresolved.

The second case is particularly awkward: the external effect happened, but the local record has no gateway ID. Retrying with a fresh provider key could create another charge. Refusing to retry leaves the result unresolved.

Give the operation an identity that survives the attempt

The lab uses order-42-payment-1 as the business operation ID. The store records its amount and a provider key derived from that ID. Every attempt loads those same values.

There can be several job executions for one operation. A new payment operation gets a different identity. Re-enqueuing the same operation should keep its existing identity; generating a new random key each time would defeat that relationship. A random key generated once and persisted would also work. If you move this operation into Sidekiq, pass its ID as a simple string and load its stored state inside the job. Sidekiq serializes arguments as JSON; a model object is not a substitute for the operation ID.

The repaired method is short:

Ruby
def perform(operation_id, amount_cents, crash_at: nil)
row = @store.resume_or_create!(operation_id, amount_cents)
return :already_completed if row[:status] == :completed
charge_and_complete(row, crash_at)
end

resume_or_create! returns the existing record when it finds one. It rejects a changed amount, because the caller cannot silently turn an old operation into a different request. A completed record stops the job. A pending record continues to the gateway with its stored key.

The fake gateway remembers the result for each key. After an interruption following the charge, the retry makes a second request and receives the same result. The counters read two requests, one charge. The job can then save the gateway ID and mark the operation completed.

For the earlier interruption, the gateway has no stored result yet. The retry makes the first charge and completes the record.

Run the counterexample

Download example.rb into an empty directory. The recorded environment is Ruby 3.3.2 with Minitest 6.0.6. The file activates that exact Minitest version; if it is missing, install it with gem install minitest --version 6.0.6 --no-document using the same Ruby environment. Then run:

Terminal
ruby example.rb --seed 42

The local run completed eight tests and 36 assertions, with no failures, errors or skips. No database, network connection or payment account is used.

Two passing tests deliberately demonstrate the broken behavior. They assert that retrying leaves the record pending, with either zero charges or one charge depending on the interruption point. The repaired-job tests assert completion in both cases, including reuse of the gateway result. Other tests cover completed retries, changed amounts and distinct operations.

A useful next experiment is to remove the completed-record shortcut. The provider can still prevent another charge, but the test for completed retries should fail. Avoiding another charge and avoiding another gateway request are separate properties, and the test checks both.

What this experiment leaves open

The store is a hash. Its lookup-or-create method is not an atomic database upsert, and two simultaneous workers are outside this experiment. Replacing a job object also does not test what survives an actual process crash. A production implementation needs durable state and a concurrency design you can test against the chosen database.

The fake gateway retains successful results forever. Real providers have additional rules. Stripe's documentation, for example, describes parameter matching, stored error responses and key pruning after at least 24 hours. An operation that stays unresolved beyond the provider's retention window needs a reconciliation strategy before another charge is attempted.

Queue reliability is another boundary. Sidekiq recommends jobs that tolerate repeated execution, while its reliability documentation explains how an in-progress job can be lost with basic fetching. This lab exercises application retry logic; it neither starts Sidekiq nor proves delivery guarantees.

When reviewing a job that calls another service, stop it mentally after each state change. Write down what the next attempt can actually observe. If it sees only a pending record, it still has work to resolve.

Found a mistake or tried a different approach?

Send Alex a note ↗