My refund handler checked the ledger before paying. It still paid twice.

# python# testing# api# ai
My refund handler checked the ledger before paying. It still paid twice.Jigon Yoo

A refund handler commits the refund, then loses its reply. Maybe the HTTP response timed out. Maybe...

A refund handler commits the refund, then loses its reply. Maybe the HTTP response timed out. Maybe the worker died before it acknowledged the queue message. Maybe it was an AI agent's tool call, and the agent saw a timeout and called the tool again with the same arguments. The sender did nothing wrong by retrying. The question is what your handler does with the second delivery.

The usual fix is to look first: if the ledger already has a refund for this order, return early. I wanted to see how far that fix goes, so I wrote a small pytest plugin that delivers the same event again in four different ways and counts how many times the effect landed. It is public and runs offline: github.com/jigonyoo/replay-twice.

Four ways to deliver the same event twice

Scenario What happens
duplicate the same event arrives twice, one after the other
timeout_retry the handler finishes, its reply is replaced by a timeout, and a scripted agent retries the same tool call with the same arguments
crash_before_ack the handler finishes, the acknowledgement is lost, and the event is redelivered: the queue's or webhook sender's view (the lost ack is simulated with an exception; no process is killed)
concurrent eight deliveries at once: threads released together by a barrier, or eight tasks on one event loop for async def handlers

The verdict does not come from the handler's return value. You give the drill an effects(key) function that counts what was actually committed, such as rows in the refund ledger for that order, and a trial passes only if the count is exactly one.

Four handlers, one ledger

The repository ships SQLite versions of a refund tool and a payment.succeeded webhook, each written four ways:

  • naive writes the refund every time.
  • check_then_act reads the ledger, and writes only if nothing is there.
  • guarded_raises claims the order id in a table with a unique key, in the same transaction as the refund. A second claim raises an integrity error.
  • guarded makes the same claim with INSERT OR IGNORE and returns “already applied” when the claim was already taken.

The result

Refund flow, three runs. The sequential scenarios are one trial per run; the concurrent scenario is 200 trials per run, each with eight overlapping deliveries. All three runs gave the same verdict in every cell; how many times the refund landed inside a race varied from run to run.

Handler duplicate timeout_retry crash_before_ack concurrent (200 trials)
naive paid twice paid twice paid twice 200 paid twice or more
check_then_act once once once 200 paid twice or more
guarded_raises once, but raised once, but raised once, but raised once, raised in all 200
guarded once once once once in all 200

The webhook flow and an async version of the refund tool gave the same verdicts, cell for cell. Measured with replay-twice 0.2.0 on Linux, Python 3.11, eight threads, seed 0.

Why looking first is not enough

Reading before writing handles every retry that arrives after the first delivery has finished. That covers three of the four scenarios, and it is why the pattern survives code review: test it by calling the handler twice and it passes.

It fails when two deliveries overlap. Both read an empty ledger, both decide nothing has been paid, both write. A webhook sender that retries quickly, two workers pulling the same message, or an agent loop that fires a tool call while the previous one is still running will all produce that overlap.

The read here takes no lock, and adding one is harder than it sounds. At the default isolation level of most databases, wrapping the read and the write in one transaction still lets both deliveries see an empty ledger, and SELECT … FOR UPDATE cannot lock a row that does not exist yet. A unique constraint on the key is the guard the database will actually enforce.

One honest caveat about the 200 out of 200. The example handler sleeps for one millisecond between its read and its write, and the drill releases all eight deliveries at the same instant; both widen the gap on purpose, so this is not a failure rate you should expect in production. It shows that the gap exists, and that a test which only calls the handler twice in sequence will never find it.

Once is not the whole answer either

guarded_raises never paid twice. It still has a bug. Every redelivery hits the unique key and raises, which in a web handler usually becomes a 500, and the sender, seeing an error, retries again. The money moved once; the retries do not stop.

So the drill reports handler errors separately from the effect count. By default it passes with a warning; with redelivery_errors="fail" it fails. The difference between guarded_raises and guarded is one statement and an early return: claim the key with INSERT OR IGNORE and return “already applied” instead of raising. A production handler would usually store the first result and return that.

with connection:
    claimed = connection.execute(
        "INSERT OR IGNORE INTO claims VALUES (?, ?)", (kind, key)
    ).rowcount
    if not claimed:
        return "already_applied"
    connection.execute("INSERT INTO ledger VALUES (?, ?, ?)", (kind, key, amount))
return "applied"
Enter fullscreen mode Exit fullscreen mode

Here the claim and the ledger row commit in one transaction, so a crash between them cannot leave a claim with no refund behind it. A real refund is usually a call to a payment API, which cannot join your database transaction. The same idea still applies: derive the key from the order rather than generating a new one per attempt, and pass it as the provider's idempotency key, so a retry gets the first refund back instead of creating a second. Providers keep these keys for a limited time, so check how long. The local claim then needs a pending and a done state: a crash between the claim and the API call should be retried with the same key, not left as a claim with no refund. The drill does not care how you do it; it only counts what landed.

What this does not test

Everything above runs in one process against SQLite. The drill does not cover several servers racing each other, your real database's isolation level, distributed locks, broker delivery policies, effects committed after the handler returns, duplicates on the payment provider's side, real process crashes and power loss, network partitions, effects your observer cannot see, or a handler that never returns (use your test runner's timeout). It also cannot help if the key itself is wrong: two partial refunds on one order need two identities. A pass means the schedules that ran were safe, not that every schedule is.

Try it on your handler

pip install "replay-twice @ git+https://github.com/jigonyoo/replay-twice"
Enter fullscreen mode Exit fullscreen mode
from uuid import uuid4
from replay_twice import ReplayCase, SCENARIOS

def test_refund_happens_once(replay, db):
    def fresh_case():
        order_id = uuid4().hex
        return ReplayCase(
            handler=refund_tool,
            effects=lambda key: db.count_refunds(key),
            event={"order_id": order_id, "amount": 800},
            key=order_id,
        )
    replay(case_factory=fresh_case, scenarios=SCENARIOS)
Enter fullscreen mode Exit fullscreen mode

Point it at a test database, not production: the handler is really called. By default the concurrent drill runs once; add --replay-repeats=200 to run it as many times as the table above. async def handlers work too. The table above reproduces with replay-twice report in about four minutes, and the raw observations are committed in the repository.

Written with AI assistance. Every number comes from the committed reports in the repository.