Dinesh’sLearning Lab
← The learning library
systems · intermediate · 13 min read
Lesson 8 of 10 in this path ↗

08 · Retries: a timeout is not a verdict

Model a lost acknowledgement, define idempotent effects, and separate delivery, durable processing and response replay.

Editorial review: · What review means

Stored in this browser only. No account, no sync. Clearing browser data removes your record.

By the end, you should be able to

  • Explain the ambiguous outcome after a timeout
  • Distinguish idempotency from identical responses
  • Design a deduplication boundary and state its limits

Bring with you

  • 07 · Relations, joins and all-or-nothing changes

Listen to this article

Browser / device speech · no paid TTS integration. Voice quality depends on your device.

Choose a local device voice to avoid a remote speech service. This site adds no TTS service, account or API calls.

Checking browser speech support…

Pause saves your segment; resume repeats that short segment. Changing voice or speed pauses playback. Stop resets to the beginning. Progress counts finished text segments, not audio time. Leaving or hiding this page stops or pauses speech.

What gets read aloud?

Reads the article body as it appears when you press Listen. Navigation, controls and closed sections are skipped. Expand a section, then Stop and Listen to include it. Code and equations get brief notices; figures use available labels or captions, not their visual details. This narration does not teach omitted mathematics or replace reading examples on the page.

For better sound at no added site cost, try installed English voices, including enhanced voices offered by your device. We cannot guarantee a best voice on every browser. Use Stop or your device’s audio controls if its speech engine misbehaves.

In this article · 5 sections

The notebook accepts an event, stores it, and returns success. Now lose the reply. The client sees a timeout; the server has already changed state. Retrying is reasonable, but a naive “increment on every arrival” implementation counts twice. Declaring that the transport is reliable does not settle the application-level question: did the requested effect occur?

This lesson uses deterministic, in-memory models. It opens no sockets and measures no network. Its purpose is to make one failure history visible, not to simulate every real distributed-system behavior.

Separate the messages from the effect

A request has a method, target, fields and possibly a body. A response has a status and fields and may carry a representation. The status communicates what the server reports; absence of a response does not prove absence of a committed effect. The failure could happen before the server receives the request, during processing, or after a successful commit but before the reply is observed.

RFC 9110 section 9.2.2 defines an idempotent method by its intended server effect under multiple identical requests. It names PUT, DELETE and safe methods as idempotent. The responses may differ, and the server may log each attempt separately. Deleting an already deleted resource can return a different status without violating the intended repeated effect.

The RFC says a client should not automatically retry a non-idempotent request unless it knows the request semantics are idempotent or knows the original request was never applied. A POST endpoint can implement an application-specific idempotency contract, but POST itself does not confer that property. “Safe” is another distinction: safe methods have essentially read-only requested semantics; idempotent DELETE is not thereby safe.

Execute the failure the explanation claims

# Model 1: a reply is lost AFTER the first effect.
naive_total = 0
 
def naive_receive(amount):
    global naive_total
    naive_total += amount
    return naive_total
 
lost_reply = naive_receive(5)  # server committed; client never observes this value
retry_reply = naive_receive(5)
assert naive_total == 10
assert retry_reply == 10
 
# Model 2: sequential in-memory processing with a stable request identity.
records = {}
total = 0
 
def receive(key, amount):
    global total
    if key in records:
        original_amount, original_response = records[key]
        if original_amount != amount:
            raise ValueError("same key with a different request")
        return original_response.copy()
    total += amount
    response = {"accepted_amount": amount}
    records[key] = (amount, response.copy())
    return response
 
lost_reply = receive("e1", 5)
retry_reply = receive("e1", 5)
assert total == 5
assert retry_reply == {"accepted_amount": 5}
receive("e2", 5)  # same amount, different event: must count separately
assert total == 10
try:
    receive("e1", 9)
except ValueError:
    pass
else:
    raise AssertionError("key reuse conflict was silently accepted")
assert total == 10

This model actually repeats a request after the first effect, unlike a “retry” example that only retries attempts rejected before processing. Deduplication cannot demonstrate value if no duplicate effect is ever attempted. The test checks three distinct cases: same key/same content, different key/same content, and same key/different content.

The notebook's event ID is a logical identity chosen before the first attempt and reused for retries. It is not a fresh random number on every attempt. In a multi-user API, the key would normally be scoped to an authenticated caller and operation so one user's key cannot collide with another's. A key is not an authorization credential.

Find the hole in the model

The second model still has a dangerous interval between changing total and writing records[key]. A crash there loses the deduplication record after applying the effect. Concurrent callers can also both observe the key as absent. The example's sequential execution and in-memory state deliberately exclude those failures; it is not a production exactly-once implementation.

If the effect and the request record live in one transactional database, store them in one transaction with a uniqueness constraint on the scoped key. A retry checks the stored canonical request, then returns the stored response or a defined conflict. The SQLite transaction contract supports grouping changes within one database; it does not group a database change and a remote charge into one atomic operation. An outbox/inbox or provider-supported idempotency protocol may be necessary for external effects, with its own retention and recovery rules.

The capstone implements the local transactional boundary and injects failures before commit. It does not claim a distributed transaction with another service.

A retry policy spends resources

Unlimited retries can turn a transient failure into sustained overload. Bound attempts and elapsed time, distinguish invalid requests from transient failures, and introduce delay rather than an immediate tight loop. Randomized delay can reduce synchronized retries; it does not make a non-idempotent effect safe. A server's Retry-After instruction and the endpoint's documented semantics matter more than a generic list of “retryable” status codes.

A deduplication record also needs a retention policy. Delete it after one hour, then replay a request after two hours, and the old identity can look new. The guarantee is limited to the retention window and the stable scope of the key. Keeping records forever has a storage and privacy cost. A cryptographic hash can compact a request fingerprint, but hashing a string is not a substitute for specifying canonicalization: field ordering, defaults and relevant identity fields need a contract.

Exercises

  1. Draw the timeline for “effect committed, acknowledgement lost, retry arrives.” Mark what the client knows at each point.
  2. Move records[key] before the effect. Did that remove the crash gap?
  3. Should a duplicate response include the current global total or the original result?
  4. A GET handler increments a business counter as its primary action. Is its method label enough to make the action safe?
Answer sketches
  1. The server knows its commit happened; after losing the reply the client cannot distinguish that history from a pre-processing failure. The retry needs a stable identity or an operation whose repeated intended effect is harmless.
  2. No: a crash after recording but before applying the effect now causes the retry to skip a missing effect. Ordering two independently committed writes changes which failure is possible; it does not create atomicity.
  3. Define the API. Replaying a stored response is useful for a stable operation result, but returning current state can also be valid. Neither automatically changes whether the underlying effect is idempotent.
  4. No. Safe method semantics describe requested behavior, not a magic protection applied to any handler named GET. Incidental logging is different from requesting a state-changing business operation.

Exit artifact: a failure timeline, a same-key conflict test, and a written boundary around the guarantee. No network fault-injection, transport verification, payment processing or concurrency correctness is claimed here.

Next: Bounded work and resources.

Pause / Recall / Apply

Can you explain it without the page?

Close the example. Reconstruct the core idea, then change one assumption. Mark complete when you’re ready; you can always undo it.

Stored in this browser only. No account, no sync. Clearing browser data removes your record.