Design idempotent AtlasMart command processing with scoped request IDs, response replay, retention windows, payload mismatch checks, and explicit crash-window analysis instead of vague exactly-once claims.

Idempotent Commands, Deduplication Windows, Request IDs, and Exactly-Once Illusions

Retries are inevitable; duplicate business effects are not. This lesson makes ambiguous payment outcomes safe with durable deduplication, then shows where the guarantee ends when state expires or side effects cross systems.

Advanced115–150 minutesIdempotency/dedup labPython 3.13+ · standard libraryVendor-neutral · free/local mandatory pathLast reviewed: August 2026
01

Design an idempotent command handler using a scoped request identity and durable response replay.

02

Explain why deduplication retention creates a bounded guarantee rather than timeless exactly-once execution.

03

Diagnose marker-first and side-effect-first crash windows when dedup state and business effects are stored separately.

04

Prefer explicit at-least-once plus idempotent processing/reconciliation contracts over vague end-to-end exactly-once claims.

1. Ambiguous failure is why idempotency matters

AtlasMart sends CapturePayment(req-77, 50). The payment service commits the charge, but the network connection closes before the client receives the result. Retrying is necessary for reliability—but unsafe if the service cannot recognize that the retry represents the same logical command.

An operation is idempotent when applying the same intended operation multiple times has the same intended effect as applying it once. HTTP defines several request methods as idempotent, but application-level POST-like commands often need an explicit idempotency key or request ID. The server scopes that identity to the tenant/operation and stores enough durable result state to replay the original outcome.

2. The dedup record is part of the correctness contract

A useful dedup record includes request ID, tenant/account scope, operation type, a fingerprint or normalized payload identity, status, durable business-result identifier, response summary, creation time, and expiry/retention policy. Reusing one request ID with a different payload must be rejected; otherwise a retry identity becomes an accidental alias for a different command.

Field Why it exists Failure if omitted
tenant + operation scope prevents cross-tenant/cross-command collisions one customer can replay or block another logical command
request/event ID recognizes retries duplicate side effects
payload fingerprint/identity detects key reuse with different intent wrong response replayed for new command
result/status replays outcome without redoing side effect retry must rediscover/redo uncertain work
retention/expiry bounds storage without policy state grows forever; with too-short policy late retries can repeat

3. Retention windows mean the guarantee is bounded

If AtlasMart retains dedup entries for 60 minutes, a retry at minute 10 can replay the original result. A retry after the entry is deleted may execute again unless another durable business identity—such as unique payment intent ID or order ID—still prevents duplication. Therefore “idempotent for 60 minutes” and “this business action can never happen twice” are different promises.

The retention horizon must exceed realistic retry/replay delays, including queue backlogs, mobile offline periods, and disaster recovery replays. Longer retention costs storage/index space and may raise privacy/governance questions.

4. Side-effect ordering creates two classic crash windows

If a service writes the dedup marker before the external side effect and crashes, a retry may see “already processed” and suppress work that never happened. If it performs the side effect before writing the marker and crashes, a retry may repeat the side effect. The cleanest solution is to place the business mutation and dedup record in one local atomic transaction when they share a datastore.

When the side effect crosses systems, use a protocol such as transactional outbox/inbox, downstream idempotency, durable workflow state, and reconciliation. “Exactly once” cannot be inferred from one component's dedup table because failures can occur between components.

5. AtlasMart lab: response replay, mismatch rejection, and expiry

Mandatory lab environment

Python 3.13+ standard library only. The generated lab was verified with Python 3.13.5. No database server, Docker, cloud account, paid feature, credential, firewall change, clock manipulation, process killing, or destructive failure injection is required. All conflicts, partitions, retries, and failures are deterministic in-memory simulations.

The lab charges twice under a naive retry, then introduces a dedup store that atomically records the business effect and response inside the simulator. It replays the original response, rejects reuse with a different amount, and demonstrates that a retry after expiry can execute again.

python · AtlasMart deterministic simulation
print("AMBIGUOUS PAYMENT TIMEOUT")
ledger = []

def naive_capture(request_id, amount):
    ledger.append((request_id, amount))
    return {"status": "captured", "amount": amount}

naive_capture("req-77", 50)
print("first capture committed; response lost")
naive_capture("req-77", 50)
print("naive ledger:", ledger, "charged total:", sum(x[1] for x in ledger))

print("\nIDEMPOTENT HANDLER WITH RESPONSE REPLAY")
ledger = []
dedup = {}

def capture(request_id, amount, now, ttl):
    row = dedup.get(request_id)
    if row and now <= row["expires_at"]:
        if row["amount"] != amount:
            raise ValueError("same request ID reused with different payload")
        return {**row["response"], "replayed": True}
    response = {"status": "captured", "amount": amount, "ledger_id": len(ledger) + 1}
    # Simulator treats business mutation + dedup record as one atomic durable boundary.
    ledger.append((request_id, amount))
    dedup[request_id] = {"amount": amount, "response": response, "expires_at": now + ttl}
    return {**response, "replayed": False}

print("first:", capture("req-77", 50, now=100, ttl=60))
print("retry:", capture("req-77", 50, now=110, ttl=60))
print("ledger:", ledger, "charged total:", sum(x[1] for x in ledger))

print("\nPAYLOAD MISMATCH IS REJECTED")
try:
    capture("req-77", 70, now=120, ttl=60)
except ValueError as e:
    print("rejected:", e)

print("\nDEDUP WINDOW IS A BOUNDED GUARANTEE")
print("after expiry:", capture("req-77", 50, now=170, ttl=60))
print("ledger now:", ledger, "charged total:", sum(x[1] for x in ledger))
print("late retry can repeat after retention unless business identity supplies a stronger invariant")

print("\nSIDE-EFFECT ORDERING FAILURE MODES")
print("marker-first then crash before charge -> retry may suppress a missing charge")
print("charge-first then crash before marker -> retry may duplicate a charge")
print("safe design needs one atomic boundary where possible, or an outbox/reconciliation protocol across boundaries")
Expected evidence

The naive path charges 100 total for one logical 50-unit payment. The idempotent path charges 50 and marks the retry as replayed. Payload mismatch is rejected. After the 60-unit retention interval expires, the same request ID executes again, proving the bounded nature of the dedup guarantee.

6. Exactly-once is a claim about a whole end-to-end path

Messaging products may offer exactly-once features within carefully defined boundaries, but an application side effect can still be repeated outside that boundary. A consumer can atomically commit a broker offset yet fail to atomically commit a payment at a third-party provider. A producer can deduplicate broker writes while a downstream database transaction retries. The useful architecture statement is precise: “delivery is at least once; handlers are idempotent by event ID for N days; business rows have a unique operation identity; reconciliation detects divergence.”

RFC 9110 also illustrates why idempotency matters for retry: idempotent request semantics permit repeating a request after communication failure without changing the intended effect. For non-idempotent application commands, an explicit key plus server-side handling provides the analogous application contract.

7. Deliberately wrong approach: random ID per retry

If a client generates a new request ID every time it retries, the server cannot know the calls represent one logical command. The key must identify the business attempt, not the transport attempt. Conversely, reusing one key forever for unrelated commands causes false deduplication. Generate once per logical operation, persist it with client/workflow state, and reuse it only for retries of the same intent.

8. Production judgment and bridge to TTL/ephemeral state

Observe dedup hit rate, key-collision/mismatch rejects, age distribution of retries, entries near expiry, duplicate side-effect incidents, outbox/inbox backlog, and reconciliation discrepancies. Protect request IDs from cross-tenant probing, avoid putting secrets in keys, and rate-limit pathological replay. Test crash points before/after durable mutations and side effects.

Dedup retention introduces an expiry problem: when may state safely disappear, how exact is expiry timing, and what happens under replication or memory pressure? Chapter 19 continues directly with TTL, lazy versus active expiry, eviction, and ephemeral-state reconstruction.

Check your understanding

  1. Why must retries reuse the same request ID?
  2. Why should a server reject the same key with a different payload?
  3. What does a 24-hour dedup retention policy actually guarantee?
  4. Why are marker-first and side-effect-first both risky across systems?
  5. What is a safer architecture statement than “exactly once”?
Review the answers

1. The ID represents one logical business attempt; changing it per transport retry prevents the server from recognizing duplicates.

2. Otherwise a key collision/reuse can replay the wrong result for a different intent.

3. Only that matching retries recognized within the retained window can be deduplicated; later retries need another invariant or can execute again.

4. Marker-first can suppress missing work after a crash; side-effect-first can duplicate work after a crash.

5. State the delivery mode, idempotency identity/window, atomic boundaries, downstream guarantees, and reconciliation path explicitly.

References

Foundational claims use primary research/specifications where practical. Product documentation is used only as a current implementation example and is not required for the mandatory labs.

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.