Idempotency and Deduplication
Date: 2026-08-16
Networks lose acknowledgements, so anything with a retry will deliver some events twice. An identifier generated at the source, deduplicated at the destination, is the only mechanism that fixes it — and it has to be generated where the event happened, not where it arrived.
What it is
An idempotent operation produces the same result whether applied once or many times. Deduplication is the mechanism achieving that for events: recognising an arrival as one you’ve already recorded, and discarding it.
Why duplicates are inevitable
Not a bug — a property of unreliable networks.
client sends event ──▶ server receives it, stores it
◀── acknowledgement LOST in transit
client sees no ack
client retries ──▶ server receives it again
→ two records, one real occurrence
The client cannot distinguish “the server never got it” from “the server got it and the reply was lost”. Its only safe move is to retry, which guarantees occasional duplicates.
This is at-least-once delivery, and it’s what almost every analytics pipeline provides. The alternatives are at-most-once (which drops events) and exactly-once (which is much harder and rarely offered). At-least-once plus deduplication is the standard compromise.
The mechanism
Generate a unique ID at the source, at the moment the event occurs:
{
"event_id": "01J8X2K4M9P3QW7R",
"event": "purchase",
"transaction_id": "ORD-44712",
"value": 48.00
}The destination keeps a record of IDs seen within a window and drops repeats.
Where the ID is generated is the whole thing. Generated at the receiving end, every retry produces a new ID and deduplication does nothing — you’ve added a mechanism that can’t work. It must be created once, by the thing that observed the event, and carried through every retry unchanged.
Two different keys
Worth separating, because they solve different problems:
| Key | Catches | Window |
|---|---|---|
| Event ID | Transport duplicates — retries, double sends | Hours |
Business key — transaction_id | The same order recorded twice by different paths | Indefinite |
The second matters more in retail. An order that fires client-side and from a server webhook produces two events with different event IDs and the same transaction ID. Only the business key catches that — and it’s the most common real cause of duplicated revenue. See Double Counting.
Deduplicate on the transaction ID, not on user plus timestamp proximity. A customer legitimately placing two orders minutes apart is a real scenario in retail, and a proximity rule silently deletes the second one.
The window
Deduplication needs a lookback, and it’s a storage-versus-correctness trade:
- Too short — a retry arriving after a long outage is admitted as new
- Too long — the store of seen IDs grows unboundedly
Hours is typical for event IDs. Business keys should be checked against the full history where the volume allows, because a duplicate order recorded a week later is still a duplicate.
Where it’s needed
- Any client with retry logic. Adding retries without IDs trades lost events for inflated ones — Event Batching and Delivery
- Webhooks. Almost every provider retries on non-2xx, and several deliver duplicates even on success
- Server-side tagging, where one inbound event fans out and any leg may retry — Server-Side Tag Management
- Backfills and reprocessing, where a replay must not duplicate what’s already stored
Checking it
select count(*) - count(distinct transaction_id) as duplicates
from purchases
where date >= current_date - 30Anything above zero is real revenue inflation. Run it as a standing check alongside the order-system reconciliation in Guide - Auditing a Tracking Plan — it’s cheap, and duplicates are one of the few tracking faults that make numbers look better, which is why nobody reports them.