Identity Stitching
Date: 2026-08-16
Joining anonymous activity to a person once they identify themselves. Every number about people — users, retention, lifetime value, pre-login attribution — is downstream of how well this works, and it usually works worse than assumed.
Identity stitching is linking the anonymous IDs a person accumulated before identifying themselves to their user ID, so their earlier activity is attributed to the same person.
The two identifiers
- Anonymous ID — generated on first visit, stored on the device (cookie or local storage), scoped to one browser on one machine. Sometimes called the device ID or distinct ID
- User ID — your application’s identifier for a person, known only once they log in, register or complete a purchase
Stitching is the operation that says these are the same person, and then decides what to do with the events that predate the knowing.
The trigger is an explicit call at the moment of identification — identify(userId) or alias(anonymousId, userId), depending on the SDK. Everything after that point carries both identifiers. The question is what happens to everything before it.
AS COLLECTED THE PROBLEM
event anonymous_id user_id name time
e1 a_77 – page_view 09:02 ← who is this?
e2 a_77 – add_to_cart 09:06 ← who is this?
e3 a_77 u_512 login 09:11 ← now we know
e4 a_77 u_512 purchase 09:14
The stitch is one row in a mapping table, written at e3:
anonymous_id user_id first_linked
a_77 u_512 09:11
Applying it resolves the first two rows retrospectively:
RESOLVED
event anonymous_id user_id name time
e1 a_77 u_512 page_view 09:02 ← recovered
e2 a_77 u_512 add_to_cart 09:06 ← recovered
e3 a_77 u_512 login 09:11
e4 a_77 u_512 purchase 09:14
That recovery is the entire operation. Everything below is about when it’s applied, what it costs, and how it goes wrong.
Two strategies
| Retroactive rewrite | Mapping table | |
|---|---|---|
| What happens | Historical anonymous events are updated to carry the user ID | An anonymous_id → user_id table is maintained; events are joined to it at query time |
| Query cost | Cheap — the join is already done | Every user-level query pays for the join |
| Reversible | No | Yes, the raw stream is untouched |
| Handles a wrong merge | Requires reprocessing, if it’s possible at all | Delete the row |
| Handles a later merge | Needs another rewrite pass | Automatic — old queries re-run correctly |
| Typical of | Product analytics SDKs | Warehouse-First Analytics |
The rewrite is the one most tools do by default, and it’s the lossy one. It destroys the record of what was known at the time, which matters more than it sounds: it means you can never reconstruct what a report said last quarter, and a bad merge is permanent damage rather than a bad join.
Deterministic vs probabilistic
Deterministic stitching joins on a shared identifier that both sessions definitely carried — a login, a hashed email in a link, an order confirmation. If the identifier matches, the join is correct by construction.
Probabilistic stitching infers the join from circumstantial signals: same IP address, same user agent, similar timing, same location. It’s a statistical guess dressed as a fact — no shared identifier ever existed.
In plain terms: deterministic means the two sessions carried the same badge. Probabilistic means they looked similar enough that a model decided to call them the same person, and it is sometimes wrong in both directions — splitting one person into two, and merging two people into one.
The error rates are rarely published and never verifiable against ground truth, since checking would require knowing the answer you’re trying to infer. Treat vendor accuracy figures for probabilistic matching as marketing. Also note that inferring a link between two devices is a stronger processing operation than recording either alone, with consequences under UK GDPR and PECR for Analytics.
What it does to user counts
One person: work laptop (Chrome), home laptop (Chrome and Safari), phone, tablet.
- No stitching — five anonymous IDs, counted as 5 users
- Login on phone and home Chrome only — those two join to one; the rest stay separate: 1 + 3 = 4 users
- Login everywhere — 1 user
The same activity produces a headline user count varying by 5×, and there is no configuration that makes it “correct” — only configurations with different biases. This is the root of most of User Counting.
Second-order consequence worth stating: stitching improves over the lifetime of an account, because logins accumulate. So a cohort measured in month one looks like more users doing less each than the same cohort measured in month six. That’s a measurement artefact that reads exactly like an engagement improvement. See Cohort Analysis.
Failure modes
-
Shared devices — the family tablet, the office machine, the in-store kiosk. Two people log in; the anonymous ID belongs to whoever came first, and both sets of activity merge into one profile. Common enough on retail sites to matter, and invisible in aggregate
-
Identity collapse — the nasty one. If merges are transitive, a shop-floor tablet accumulates links and the chain propagates:
mapping table transitive resolution a_11 → u_1 (Monday) u_1 ─┐ a_11 → u_2 (Tuesday) u_2 ─┼─ all one profile a_11 → u_3 (Wednesday) u_3 ─┘The result is a single profile with implausible activity. Because it’s one profile, it never shows up as an anomaly in user counts — it shows up as a “power user” nobody questions, and it quietly skews every average it’s included in. Guard by capping merge chains and alerting on profiles exceeding a plausible event count
-
Irreversibility — most merges cannot be undone once written. Under rewrite semantics, a wrong merge is permanent
-
Anonymous ID churn — cookie clearing, private browsing, and browser-imposed lifetime caps on client-set storage all reset the anonymous ID. The same person on the same device becomes several apparent users over a month, and the churn is heavier on Safari than Chrome, so it varies systematically by device type [CHECK: current ITP cookie lifetime cap for JS-set cookies]. See Browser Privacy Restrictions
-
Logged-out majority — on most ecommerce sites, most sessions never authenticate. Stitching is a minority operation, and the stitched population is a biased sample: more loyal, more likely to convert, more likely to be measured properly. Analysis restricted to stitched users is analysis of your best customers, whatever it’s labelled
-
Server events without an anonymous ID — a backend
purchaseevent sent with the user ID but no anonymous ID cannot join to the browsing session that preceded it. Very common, and it silently orphans the whole pre-purchase path from the conversion
Why it constrains attribution
A conversion is credited to touchpoints that happened before it. Those touchpoints were almost always anonymous; the conversion is almost always identified. The credit only reaches them if the stitch held.
Where it fails, the path collapses to whatever happened after the identifier existed — which is typically the final session — and that credit lands on direct or on branded search. So broken stitching doesn’t produce obviously missing data. It produces a plausible, wrong, systematically last-touch-biased picture that flatters exactly the channels which need it least. See Attribution Models and Direct Traffic and Lost Referrers.
The practical rule: before arguing about attribution models, check what proportion of conversions are stitched to a pre-login session at all. If it’s low, the model choice is a rounding error next to the identity problem.