Tags: analytics concept

Identity Stitching

Date: 2026-08-16


Joining anonymous activity to a person once they identify themselves. Every number about people — users, retention, lifetime value, pre-login attribution — is downstream of how well this works, and it usually works worse than assumed.


Identity stitching is linking the anonymous IDs a person accumulated before identifying themselves to their user ID, so their earlier activity is attributed to the same person.

The two identifiers

  • Anonymous ID — generated on first visit, stored on the device (cookie or local storage), scoped to one browser on one machine. Sometimes called the device ID or distinct ID
  • User ID — your application’s identifier for a person, known only once they log in, register or complete a purchase

Stitching is the operation that says these are the same person, and then decides what to do with the events that predate the knowing.

The trigger is an explicit call at the moment of identification — identify(userId) or alias(anonymousId, userId), depending on the SDK. Everything after that point carries both identifiers. The question is what happens to everything before it.

AS COLLECTED                                    THE PROBLEM

event  anonymous_id  user_id  name         time
e1     a_77          –        page_view    09:02   ← who is this?
e2     a_77          –        add_to_cart  09:06   ← who is this?
e3     a_77          u_512    login        09:11   ← now we know
e4     a_77          u_512    purchase     09:14

The stitch is one row in a mapping table, written at e3:

anonymous_id  user_id  first_linked
a_77          u_512    09:11

Applying it resolves the first two rows retrospectively:

RESOLVED

event  anonymous_id  user_id  name         time
e1     a_77          u_512    page_view    09:02   ← recovered
e2     a_77          u_512    add_to_cart  09:06   ← recovered
e3     a_77          u_512    login        09:11
e4     a_77          u_512    purchase     09:14

That recovery is the entire operation. Everything below is about when it’s applied, what it costs, and how it goes wrong.

Two strategies

Retroactive rewriteMapping table
What happensHistorical anonymous events are updated to carry the user IDAn anonymous_id → user_id table is maintained; events are joined to it at query time
Query costCheap — the join is already doneEvery user-level query pays for the join
ReversibleNoYes, the raw stream is untouched
Handles a wrong mergeRequires reprocessing, if it’s possible at allDelete the row
Handles a later mergeNeeds another rewrite passAutomatic — old queries re-run correctly
Typical ofProduct analytics SDKsWarehouse-First Analytics

The rewrite is the one most tools do by default, and it’s the lossy one. It destroys the record of what was known at the time, which matters more than it sounds: it means you can never reconstruct what a report said last quarter, and a bad merge is permanent damage rather than a bad join.

Deterministic vs probabilistic

Deterministic stitching joins on a shared identifier that both sessions definitely carried — a login, a hashed email in a link, an order confirmation. If the identifier matches, the join is correct by construction.

Probabilistic stitching infers the join from circumstantial signals: same IP address, same user agent, similar timing, same location. It’s a statistical guess dressed as a fact — no shared identifier ever existed.

In plain terms: deterministic means the two sessions carried the same badge. Probabilistic means they looked similar enough that a model decided to call them the same person, and it is sometimes wrong in both directions — splitting one person into two, and merging two people into one.

The error rates are rarely published and never verifiable against ground truth, since checking would require knowing the answer you’re trying to infer. Treat vendor accuracy figures for probabilistic matching as marketing. Also note that inferring a link between two devices is a stronger processing operation than recording either alone, with consequences under UK GDPR and PECR for Analytics.

What it does to user counts

One person: work laptop (Chrome), home laptop (Chrome and Safari), phone, tablet.

  • No stitching — five anonymous IDs, counted as 5 users
  • Login on phone and home Chrome only — those two join to one; the rest stay separate: 1 + 3 = 4 users
  • Login everywhere — 1 user

The same activity produces a headline user count varying by 5×, and there is no configuration that makes it “correct” — only configurations with different biases. This is the root of most of User Counting.

Second-order consequence worth stating: stitching improves over the lifetime of an account, because logins accumulate. So a cohort measured in month one looks like more users doing less each than the same cohort measured in month six. That’s a measurement artefact that reads exactly like an engagement improvement. See Cohort Analysis.

Failure modes

  • Shared devices — the family tablet, the office machine, the in-store kiosk. Two people log in; the anonymous ID belongs to whoever came first, and both sets of activity merge into one profile. Common enough on retail sites to matter, and invisible in aggregate

  • Identity collapse — the nasty one. If merges are transitive, a shop-floor tablet accumulates links and the chain propagates:

    mapping table                    transitive resolution
    a_11 → u_1   (Monday)            u_1 ─┐
    a_11 → u_2   (Tuesday)           u_2 ─┼─ all one profile
    a_11 → u_3   (Wednesday)         u_3 ─┘
    

    The result is a single profile with implausible activity. Because it’s one profile, it never shows up as an anomaly in user counts — it shows up as a “power user” nobody questions, and it quietly skews every average it’s included in. Guard by capping merge chains and alerting on profiles exceeding a plausible event count

  • Irreversibility — most merges cannot be undone once written. Under rewrite semantics, a wrong merge is permanent

  • Anonymous ID churn — cookie clearing, private browsing, and browser-imposed lifetime caps on client-set storage all reset the anonymous ID. The same person on the same device becomes several apparent users over a month, and the churn is heavier on Safari than Chrome, so it varies systematically by device type [CHECK: current ITP cookie lifetime cap for JS-set cookies]. See Browser Privacy Restrictions

  • Logged-out majority — on most ecommerce sites, most sessions never authenticate. Stitching is a minority operation, and the stitched population is a biased sample: more loyal, more likely to convert, more likely to be measured properly. Analysis restricted to stitched users is analysis of your best customers, whatever it’s labelled

  • Server events without an anonymous ID — a backend purchase event sent with the user ID but no anonymous ID cannot join to the browsing session that preceded it. Very common, and it silently orphans the whole pre-purchase path from the conversion

Why it constrains attribution

A conversion is credited to touchpoints that happened before it. Those touchpoints were almost always anonymous; the conversion is almost always identified. The credit only reaches them if the stitch held.

Where it fails, the path collapses to whatever happened after the identifier existed — which is typically the final session — and that credit lands on direct or on branded search. So broken stitching doesn’t produce obviously missing data. It produces a plausible, wrong, systematically last-touch-biased picture that flatters exactly the channels which need it least. See Attribution Models and Direct Traffic and Lost Referrers.

The practical rule: before arguing about attribution models, check what proportion of conversions are stitched to a pre-login session at all. If it’s low, the model choice is a rounding error next to the identity problem.