Tags: analytics concept

Anonymous and Identified Users

Date: 2026-08-16


Two populations in one dataset, measured differently and behaving differently. Any analysis restricted to identified users is analysis of your best customers, whatever it’s labelled.


What it is

An anonymous user is a device with an identifier and no known person attached. An identified user is one you’ve matched to an account, an email, or an order.

The transition happens at login, registration or purchase, and joining the two halves is Identity Stitching.

They are not comparable populations

The distinction that matters analytically. Identified users are self-selected — they signed in, which means they’ve bought before, intend to buy, or value something behind the login.

                    anonymous        identified
share of traffic       ~85%              ~15%
conversion rate         low              much higher
repeat rate              —               by definition higher
measurement quality     poor             good

That correlation is the trap: the population you can measure best is the population that behaves least like the average. Any “our users do X” claim drawn from identified users is a claim about your loyalists.

For a retail site most sessions never authenticate, so the identified set is a minority — and a flattering one.

What each supports

AnonymousIdentified
Session-level behaviourYesYes
Cross-deviceNoYes
Durable over monthsNo — identifiers churnYes
Lifetime value, cohortsNoYes
Joins to CRM, orders, emailNoYes
Represents your trafficYesNo

Anonymous data is broad and shallow; identified data is narrow and deep. Most useful questions need one or the other specifically, and knowing which is half the analysis.

Where it goes wrong

  • Reporting identified-user metrics as site metrics. “Our conversion rate is 9%” when that’s the logged-in rate and the site rate is 3%
  • Cohort and retention analysis on identified users only, then generalised. Structurally overstates retention, because you’ve excluded everyone who never came back far enough to sign in — Survivorship Bias
  • Comparing periods where the login rate changed. A checkout change that increases account creation makes the identified population larger and less selective, so its conversion rate falls with no behavioural change — a mix shift, not a decline — Simpson’s Paradox
  • Assuming identified means one person. Shared accounts, households and office logins are common in retail — Identity Stitching
  • Forgetting that identification improves over time, so a cohort measured at month one and month six isn’t measuring the same set — Cohort Analysis

Practical handling

  • Report the split. What proportion of sessions, and of conversions, are identified? It’s the context every other number needs
  • State the population on every metric. “Conversion rate (all sessions)” versus “(identified)” — two different numbers with the same name is Metric Drift waiting to happen
  • Use anonymous data for behaviour, identified data for value. Funnels and on-site behaviour from the whole population; lifetime value, retention and cohorts from the identified set, labelled as such
  • Watch the boundary. Users crossing from anonymous to identified mid-session are the ones whose data is most likely to be split across two records — that’s the stitch, and it’s where conversions get orphaned

The privacy dimension

Anonymous is a weaker claim than it sounds. A persistent device identifier plus behavioural history is pseudonymous, not anonymous, and remains personal data under UK GDPR — the identifier alone doesn’t name someone, but it singles them out.

Which means “we only collect anonymous data” is usually inaccurate as a compliance position. See Pseudonymisation and Anonymisation and UK GDPR and PECR for Analytics.