Tags: analytics concept

Pseudonymisation and Anonymisation

Date: 2026-08-17


Two words used interchangeably that mean opposite things legally. Pseudonymised data is still personal data with every obligation attached; anonymous data is outside data protection law entirely. Almost everything called “anonymised” in analytics is pseudonymised, and the gap is where the compliance failures sit.


Pseudonymisation replaces identifiers with a substitute that can still be linked back to the person with extra information; anonymisation removes the possibility of identification altogether, by anyone, by any reasonably likely means.

The distinction

PSEUDONYMISED                        ANONYMISED

alex@example.com                     alex@example.com
      ↓ hash                               ↓ aggregate
7f3a9b2c…                            "1,247 users in SW1 bought
      ↓                                     this week"
still joins to one person                  ↓
still tracks them over time          cannot be traced to anyone
still targets them                   cannot be re-identified
                                     cannot be reversed

STILL PERSONAL DATA                  NOT PERSONAL DATA
  · lawful basis needed                · no lawful basis needed
  · deletion requests apply            · no deletion requests
  · retention limits apply             · keep indefinitely
  · breach notification applies        · no breach exposure

The test isn’t whether the identifier looks like a name. It’s whether the data can single out an individual — directly or by combination with anything else reasonably available. A user ID that follows one person across a year singles them out perfectly, whatever it’s made of.

Why hashing isn’t anonymisation

The most common misunderstanding in analytics, and it’s worth being precise about.

// this is PSEUDONYMISATION, not anonymisation
const id = crypto.createHash('sha256')
  .update(email.trim().toLowerCase())
  .digest('hex');

Three reasons it doesn’t anonymise:

  • It’s deterministic, so the same input always gives the same output — which is the entire point, because that’s what makes the join work. A stable join key is a stable identifier for a person
  • The input space is small enough to brute-force. Email addresses, phone numbers and postcodes are enumerable. Hashing a UK mobile number leaves roughly 10⁸ possibilities — trivially reversible with a modern GPU and a rainbow table
  • Anyone else who hashes the same email gets the same value, which is exactly why ad platforms accept hashed emails for matching. If it were anonymous, the matching wouldn’t work — Hashing

Adding a secret salt helps — it stops anyone without the salt from reversing or cross-matching — but the data remains pseudonymised for you, because you hold the salt and can still single out an individual.

What actually anonymises

Anonymisation means accepting a real loss of utility. Anything that preserves per-person analysis hasn’t anonymised.

TechniqueWhat it doesCost
AggregationReport counts, never rowsNo user-level analysis at all
k-anonymitySuppress any group smaller than k peopleSmall segments disappear
GeneralisationPostcode → region; age → bandPrecision
Noise / differential privacyAdd calibrated randomness so no individual affects the output detectablySmall numbers become unreliable — Privacy-Preserving Measurement
Irreversible deletion of the keyDrop the identifier column entirelyCannot join, cannot delete on request, cannot correct

k-anonymity is the practical one. A dataset is k-anonymous when any combination of the retained attributes matches at least k people — so no row is unique.

✗  age 34, postcode SW1A 1AA, bought a wheelchair ramp   → one person
✓  age 30–39, region London, category mobility aids      → k = 340

The re-identification risk is combinatorial, and this is what people underestimate. Individually harmless fields become identifying together — gender, birth date and postcode district identify a large fraction of a population between them. Three or four “anonymous” attributes is usually enough to single someone out.

The practical position for analytics

Most analytics data is pseudonymised and should be treated as personal data. That’s the honest default. What to do with it:

  • Separate the identity from the behaviour. Keep the mapping (person ↔ pseudonymous ID) in one governed place; keep the event stream keyed on the pseudonym. It limits who can re-identify, which is a genuine reduction in risk even though it isn’t anonymisation
  • Set retention on the raw stream — 14 months is a common choice — and keep aggregates beyond it. Aggregates aren’t personal data, so they can be kept indefinitely, which is how you get long trends without long retention of identifiable rows — Data Retention, Event Streams vs Aggregates
  • Rotate the salt periodically if you use one, so an identifier’s lifetime is bounded. This does reduce linkability over time, at the cost of breaking longitudinal analysis across the rotation
  • Design deletion in. Being able to delete one person’s data across the whole estate is a hard requirement of pseudonymised data and effectively impossible to retrofit — UK GDPR and PECR for Analytics
  • Never claim “anonymised” in a privacy notice for data that isn’t. It’s the kind of statement that turns a technical issue into a misrepresentation

Where people get it wrong

  • “We only use hashed emails, so it’s anonymous.” No — and it’s the exact case ad platforms rely on being non-anonymous
  • “We removed names, so it’s anonymous.” Names were never the identifier that mattered
  • “IP addresses aren’t personal data.” They generally are, in the UK and EU. Truncation reduces but doesn’t eliminate the issue — PII in Analytics
  • “It’s anonymous because we can’t identify them.” The test is whether anyone reasonably could, using means reasonably likely to be used — including by combining with other data
  • Anonymising after collection. The data was personal at collection, so the obligations attached then. Anonymisation limits future risk; it doesn’t retrospectively make the collection lawful

Where it interacts

  • PII in Analytics — personal data arriving by accident, which is the more common problem than deliberate identifier design
  • Privacy-Preserving Measurement — the techniques that get genuine anonymity, and what they cost in precision
  • Consent Management — pseudonymised data still needs a lawful basis, so this doesn’t remove the consent question
  • Data Retention — the obligation that applies to pseudonymised data and doesn’t apply to anonymous aggregates