PII in Analytics
Date: 2026-08-16
Personal data gets into event streams by accident, not by decision — a URL parameter, a form field, an autocaptured input. Once it’s there it’s in every vendor’s system, and deletion means finding it in all of them.
What it is
Personal data in analytics is any field that identifies or could identify a person. Under UK GDPR that’s broader than the American “PII” framing suggests: it includes anything singling someone out, so a persistent device identifier counts even with no name attached.
The practical set: names, emails, phone numbers, addresses, postcodes, order numbers tied to a person, IP addresses, and free-text fields that could contain any of the above.
How it gets in
Almost never deliberately. Five routes, all common:
1 URL parameters /reset?email=alex@example.com
captured automatically as page path
2 Form autocapture input values swept up by a replay
or autocapture tool
3 Data layer pushed "for one tag", visible to all
4 Free-text fields search terms, feedback, notes:
"call me on 07700 900123"
5 Session replay the whole page, including what
the user typed
See: Autocapture · Session Replay
Route 1 is the most common and the least noticed. A password reset or email confirmation link with the address in the query string becomes a page path in your analytics, then in your warehouse, then in your BI tool.
Why it matters more than it looks
- It spreads instantly. Anything in the data layer is readable by every tag on the page, including vendors you never audited — Third-Party Scripts
- Deletion becomes multi-system. A subject access or erasure request means finding that person in analytics, the warehouse, every vendor and every export
- Most vendor terms prohibit it. Sending personal data to a tool not contracted for it is a contractual breach as well as a compliance one
- It’s usually undiscovered. Nothing breaks. It sits there until an audit or an incident
Preventing it
- Strip URL parameters at collection. Allowlist the ones you want rather than blocklisting the ones you don’t — a blocklist never anticipates the next one
- Never put personal data in the data layer. Hash before pushing if a vendor needs an identifier, and treat “just for the one tag” as false
- Verify input masking is on for replay and autocapture tools rather than assuming the default — Session Replay
- Use surrogate identifiers. A hashed user ID rather than an email; a random order reference rather than a customer name
- Filter server-side as a backstop, before events reach vendors — Server-Side Tag Management
- Design URLs so tokens aren’t identifiers — a one-time token in a reset link, not an address — URL Structure
Finding what’s already there
Worth running once, then quarterly:
scan event properties and page paths for
@ email addresses
\d{11} phone-shaped numbers
postcode patterns
known staff domains
Do it in the warehouse where you can query everything, not in a vendor UI where you can only sample. Add it to the property checks in Guide - Auditing a Tracking Plan — it’s step 4, and the note there says explicitly that a hit is a disclosure incident rather than a data-quality finding.
If you find it
Treat it as an incident, in this order:
- Stop the source — fix the collection before cleaning up, or you’ll clean the same thing twice
- Establish the spread — which vendors received it, over what period
- Delete where you can — your warehouse first, then each vendor’s deletion mechanism
- Record it — what, when, how long, what was done. That record is what a regulator asks for
- Add a check so it can’t recur silently — Schema Enforcement
The distinction that decides severity: was it collected, or was it also shared? Personal data in your own warehouse is a problem you can fix. Personal data sent to five vendors is five problems, four of which you don’t control.