Data Retention
Date: 2026-08-16
A policy and a destructive operation at the same time. Expiry is a legal obligation and it silently truncates every year-on-year comparison you’ll want to make — so the window is a decision, not a default.
What it is
Data retention is how long event data is kept before deletion or aggregation. UK GDPR requires that personal data not be kept longer than necessary for the purpose it was collected for, which makes an indefinite default indefensible.
The tension is direct: compliance pushes the window down, analysis pushes it up.
What you lose as the window shortens
| Window | What becomes impossible |
|---|---|
| 14 months | Two-year seasonal comparison; anything about last Black Friday but one |
| 12 months | Year-on-year with any margin. A 12-month window can’t compare complete years |
| 6 months | Annual seasonality; cohort retention past six months |
| 2 months | Almost all longitudinal analysis |
The 14-month figure matters because several tools default to it, and it’s the shortest window that supports a genuine year-on-year comparison with room to run the query. [CHECK: current default and configurable retention periods in GA4 — these have changed and are tier-dependent.]
In plain terms: the day your retention window expires, a slice of your history is gone permanently. Nothing recovers it, and you find out when someone asks a question you can no longer answer.
The move that resolves it
Retention limits apply to personal data. They don’t apply to genuinely aggregated data with no ability to single anyone out.
raw events 14 months personal, expires
│
▼ before expiry
daily aggregates indefinite sessions, orders, revenue
by dimension by date × channel × device
Aggregate before you expire, not after. Roll raw events into daily tables carrying the dimensions you’ll want — date, channel, device, template, country — and keep those indefinitely. You lose the ability to ask new user-level questions about old periods, and you keep every trend line.
The discipline this requires: deciding in advance which dimensions matter, because the aggregate can only contain what you thought to put in it. That decision is worth more attention than the retention setting itself — see Warehouse-First Analytics.
Note that aggregation must be genuine. A table with one row per user, however pseudonymised, is still personal data — Pseudonymisation and Anonymisation.
Setting the policy
- Different windows for different data. Raw event streams shorter, aggregates indefinite, order records governed by finance and tax obligations rather than by analytics
- Write down the purpose for each. “Necessary for” is the legal test, and it needs an answer per dataset
- Apply it everywhere. Retention set in one tool while the same events sit forever in your warehouse and in four vendors’ systems is a policy in name only
- Automate deletion. A policy nobody executes is worse than none — you’ve documented an obligation you’re not meeting
- Handle erasure requests separately. Retention is scheduled bulk deletion; a subject request is targeted and immediate, and needs the ability to find one person across every system — PII in Analytics
Before you shorten it
The mistake that’s genuinely unrecoverable: reducing a retention window without exporting first. Tools apply the new setting to existing data, and the deletion is immediate and permanent.
1 decide the new window
2 export everything older than it, aggregated
3 verify the export
4 then change the setting
Skipping step 2 destroys history in an afternoon that took years to accumulate. Worth stating loudly, because “tighten retention” arrives as a compliance ticket with no mention of the analytical cost, and whoever actions it usually isn’t the person who’ll need the data.