Tags: analytics commerce concept
Customer Data Platforms
Date: 2026-08-17
A product category that collects customer data from everywhere, resolves it to one profile per person, and pushes segments out to the tools that act on them. The category is real; the marketing definition is broad enough to include products that share almost nothing, so the useful question is which of the four jobs you’re actually buying.
A customer data platform (CDP) is software that ingests customer events and records from many sources, merges them into persistent per-person profiles, and syncs audiences from those profiles to other tools.
The four jobs, which are separable
1 COLLECT one SDK, many destinations. events in, fan out to
analytics, ads, email, warehouse
← this is what most "CDPs" mostly are
2 RESOLVE merge anonymous IDs, emails, order records and device
IDs into one profile per person
← Identity Stitching. the genuinely hard part
3 SEGMENT "bought twice, not in 90 days, opted in to email"
computed and kept fresh
4 ACTIVATE push those segments into email, ads, on-site
personalisation — and keep them synced
← the part that produces revenue
Most organisations need 1 and 4, already have the data for 2, and can do 3 in a warehouse. Being clear about which one is the bottleneck is what stops a CDP purchase from being an expensive way to solve a problem you didn’t have.
The two architectures
The distinction that matters most, and the one vendor comparisons obscure:
| Packaged CDP | Composable / warehouse-native | |
|---|---|---|
| Where data lives | The vendor’s store | Your warehouse |
| Identity resolution | Their black box | Your models, inspectable |
| Segments defined in | Their UI | SQL / your modelling layer |
| Activation | Built-in connectors | Reverse ETL to the same destinations |
| Time to value | Weeks | Months |
| Cost model | Per profile / per event — grows with success | Warehouse compute + a sync tool |
| Lock-in | High — the profiles are theirs | Low — the data never left |
| Suits | Small data team, urgent need | Existing warehouse and a data team |
If you already have a warehouse and a modelling layer, the composable route is usually right — a CDP’s store becomes a second copy of data you already own, with a second set of definitions that will drift from the first — Warehouse-First Analytics, Reverse ETL.
The strongest argument for a packaged CDP is having no data team and needing activation this quarter.
Identity resolution is the part that’s hard
Everything else is plumbing. This is where CDPs earn their price or fail quietly.
inputs resolved profile
web anon_id a1f3… → 12 sessions
web anon_id c9b2… → 4 sessions person 4821
email alex@… → 6 opens · 3 devices
order #88214 → £180 · 2 emails
loyalty card 4429 → 14 in-store purchases · £2,410 lifetime
app install d31a… → 3 sessions · last seen 2 days ago
Ask any vendor these, because the answers differ enormously and the defaults are rarely what you want:
- What merges two profiles — email match, hashed email, device, probabilistic? Which are on by default?
- Can a merge be undone? Shared devices and family email addresses produce wrong merges, and an irreversible merge is a data quality problem you can’t fix
- What happens to history on merge — is it re-attributed retroactively, and do historical reports change?
- How is a conflict resolved when two profiles disagree on a field? Last-write-wins is the usual default and is frequently wrong
- Can you inspect why two identities merged? A black box here means you can’t debug a wrong profile, and wrong profiles reach customers
The failure modes
- Bought as an analytics tool. A CDP is an activation tool. Using it for analysis means analysing a processed, resolved, partially-modelled dataset with no access to what it did — Event Streams vs Aggregates
- A second source of truth. Segment definitions in the CDP that don’t match the warehouse produce two “active customer” counts, and the resulting meetings are about which is right — Metric Design, Self-Serve Analytics
- Cost scaling with success. Per-profile pricing means the bill grows with every anonymous visitor. Model the cost at three times current traffic before signing — [CHECK: pricing models vary by vendor and change; verify the billing axis and get a quote rather than relying on published figures]
- Consent not propagating. The CDP fans data out to a dozen destinations, so a consent decision has to travel with every event to every one of them. A CDP that collects before consent, or forwards to a destination the user declined, has industrialised a compliance failure — Consent Management
- Garbage in. Identity resolution can’t fix an event taxonomy that doesn’t identify things consistently. The CDP amplifies whatever discipline exists upstream — Event Taxonomy Design, Tracking Plans
The compliance surface
A CDP concentrates personal data by design, which raises the stakes on everything:
- It’s a processor and you’re the controller. Contracts, a record of processing, and a documented lawful basis — UK GDPR and PECR for Analytics, Legitimate Interest vs Consent
- Deletion must propagate to every downstream destination, not just the CDP’s own store. Ask specifically how, and test it
- Data residency where contracts require it — Data Residency
- Profiles are personal data even when keyed on a hashed identifier — Pseudonymisation and Anonymisation
Where it interacts
- Identity Stitching — the core capability, and the thing to evaluate a vendor on
- Warehouse-First Analytics — the alternative architecture, and increasingly the default for organisations with a data team
- RFM Segmentation and Lifecycle Stages — the segment definitions a CDP exists to compute and activate
- Server-Side Tag Management — overlapping collection capability, and a common source of two systems doing one job