Warehouse-First Analytics
Date: 2026-08-16
Land raw events in a warehouse you own, and model them afterwards. It converts every definition from a permanent decision made at collection time into a query you can change and re-run over history.
What it is
Warehouse-first analytics treats your own data warehouse as the system of record, with vendor tools as destinations rather than sources.
VENDOR-FIRST WAREHOUSE-FIRST
events ──▶ vendor tool events ──▶ your warehouse
│ │
▼ ├──▶ vendor tools
vendor's model ├──▶ BI
vendor's session rules └──▶ reverse ETL
vendor's attribution
vendor's retention raw events kept,
modelled at query time
What it buys
The gains are all versions of one property: decisions become reversible.
- Definitions can be restated. Change a session timeout, a channel mapping, an attribution model — and re-run it over all history. In a vendor tool those are applied at ingestion and create a permanent seam — Metric Drift
- Sessionisation stops being a verdict. Computed as a window function at query time, so you can ask what last year looked like at 45 minutes — Sessionisation
- Channel Taxonomy fixes are retroactive. Correct a mapping and every historical report corrects with it. This is the one people feel most
- No sampling. Full data, every query, however deep the segment — Data Sampling
- Joins to everything you own. Orders, costs, margin, returns, CRM, stock. A vendor tool can’t tell you conversion by contribution margin; a warehouse can — Contribution Margin
- Retention under your control, with aggregation before expiry — Data Retention
- One definition, many tools. Every downstream system reads the same modelled tables, so “conversion rate” means one thing
What it costs
Real, and consistently understated in the pitch:
- Infrastructure and query spend, which scales with volume and with how carelessly people query
- Engineering. Pipelines, models, tests, monitoring — Data Quality Monitoring
- Latency. Vendor tools show data in minutes; a warehouse pipeline is often hourly or daily. For debugging a live issue that matters
- A skills dependency. Everything is SQL. A marketing team that self-served in a vendor UI now needs someone to build them a dashboard
- Modelling is a discipline, not a task. Raw events aren’t usable directly; someone has to build and maintain the models everyone queries
Where it’s genuinely worth it
- When definitions matter and keep changing — attribution, sessions, channels
- When you need to join to commercial data the vendor doesn’t have. This is the strongest single argument for retail: analytics that can’t see margin can’t tell you whether growth was profitable
- When sampling is biting on segments you care about
- When several tools disagree and you need one source to arbitrate — Tool Discrepancies
- When retention limits are about to destroy history
Where it isn’t
- Small volumes. If a vendor tool answers your questions unsampled, the warehouse is overhead
- No engineering capacity. A half-built pipeline is worse than a working vendor tool, because it looks authoritative and isn’t
- Real-time needs. Debugging a broken tag wants a live debug view, not a nightly model
The pragmatic shape
Not a migration. Most teams run both:
- Vendor tools stay for self-serve, real-time debugging and the reports marketing already uses
- Raw events land in the warehouse in parallel — often via the vendor’s own export, which is the cheapest starting point
- Modelled tables define sessions, channels and attribution once, in SQL, in version control
- BI reads the models, not the raw events
- Reconciliation between the two runs as a standing check
Step 2 alone is worth doing early even with no immediate plan for the data, because you cannot backfill events you never landed. Starting the export now costs little and means the history exists when you need it — which is the same argument as attaching properties in Events and Properties, one layer up.