Event Streams vs Aggregates
Date: 2026-08-17
Keeping every event as an immutable row, versus keeping pre-computed counts. Aggregates are cheap, fast and answer only the questions you anticipated; streams are expensive and answer questions you haven’t thought of yet — including the one you’ll be asked next quarter.
An event stream stores one immutable row per thing that happened; an aggregate stores counts or sums already rolled up to a fixed grain, discarding the rows.
The two shapes
EVENT STREAM AGGREGATE
user event time props date event count
u1 page_view 10:00 {sku:A} 2026-08-17 page_view 41,882
u1 add_to_cart 10:12 {sku:A} 2026-08-17 add_to_cart 6,204
u2 page_view 10:14 {sku:B} 2026-08-17 purchase 1,247
u1 purchase 10:31 {£48.00}
… 3 rows/day
41,882 rows/day storage: trivial
storage: real query: instant
query: needs compute
The aggregate cannot answer “how many people who viewed A also bought B”. That information was destroyed at write time, and no amount of querying recovers it. That’s the entire trade.
What each makes impossible
| Question | Stream | Daily aggregate |
|---|---|---|
| How many purchases yesterday? | ✓ | ✓ instantly |
| What’s the conversion rate? | ✓ | ✓ if both counts kept |
| Conversion by device and channel and new/returning? | ✓ | ✗ unless pre-computed for that exact combination |
| What did users do between add-to-cart and purchase? | ✓ | ✗ |
| Redefine “session” as 45 minutes and re-run last year | ✓ | ✗ |
| Cohort retention for users acquired in March | ✓ | ✗ |
| Same as above, but excluding a bot pattern found today | ✓ | ✗ |
The last three are the ones that matter commercially, and they’re all the same underlying property: a stream lets you change the definition and re-run history; an aggregate has already applied the definition permanently.
The combinatorial problem with pre-aggregation
The tempting fix — pre-compute more combinations — doesn’t scale.
dimensions you might slice by
device (3) × channel (8) × new/returning (2) × country (12)
× product category (40) × day (365)
pre-computed rows for one metric 3 × 8 × 2 × 12 × 40 × 365 = 8,409,600
add one dimension, say membership tier (4) → 33,638,400
Every new dimension multiplies the aggregate table, and you still can’t answer a question involving a dimension you didn’t include. Meanwhile the raw stream for the same period might be 15 million rows and answers all of them. Past a certain number of dimensions, the aggregate is larger than the stream it summarises — which is the point at which pre-aggregation has stopped being an optimisation.
The shape that actually works
Not a choice — a layering. Keep the stream, derive the aggregates, treat the aggregates as disposable.
RAW EVENTS immutable, append-only, never edited
│ ← the source of truth. keep it as long as
│ retention policy and cost allow
↓
CLEANED / MODELLED bots removed, sessions assigned, identities
│ stitched, currency normalised
│ ← definitions live HERE, and can be changed
↓ and re-run — Warehouse-First Analytics
DAILY AGGREGATES pre-computed rollups for dashboards
← rebuildable from above. deleting one is safe
The rule that keeps this honest: aggregates must be reproducible from the stream. The moment a number exists only in an aggregate — because the raw data expired, or the aggregate was hand-corrected — you’ve lost the ability to audit or restate it, and every discrepancy becomes unresolvable.
Raw events are immutable. A correction is a new event or a change in the modelling layer, never an update to a past row. This is what makes “what did we believe on 3 March” answerable — Immutability.
When aggregates are genuinely right
Not a compromise, actually correct:
- Real-time dashboards where the query must return in milliseconds
- Very high volume where storing every event is genuinely unaffordable — ad impressions, IoT telemetry
- Long retention of summaries past the point where raw personal data must be deleted. Aggregates that can’t identify anyone fall outside personal-data retention limits, so keeping five years of daily totals alongside 14 months of raw events is both cheaper and more compliant — Data Retention, Pseudonymisation and Anonymisation
- Third-party data you can’t get raw. Ad platforms give you aggregates whether you like it or not — Walled Garden Reporting
The cost of the stream, honestly
- Storage grows linearly and forever. Manageable in columnar storage, and it’s rarely the binding constraint
- Query cost is the real expense. Scanning a year of events per dashboard load is what produces alarming warehouse bills — which is precisely what the aggregate layer is for
- Partition and cluster by date, or every query scans everything — Query Planning, Indexing
- Personal data at volume carries deletion obligations across the whole stream, which is a real engineering task rather than a policy line — UK GDPR and PECR for Analytics
Where it interacts
- Warehouse-First Analytics — the architecture this argues for, and the reason definitions become changeable
- Sessionisation — the clearest example: sessions computed at collection time are permanent, sessions computed in modelling can be redefined and re-run
- Data Sampling — what vendors do when they can’t afford the stream, and the reason their numbers can’t be reconciled to yours
- Metric Design — a stream lets a definition be corrected retrospectively, which changes how much pressure there is to get it right first time