Tags: analytics concept

Event Streams vs Aggregates

Date: 2026-08-17


Keeping every event as an immutable row, versus keeping pre-computed counts. Aggregates are cheap, fast and answer only the questions you anticipated; streams are expensive and answer questions you haven’t thought of yet — including the one you’ll be asked next quarter.


An event stream stores one immutable row per thing that happened; an aggregate stores counts or sums already rolled up to a fixed grain, discarding the rows.

The two shapes

EVENT STREAM                              AGGREGATE

user   event        time      props       date        event       count
u1     page_view    10:00     {sku:A}     2026-08-17  page_view   41,882
u1     add_to_cart  10:12     {sku:A}     2026-08-17  add_to_cart  6,204
u2     page_view    10:14     {sku:B}     2026-08-17  purchase     1,247
u1     purchase     10:31     {£48.00}
…                                         3 rows/day
41,882 rows/day                           storage: trivial
storage: real                             query: instant
query: needs compute

The aggregate cannot answer “how many people who viewed A also bought B”. That information was destroyed at write time, and no amount of querying recovers it. That’s the entire trade.

What each makes impossible

QuestionStreamDaily aggregate
How many purchases yesterday?✓✓ instantly
What’s the conversion rate?✓✓ if both counts kept
Conversion by device and channel and new/returning?✓✗ unless pre-computed for that exact combination
What did users do between add-to-cart and purchase?✓✗
Redefine “session” as 45 minutes and re-run last year✓✗
Cohort retention for users acquired in March✓✗
Same as above, but excluding a bot pattern found today✓✗

The last three are the ones that matter commercially, and they’re all the same underlying property: a stream lets you change the definition and re-run history; an aggregate has already applied the definition permanently.

The combinatorial problem with pre-aggregation

The tempting fix — pre-compute more combinations — doesn’t scale.

dimensions you might slice by
  device (3) × channel (8) × new/returning (2) × country (12)
  × product category (40) × day (365)

pre-computed rows for one metric   3 × 8 × 2 × 12 × 40 × 365  =  8,409,600

add one dimension, say membership tier (4)  →  33,638,400

Every new dimension multiplies the aggregate table, and you still can’t answer a question involving a dimension you didn’t include. Meanwhile the raw stream for the same period might be 15 million rows and answers all of them. Past a certain number of dimensions, the aggregate is larger than the stream it summarises — which is the point at which pre-aggregation has stopped being an optimisation.

The shape that actually works

Not a choice — a layering. Keep the stream, derive the aggregates, treat the aggregates as disposable.

RAW EVENTS            immutable, append-only, never edited
      │               ← the source of truth. keep it as long as
      │                 retention policy and cost allow
      ↓
CLEANED / MODELLED    bots removed, sessions assigned, identities
      │               stitched, currency normalised
      │               ← definitions live HERE, and can be changed
      ↓                 and re-run — Warehouse-First Analytics
DAILY AGGREGATES      pre-computed rollups for dashboards
                      ← rebuildable from above. deleting one is safe

The rule that keeps this honest: aggregates must be reproducible from the stream. The moment a number exists only in an aggregate — because the raw data expired, or the aggregate was hand-corrected — you’ve lost the ability to audit or restate it, and every discrepancy becomes unresolvable.

Raw events are immutable. A correction is a new event or a change in the modelling layer, never an update to a past row. This is what makes “what did we believe on 3 March” answerable — Immutability.

When aggregates are genuinely right

Not a compromise, actually correct:

  • Real-time dashboards where the query must return in milliseconds
  • Very high volume where storing every event is genuinely unaffordable — ad impressions, IoT telemetry
  • Long retention of summaries past the point where raw personal data must be deleted. Aggregates that can’t identify anyone fall outside personal-data retention limits, so keeping five years of daily totals alongside 14 months of raw events is both cheaper and more compliant — Data Retention, Pseudonymisation and Anonymisation
  • Third-party data you can’t get raw. Ad platforms give you aggregates whether you like it or not — Walled Garden Reporting

The cost of the stream, honestly

  • Storage grows linearly and forever. Manageable in columnar storage, and it’s rarely the binding constraint
  • Query cost is the real expense. Scanning a year of events per dashboard load is what produces alarming warehouse bills — which is precisely what the aggregate layer is for
  • Partition and cluster by date, or every query scans everything — Query Planning, Indexing
  • Personal data at volume carries deletion obligations across the whole stream, which is a real engineering task rather than a policy line — UK GDPR and PECR for Analytics

Where it interacts

  • Warehouse-First Analytics — the architecture this argues for, and the reason definitions become changeable
  • Sessionisation — the clearest example: sessions computed at collection time are permanent, sessions computed in modelling can be redefined and re-run
  • Data Sampling — what vendors do when they can’t afford the stream, and the reason their numbers can’t be reconciled to yours
  • Metric Design — a stream lets a definition be corrected retrospectively, which changes how much pressure there is to get it right first time