Analytics MOC
Date: 2026-08-16
Analytics is the chain from a thing happening in a browser to a number someone acts on. Every link in it — collection, identity, storage, definition, attribution, analysis — can break quietly, and the number still appears.
Ordered by dependency. Privacy decides what may be collected; collection decides what identity can do; identity decides what attribution can do; everything upstream decides whether the analysis means anything. Read down, not around.
Unwritten links are dimmed by Obsidian — that’s the gap list.
Working from a broken number rather than a topic? Jump to When a number looks wrong at the end.
The unit
What everything else is made of.
- Events and Properties — the atom: a named thing that happened, plus the key-value context carried with it
- Page Views vs Events — why the pageview-centric model broke on single-page apps, and what the event model replaced it with
- Event Taxonomy Design — the event-versus-property decision, one naming convention, and why an event can never be renamed
- Ecommerce Event Schema — the inherited convention:
view_item→add_to_cart→begin_checkout→purchase, and the item array every tool expects to find - Tracking Plans — the contract between whoever wants the number and whoever writes the code, and the enforcement without which it’s a wish
- The Data Layer — the page-level object that decouples what the site knows from what any tag reads, its push-queue mechanism, and the state that persists when it shouldn’t
- Checkout Instrumentation Constraints — what stays measurable when checkout isn’t your page: sandboxed pixels, no arbitrary scripts, a fixed event set
Privacy and governance
Constrains everything below it. Not a compliance afterthought — it decides what can be collected at all.
- Consent Management — consent as a gate on collection: categories, storage of the decision, and re-prompting
- UK GDPR and PECR for Analytics — the two regimes that actually apply to a UK site, and where they differ [CHECK: current ICO position on analytics cookies]
- Legitimate Interest vs Consent — the lawful-basis fork, and why analytics rarely gets to pick the easy one
- PII in Analytics — what counts as personal data once it’s in an event property, and how it leaks in by accident
- Pseudonymisation and Anonymisation — the difference the regulation cares about, and why most “anonymous” analytics isn’t
- Data Retention — expiry as a policy and as a destructive operation on your own trend lines
- Data Residency — where the data physically lands, and when that’s a contractual requirement
- Privacy-Preserving Measurement — aggregate reporting, noise injection, on-device attribution; the direction of travel
Collection
How the event gets from the page to somewhere it persists. Determines the ceiling on data quality — nothing downstream recovers what was never sent.
- Client-Side vs Server-Side Tracking — the core tradeoff: context and cost versus completeness and control
- Tag Managers — the indirection layer that lets marketing deploy without a release, and what that costs
- Server-Side Tag Management — moving the container off the page: one first-party request in, many vendor requests out
- Server-Side Conversion APIs — sending conversions to ad platforms from your backend instead of the browser
- Offline Conversion Imports — sending outcomes that happen after the website (qualified lead, closed deal, refund) back to ad platforms, so bidding optimises to revenue rather than form fills
- Autocapture — recording every click and pageview by default, then defining events retrospectively
- Event Batching and Delivery — queuing, retries,
sendBeacon, and the events lost when the tab closes mid-flight - Ad Blockers and Tracking Loss — the missing share, and why it’s missing non-randomly
- Instrumentation Debugging — the tooling: debug views, network inspection, preview modes, local event validation
Identity
Turning a stream of events into a story about a person. The step most silently wrong.
- Sessionisation — cutting the event stream on an arbitrary timeout, and why session conversion rate moves when the timeout does
- Identity Stitching — joining anonymous activity to a known person, and why every user-level number including attribution sits downstream of it
- Anonymous and Identified Users — the two populations, why they’re counted differently, and what breaks at the boundary
- User Counting — the least reliable number in any tool, and why two tools never agree on it
- Cross-Device Tracking — deterministic versus probabilistic device graphs, and the accuracy claims to distrust
The pipeline
Where events go once collected, and what has to hold for them to still be countable at the other end.
- Event Streams vs Aggregates — raw immutable log versus pre-computed rollup, and what each makes impossible
- Warehouse-First Analytics — landing raw events in a warehouse and modelling after the fact, rather than at collection time
- Reverse ETL — pushing modelled data back out to the tools that act on it, which is what closes the loop
- Customer Data Platforms — one collection point fanning out to many destinations, plus a profile store
- Schema Enforcement — rejecting or quarantining malformed events at the edge rather than discovering them in a dashboard
- Data Quality Monitoring — alerting on the pipeline itself, not just on the application that feeds it
- Idempotency and Deduplication — at-least-once delivery means duplicates; event IDs are how you survive them
- Cardinality — why high-cardinality properties get truncated, sampled or dropped, and what disappears with them
Metrics
Definition is the discipline. Most metric arguments are definition arguments in disguise.
- Metric Design — what separates a metric that changes a decision from one that decorates a slide, and writing the definition down so it survives: numerator, denominator, filters, timezone, owner
- Leading and Lagging Indicators — trading predictive speed against the thing you actually care about
- North Star Metric — the single-number framing, its uses, and the behaviour it distorts
- Vanity Metrics — numbers that only go up, and the test that identifies them
- Metric Drift — the definition changing under a trend line that keeps rendering as though it didn’t
- Engagement Metrics — bounce, engaged sessions, time on page, scroll depth, and what each actually measures
- Revenue Metrics — gross versus net, refunds, tax, discounts and currency conversion, all of which move the number
Attribution
Assigning credit for an outcome to the things that preceded it. Depends on identity holding up across sessions and devices.
- Attribution Models — last click, first click, linear, positional, decay, data-driven; the same order paying paid search anywhere from £0 to £200
- Attribution Windows — the lookback period as a modelling choice, and how it manufactures or hides conversions
- Channel Taxonomy — the classification that turns raw sources into channels, and where the rules live
- UTM Governance — campaign parameter conventions, and the naming chaos when nobody owns them
- Direct Traffic and Lost Referrers — how referrer data goes missing, and what “direct” really contains
- Multi-Touch Attribution — spreading credit across a path, and why the models disagree by design
- View-Through Attribution — crediting impressions nobody clicked, and the correlation problem it creates
- Self-Reported Attribution — the “how did you hear about us?” field as a check on click-based models, and the recall bias it brings with it
- Modelled Conversions — vendors estimating what consent and blocking removed, and what you’re trusting when you use them
- Walled Garden Reporting — why every ad platform claims more conversions than you recorded, all at once
- Marketing Mix Modelling — regressing outcomes on spend at aggregate level, with no user-level data at all
Analysis methods
What you do once the data is trustworthy. Each of these has a failure mode more interesting than its procedure.
- Funnel Analysis — ordered step conversion, and the choices (window, ordering, re-entry) that move the result
- Cohort Analysis — grouping by shared start point so mix changes stop masquerading as behaviour changes
- Retention Curves — the shape that shows whether anything sticks, and how to read its flattening
- Segmentation (analysis) — cutting the population to find the group the average was hiding
- Metric Decomposition — splitting a moved number into its factors, and into mix versus rate across segments; where most “conversion fell” investigations end
- Path Analysis — actual route-taking through a site, and why the output is usually unreadable
- Seasonality — the recurring pattern under the trend, and comparing like periods rather than adjacent ones
- Anomaly Detection — separating a real break from ordinary variance, without alerting on every Monday
- Data Sampling — when a tool stops counting everything, and what that does to small segments
- Benchmarking — comparing against industry figures, and why the comparison is nearly always invalid
Behavioural and qualitative
Different evidence class. Explains why, never how many.
- Session Replay — watching reconstructed sessions, its privacy obligations, and the sampling bias in what you choose to watch
- Heatmaps — aggregated click, move and scroll rendering, and the layout conditions that invalidate it
- Form Analytics — field-level abandonment, time-per-field and error rates
- Voice of Customer Data — surveys, feedback widgets and support tickets as an analytics input
Delivery
The last link. A correct number nobody can find or interpret has not been delivered.
- Dashboard Design — designing for a decision rather than for coverage
- Self-Serve Analytics — letting others query without letting them redefine your metrics
- Alerting on Metrics — thresholds versus deviation, and the alert-fatigue failure
- Annotation and Change Logs — recording deploys, campaigns and outages against the timeline so spikes stay explicable
Failure modes
Worth their own notes because they’re the recurring diagnoses, not one-off bugs.
- Double Counting — the same event recorded twice, and the four ways it usually happens
- Bot and Internal Traffic — filtering non-human and staff traffic, and what filtering breaks
- Timezones and Date Boundaries — one report’s Monday being another’s Sunday, and the reconciliation that follows
- Tool Discrepancies — why GA4, the ad platform and the warehouse never agree, and which differences are expected
Guides
- Guide - Instrumenting a Site — greenfield: choosing the stack, taxonomy, plan, data layer, consent, and QA before launch
- Guide - Auditing a Tracking Plan — checking that what’s specified, what fires and what lands are the same set
- Guide - Diagnosing a Metric Movement — the ordered checklist that separates a real change from a tracking break
- Guide - Replatforming and Migrations — running two systems in parallel, and the measurement continuity most migrations lose
Platforms
Tool shape and quirks only — the concepts live above.
- GA4 · PostHog · Segment · Amplitude · Google Tag Manager · BigQuery · Looker Studio
When a number looks wrong
Entry from the broken figure rather than the topic. Cheapest cause first — the boring explanation is nearly always the right one.
| Symptom | Usual cause |
|---|---|
| Fewer orders than the order system | Timezone boundaries first; the rest is consent denial, blockers and bots. A residual gap is normal — Timezones and Date Boundaries · Ad Blockers and Tracking Loss |
| More orders than the order system | Double-firing on refresh or back-navigation, then test orders and staff traffic — Double Counting · Bot and Internal Traffic |
| Revenue 100× too high | Currency subunits: pence into a field expecting pounds — Revenue Metrics |
| Conversion rate moved with no site change | The denominator moved. Sessionisation rules changed, not behaviour — Sessionisation · Metric Drift |
| Users doubled, sessions flat | Identity broke: cookie lifetime, consent, or a stitching regression — Identity Stitching · User Counting |
| A metric stepped once and stayed | A definition changed. Find the deploy or container publish on that date — Metric Drift · Annotation and Change Logs |
| Everything fell ~30% on one date | Consent banner change, or tags no longer firing under denial — Consent Management |
| Gradual decline nobody can date | Cookie lifetime caps eroding returning-user recognition. There’s no single event to find — Browser Privacy Restrictions |
| Ad platforms claim more conversions than exist | Each attributes in its own model and window, and none sees the others. Not reconcilable — Walled Garden Reporting · View-Through Attribution |
| A channel suddenly looks terrible | Check the stitch rate before the model: lost joins collapse credit onto the last session — Identity Stitching · Attribution Models |
| A funnel step has more people than the step before | Re-entry, out-of-order completion, or too wide a window — Funnel Analysis |
| Two tools disagree on sessions or users | Different timeouts, identity models and bot filters. The definitions differ, not the traffic — Tool Discrepancies · Sessionisation |
| A number nobody can reproduce | No written definition: numerator, denominator, filters, timezone, owner — Metric Design |
The order to work in, when several could be true at once:
- Timezone and date range — the largest source of fake discrepancies
- Definition — is it the same metric in both places
- Collection — is the event firing, once, at the right moment
- Identity — are sessions and users counted the way you think
- Attribution — only meaningful once 1–4 hold
- Behaviour — the last explanation to reach for, not the first
Most investigations that stall started at 5 or 6. Whole-surface procedure: Guide - Auditing a Tracking Plan.
Borders
Filed elsewhere, needed constantly here.
- Statistics — Sample Ratio Mismatch · Simpson’s Paradox · Correlation and Causation · Confidence Intervals · Regression to the Mean · Survivorship Bias · Ratio Metrics · Communicating Uncertainty
- Experimentation — Guardrail Metrics · Randomisation Unit · Experiment Assignment Tracking
- Web Development — Browser Privacy Restrictions · Cookies · Client Storage · Core Web Vitals · Real User Monitoring
- Commerce & Growth — Incrementality Testing · Geo Holdout Tests · Customer Lifetime Value · Cohort Revenue · Contribution Margin