Event Taxonomy Design
Date: 2026-08-16
A taxonomy is a compression scheme. Few event names with rich properties answer more questions than many event names, because grouping is cheap and un-grouping is impossible.
The one decision
Every instrumentation choice reduces to: is this a new event, or a property value on an existing one?
Get it wrong towards too many events and the analysis is a manual union of names nobody can enumerate. Get it wrong towards too few and the event means nothing without three filters attached.
The test: would you ever want these two things on the same chart, compared? If yes, one event, distinguished by a property.
| Instinct | Better | Why |
|---|---|---|
header_cta_clicked, footer_cta_clicked | cta_clicked + position: header | footer | You will want them ranked against each other |
signup_completed, purchase_completed | Separate events | Genuinely different outcomes, never summed |
viewed_product_shoes | product_viewed + category: shoes | Category is unbounded — see Cardinality |
checkout_step_2 | checkout_step_viewed + step: 2, step_name: delivery | Step count changes; the event name shouldn’t |
The difference is visible the moment you look at what lands:
// FOUR events. Four names you must already know to find them.
{ "event": "clicked_header_cta" }
{ "event": "Clicked Footer CTA" }
{ "event": "cta_click_sidebar" }
{ "event": "ctaClicked", "location": "modal" }
// ONE event. Four rows.
{ "event": "cta_clicked", "position": "header", "text": "Book now" }
{ "event": "cta_clicked", "position": "footer", "text": "Book now" }
{ "event": "cta_clicked", "position": "sidebar", "text": "Get a quote" }
{ "event": "cta_clicked", "position": "modal", "text": "Get a quote" }And in what you can then ask:
-- with the taxonomy: one query, complete, self-documenting
select position, count(*) from events
where event_name = 'cta_clicked' group by 1;
-- without it: you must already know every name that exists,
-- and any name added since you last looked is silently missing
select event_name, count(*) from events
where event_name in ('clicked_header_cta','Clicked Footer CTA',
'cta_click_sidebar','ctaClicked') group by 1;The second query cannot tell you what it’s missing. That’s the real cost of a loose taxonomy — not ugliness, but the fact that incomplete answers look identical to complete ones.
Corollary: never put a variable in an event name. Anything that could take more than a handful of values is a property. Event names are a closed set you can print on one page; properties are open.
Naming
The convention barely matters. Having exactly one does.
- Object–action, past tense —
product_viewed,checkout_started,payment_failed. Sorts usefully: everything about products lands together alphabetically - Action–object —
view_product,begin_checkout. Google’s ecommerce spec uses this, so if you’re inheriting Ecommerce Event Schema you’re already committed to it for those events - One case convention —
snake_casethroughout, orTitle Casethroughout. Most tools treatproduct_viewedandProduct Viewedas two separate events, and nothing warns you
The realistic outcome is mixed: inherited vendor events in one style, your own in another. Pick a rule for what you own, write down which events are inherited and therefore exempt, and stop relitigating it.
Properties
Properties are where taxonomies actually decay, because nobody reviews them.
- One meaning per property, permanently.
valuemeaning order total in one event and item price in another is a permanent trap. If the meaning differs, the name differs - Fix the unit and record it. Pence or pounds, never both.
price_penceis ugly and correct;priceis clean and will eventually be off by 100× - Fix the type.
trueand"true"and1are three values in most query engines. Booleans stay booleans - Prefer flat. Nested objects are supported unevenly and query awkwardly; the exception is the item array in Ecommerce Event Schema, which is standardised enough to be worth the trouble
- Empty vs missing vs zero are three different states. Decide which one “the user didn’t select a size” is, and use it everywhere
Global properties — user ID, device, plan tier, consent state, experiment assignment — get attached to every event by the SDK rather than passed per call. Anything a developer has to remember to include is a property that will be missing on a third of events.
Versioning
You cannot rename an event. The old name owns the history; the new name starts empty. Both are true and both are permanent.
The options, in order of preference:
- Live with the bad name. A wrong-but-consistent name costs a sentence of explanation; a rename costs a permanent seam in every trend line
- Add a property that carries the new distinction, leave the name alone
- New event, both firing in parallel for a defined overlap, old one marked deprecated in the tracking plan with its retirement date. History stays queryable as a union across the seam
- Version in the name —
checkout_started_v2. Honest, and every query now needs anINclause forever
Deprecated never means deleted. Stopping the fire is safe; removing the event from the plan destroys the only record of what the historical data meant.
Why they rot
The taxonomy is the one part of instrumentation with no natural owner. Each feature team names its own events sensibly for that feature, and eighteen months later there are four verbs for “the user saw something” and the person asking “how many people saw a product page” cannot answer it without a list nobody has.
- The cost is deferred and the saving is immediate, which is exactly the shape of debt that always gets taken on
- Nothing breaks. Bad taxonomy produces working dashboards with quietly wrong numbers, so no incident forces the fix
- The person who names the event is never the person who queries it
The only defences that work are structural, not cultural: a plan the code is generated from, and validation that rejects unknown events at the boundary. See Tracking Plans and Schema Enforcement.
The tradeoff
A strict taxonomy adds a review step to every feature that touches instrumentation, and teams under delivery pressure will route around it. A loose one ships fast and makes the second year of data unusable.
Neither extreme survives contact. What works: a tight closed set for the events that feed metrics anyone reports on, and tolerance for a long tail of scruffy diagnostic events that nobody promises anything about. Mark which is which in the plan — the mistake is pretending the whole surface is governed to the same standard.
Autocapture is the other answer entirely: record everything, name nothing, define events retrospectively. It moves the problem from naming to definition, and the definitions rot in the same way.