Guide - Auditing a Tracking Plan
Date: 2026-08-16
Three sets have to match: what’s specified, what fires, and what lands. Every tracking defect is a gap between two of them, so the audit is three comparisons run in a fixed order — cheapest and broadest first.
Before you start
Get access, or stop. Destination tool with query access (or the warehouse), the tracking plan, the tag manager container, a staging environment you can transact on, and the order or CRM system you’ll reconcile against. Chasing access mid-audit is what turns two days into three weeks.
Decide which set is the reference. If the plan is current, the plan is correct and reality is the defect. If the plan is two years stale, reality may be more correct than the document, and the output is a corrected plan rather than a bug list. Say which before you start; auditing without deciding produces findings nobody can action.
If there is no plan, this is the wrong guide. Inventory first — steps 2 and 3 alone produce the plan you were missing. Then audit against it later.
Scope explicitly. Which surfaces (web, app, server), what date range, which destination is authoritative when tools disagree. A full audit of 200 events is weeks of work; the critical path is days.
1. Freeze the reference
- Version or snapshot the plan. The audit takes long enough that someone will edit it underneath you
- Rank events by the metric they’d break. For each reported metric, list the events it depends on. Anything supporting a number that reaches a decision-maker is critical path; everything else is the long tail
- Expect 10–25 critical-path events. That’s the audit. The tail gets an inventory and nothing more
2. Inventory what actually lands
Cheapest step, finds the most, requires no browsing. Query the destination for distinct event names over at least 30 days — long enough to cover a full weekly cycle and any monthly process.
select event_name,
count(*) as events,
count(distinct user_id) as users,
min(event_time) as first_seen,
max(event_time) as last_seen
from events
where event_time >= current_date - 30
group by 1
order by 2 desc;first_seen and last_seen are doing real work here — an event whose last_seen is three weeks ago broke three weeks ago, and that alone dates the incident.
Then produce the set diff against the plan: specified only, landed only, both.
If the destination samples queries at volume, confirm you’re querying unsampled data before trusting counts — see Data Sampling.
3. Triage the diff
| Class | Usual cause | Action |
|---|---|---|
| In plan, zero volume | Never implemented; or broke silently; or genuinely seasonal | Check last_seen. Never-seen is a build gap, stopped-seen is an incident |
| In plan, volume far below expectation | Fires on one platform only, or one template, or behind consent | Segment by device, page template and consent state before assuming it’s broken |
| Volume, not in plan | Undocumented feature work, tag manager addition, vendor auto-event, Autocapture output | Identify the owner. Either document it or turn it off — an unowned event is a future wrong answer |
| Near-duplicate names | Case, tense or word-order collision — the classic Event Taxonomy Design decay | Do not rename. Record both, pick the survivor, deprecate the other with a date |
| Both sides match | — | Goes through to step 4 |
Sudden volume changes matter as much as absence. Plot daily volume per critical event and look for step changes — they date to a deploy, a container publish or a consent banner change, and Annotation and Change Logs is what makes them attributable.
4. Check properties — critical path only
Per event, per property, over the same window:
- Presence — null rate on every property the plan marks required. Anything above ~0% needs an explanation; anything above 5% is a defect
- Type consistency — a property arriving as string on some events and number on others. Sort by distinct values and the mixed types are usually visible immediately
- Units — revenue in pounds or pence, never both. Check the distribution’s magnitude rather than the definition
- Cardinality — distinct value count against expectation. A property specified as five values holding 4,000 is a variable that leaked into a property, and it will hit collection limits: Cardinality
- Plausibility — negatives, zeros, £0.01 test orders,
undefinedas a literal string, staff email domains. See Bot and Internal Traffic - Personal data that shouldn’t be there — scan properties for anything resembling an email, name, postcode or raw address. This one is not a data quality finding, it’s a disclosure incident: PII in Analytics
5. Verify triggers by hand
The step no query replaces. The plan says when an event fires; only walking it proves that. Use the tool’s debug view plus network inspection — see Instrumentation Debugging.
For each critical-path event, four passes:
- Happy path — fires once, at the specified moment, with the specified properties populated
- Failure path — declined payment, out-of-stock, validation error. Does
purchasefire anyway on optimistic UI? This is the single most expensive defect class and it’s rarely tested - Repeat and back button — refresh the confirmation page, navigate back and forwards. Anything firing twice is Double Counting, and the fix lives in Idempotency and Deduplication
- Route change — in a single-page app, navigate between templates and check for values surviving from the previous view. The persistence trap in The Data Layer
Then the two dimensions people skip:
- Consent states. Run the whole critical path with consent granted and denied. What fires under denial is both a compliance question and the explanation for a chunk of your volume gap: Consent Management
- Safari and Chrome separately. Storage and identity behaviour differ enough to change what you see: Browser Privacy Restrictions
Where checkout is hosted or locked down, some of this simply isn’t observable and the audit records that as a known limit rather than a finding — Checkout Instrumentation Constraints.
6. Reconcile against the source of truth
Compare analytics purchases and revenue to the order system, same period, same timezone.
- Align timezones first. A large fraction of apparent discrepancies are date boundaries, not tracking: Timezones and Date Boundaries
- Expect a gap. Zero difference means double-counting or a filter you don’t know about, not success. The realistic contributors are consent denial, Ad Blockers and Tracking Loss, bots, and events lost on page unload
- Quantify it and write it down. A stable gap is a correction factor you can apply and explain. A drifting gap is an open incident
- Do the same per channel if attribution matters downstream, since the gap is rarely uniform across traffic sources: Tool Discrepancies
7. Check identity
Cheap, and it gates everything attribution-related. Two numbers:
- Stitch rate — the share of conversions joined to a pre-login session. Low means path-based analysis is measuring the last session only, whatever the reports say: Identity Stitching
- Anonymous IDs per known user per month — the churn rate. High means User Counting is inflated and cross-session anything is unreliable
If either is poor, say so at the top of the write-up. It changes how every attribution number in the business should be read, and it outranks most of the property-level findings underneath it: Attribution Models.
8. Write it up
One row per finding, ordered by which reported metric it corrupts — never by ease of fix, or the trivial ones get done and the expensive ones get discussed.
| Field | Content |
|---|---|
| Event / property | The specific thing |
| Defect | What’s wrong, in one line |
| Metric affected | The number a person acts on |
| Since | From first_seen / last_seen / volume step change |
| Severity | Corrupts a reported metric > corrupts analysis > cosmetic |
| Owner | A person |
| Fix | Code, container, or plan |
Then three closing decisions, all of which get skipped and shouldn’t be:
- Amend the plan in the same pass. The plan is a finding too. Fixing the code and leaving the document stale means re-discovering all of this next year
- Decide on history. Correct it, or annotate and leave it. Either is fine; silently doing neither means the trend line lies forever — Metric Drift
- Move at least one finding upstream. Every recurring defect class has a structural fix: Schema Enforcement for malformed events, generated clients for unspecified ones, container governance for GTM additions. An audit whose only output is a fix list guarantees the same audit next year
Cadence
Quarterly on the critical path. Immediately after: a replatform, a checkout change, a consent banner change, a major redesign, a new destination, or any tag manager change made by someone outside the process.
Related: Guide - Diagnosing a Metric Movement for a single number that moved, rather than the whole surface.