Schema Enforcement
Date: 2026-08-16
Validating events against the tracking plan at the point of collection, and doing something about the failures. It’s the only enforcement that catches events your codebase never knew about — which is most of the ones that break things.
What it is
Schema enforcement is checking each incoming event against a declared specification — the event exists, required properties are present, types and allowed values match — and routing failures somewhere visible.
It’s the last of the enforcement ladder in Tracking Plans, and the only rung that sees everything:
| Enforcement | Catches | Misses |
|---|---|---|
| Review | Whatever anyone remembers | Everything else |
| CI validation | New violations in your code | Existing decay, third parties |
| Codegen | Unspecified events at authoring time | Anything not in your codebase |
| Collection-side | Everything that arrives | Nothing |
A tag someone added in the container last Tuesday is invisible to the first three. It arrives at the fourth like everything else.
What to do with a failure
Three options, and the choice matters more than the validation itself:
Reject. Return an error, drop the event. Cleanest data, and you lose events permanently when the schema is wrong rather than the event. Dangerous in production.
Quarantine. Accept it, store it in a separate table, keep it out of reporting. The right default — nothing is lost, reporting stays clean, and the quarantine table is a live list of what’s broken.
Pass with a warning. Store it and flag it. Weakest, because a warning nobody reads is no enforcement at all — but appropriate during a rollout, before you know what will fail.
event arrives
├─ valid → main table
├─ unknown event → quarantine, alert if volume is high
├─ missing field → quarantine
└─ wrong type → quarantine
What the schema should check
Roughly in order of value:
- Event name is known. Catches typos, casing drift, and undocumented tags — Event Taxonomy Design
- Required properties present. The most common real failure
- Types correct.
"true"versustrue, string versus number - Allowed values.
currencyin a closed list;stepwithin range - Plausibility. Negative revenue, zero-value purchases, absurd quantities
- No personal data. Pattern-match for anything email- or phone-shaped. This one is a disclosure control rather than a quality check — PII in Analytics
Generating it from the plan
The schema and the tracking plan must be the same artefact, or they drift and you’re enforcing a stale specification.
tracking plan (YAML in the repo)
├─▶ generates the typed client → compile-time enforcement
└─▶ generates the collection schema → runtime enforcement
One source, two outputs. That’s what makes the plan structurally unable to drift, and it’s the version of “a plan with enforcement” that actually holds — see the enforcement ladder in Tracking Plans.
Watching the quarantine
The quarantine table is the most useful diagnostic surface in the pipeline, and it’s usually ignored:
- A new event name appearing means someone shipped something undocumented — find the owner
- A spike in one failure type means a release broke something, and the timestamp dates it
- Steady low-level failures are a property nobody fixed, quietly missing from a third of events
Alert on volume, not on individual failures. A handful of malformed events daily is normal; a thousand overnight is an incident.
Where it goes wrong
- Too strict too early. Enforcing before the plan matches reality drops legitimate events. Run in warn-only mode first, read the output, then tighten
- Nobody watching the quarantine, which turns enforcement into silent data loss with extra steps
- Schema maintained separately from the plan, so it enforces last year’s specification
- Blocking third parties entirely. Vendor tags fire events you didn’t specify; the answer is a permissive namespace for them, not rejecting everything unrecognised