Data Quality Monitoring
Date: 2026-08-17
Testing the pipeline the way you’d test code — automatically, on every run, with a failure that stops something. It’s the difference between finding out data was wrong when a report looks odd three weeks later, and finding out when the job that produced it refused to publish.
Data quality monitoring is automated tests on data as it moves through a pipeline — freshness, volume, schema, values — that alert or halt the job when one fails.
The six dimensions
A vocabulary worth having, because “the data’s wrong” covers six different failures with different fixes.
| Dimension | Question | Typical check |
|---|---|---|
| Freshness | Is it up to date? | Latest timestamp within N hours |
| Volume | Is it all there? | Row count within expected range for this weekday |
| Schema | Is it shaped right? | Expected columns, expected types, no silent additions |
| Validity | Are values legal? | Currency in a known set, quantity > 0, enum membership |
| Uniqueness | Duplicates? | Primary key actually unique — Idempotency and Deduplication |
| Consistency | Do related things agree? | Order total = sum of line items + shipping − discount |
Freshness and volume catch the most, for the least effort. A pipeline that didn’t run, or ran and produced a tenth of the usual rows, accounts for a large share of real incidents and takes ten minutes to check.
Where the checks go
COLLECTION schema validation at the SDK / collector
→ reject or quarantine malformed events at the edge
Schema Enforcement
↓
LANDING freshness + volume on raw arrival
→ "did today's data arrive, and is there enough of it"
↓
TRANSFORMATION validity, uniqueness, consistency, referential integrity
→ THE MAIN LAYER. tests run as part of the job
↓
SERVING reconciliation against source of truth
→ warehouse revenue vs the finance system
↓
CONSUMPTION the business-metric alerts — Alerting on Metrics
The last layer is where business-metric thresholds live — Alerting on Metrics — and it’s the layer you want to be catching almost nothing, because everything above it should have fired first.
The further right a problem is caught, the more it has already contaminated. A bad value caught at collection affects nothing; the same value caught at consumption is already in three dashboards, a board deck and a forecast.
Assertions as part of the job
The practice that changes things: checks live with the transformation and fail the run, rather than existing as a separate monitoring dashboard someone was supposed to look at.
-- runs after the orders model builds; a failure blocks publication
select 'null_order_id' as check, count(*) as failures
from orders where order_id is null
union all
select 'duplicate_order', count(*) - count(distinct order_id) from orders
union all
select 'negative_total', count(*) from orders where total_pence < 0
union all
select 'total_mismatch', count(*) from orders o
where abs(o.total_pence - (
select coalesce(sum(li.line_total_pence), 0)
from order_lines li where li.order_id = o.order_id
) - o.shipping_pence + o.discount_pence) > 1 -- 1p rounding tolerance
union all
select 'future_dated', count(*) from orders where created_at > now();The rounding tolerance is the line that stops this being abandoned. A consistency check with zero tolerance fails on legitimate penny rounding within a week, gets muted, and then never catches the real £4,000 discrepancy — Revenue Metrics.
Blocking versus warning
Not every failure should stop the pipeline, and getting this wrong in either direction kills the practice.
BLOCK — stop, don't publish WARN — publish, notify
data is unusable or misleading quality is degraded but usable
· zero rows · volume 15% below expected
· schema changed incompatibly · a new event name appeared
· primary key not unique · null rate on an optional field up
· revenue negative or 100× off · one dimension's cardinality growing
· yesterday's data missing
consequence: stale data, which consequence: a ticket
is honest Cardinality
A dimension’s distinct values climbing steadily is the warn-level check most worth having, because it’s how a taxonomy quietly degrades rather than breaks — Cardinality.
Stale is better than wrong. A dashboard showing yesterday’s number with a “data as of” stamp is recoverable; one showing today’s wrong number gets acted on. This is the single most important principle here and the one that’s hardest to hold when someone senior wants the figure now.
Reconciliation
The check that catches what internal consistency can’t: compare against an independent source of truth.
warehouse finance system delta
orders yesterday 4,182 4,196 −14 (−0.33%)
revenue £312,447 £313,890 −£1,443 (−0.46%)
expected drift: refunds processed after the cut-off, test orders excluded
one side only, timezone boundary — Timezones and Date Boundaries
tolerance: ±1%. beyond that, investigate.
Most of the expected drift has a boring cause, and a day boundary interpreted differently by two systems is the most common of them — Timezones and Date Boundaries.
Set and document a tolerance, and alert on the tolerance being exceeded — not on the delta existing. Analytics and finance will never agree exactly, and a team chasing perfect reconciliation will stop reconciling. What matters is that the gap is stable and explained; a gap that widens is the signal — Tool Discrepancies.
Making it stick
- Own the checks with the model. Written by whoever writes the transformation, in the same repository, reviewed in the same pull request — Code Review
- Test the tests. Break something deliberately in a non-production environment and confirm the check fires
- Record every failure, including muted ones. The pattern over months tells you where the pipeline is actually fragile
- A data contract with upstream — the schema and its guarantees, agreed and versioned, so a producer changing a field is a breaking change rather than a surprise — Schema Enforcement, Backwards Compatibility
- Publish freshness and status to consumers, so people know whether to trust today’s dashboard without asking
Where it interacts
- Schema Enforcement — prevention at collection, which is cheaper than any detection downstream
- Anomaly Detection — the same statistics applied to pipeline metrics rather than business ones
- Instrumentation Debugging — what happens after a check fails and the cause is in the browser rather than the pipeline
- Warehouse-First Analytics — a warehouse makes all of this possible, and makes it necessary, because nothing else is validating the data