Tags: analytics web-dev concept

The Data Layer

Date: 2026-08-16


The site declares its facts once, in a structured object, so tags read data instead of scraping the DOM. Without it, every measurement breaks the next time someone changes a CSS class.


A data layer is a JavaScript object on the page — conventionally window.dataLayer — that holds structured facts about the page, user and transactions for tags to read.

The problem it solves

A tag needs to know the order value. Its options:

  • Read it from the page — document.querySelector('.order-total').innerText, then strip a £, then parse. Works until a redesign, a currency change, or a stray thin space. Fails silently, always in the direction of missing revenue
  • Read it from a declared object — the application, which already knows the number as a number, publishes it

The data layer is the second. It’s an interface boundary: the site’s job is to state what happened; the tag’s job is to decide what to do about it. Neither needs to know anything about the other, which is the entire point — marketing can change tags without a deploy, and engineering can rebuild the front end without breaking measurement, as long as the contract holds.

The contract is a tracking plan like any other. Data layer key names are a taxonomy and rot exactly like event names do — see Event Taxonomy Design.

The array-as-queue mechanism

Worth understanding properly, because it explains a class of bugs that otherwise look like magic.

<!-- before anything else, including the tag manager -->
<script>window.dataLayer = window.dataLayer || [];</script>

That’s a plain array. Pushes into it are ordinary Array.prototype.push calls — nothing is listening yet:

dataLayer.push({ event: 'login', method: 'email' });

When the tag manager library loads, it replaces push on that specific array with its own function, then walks the entries already sitting there and processes them in order. From that moment, every push is intercepted synchronously as it happens.

Consequences:

  • Pushes before the library loads are not lost — they queue, and replay in order once it arrives. This is why the async loader is safe
  • window.dataLayer = [...] destroys the mechanism. Reassigning the variable throws away the array whose push was patched. Pushes to the new array go nowhere, forever. Always push; never assign
  • Order is preserved, so a push describing state must precede the push announcing the event that depends on it
  • It’s the same pattern as gtag’s arguments-object queue and most third-party snippets — recognise it and you can debug any of them

Two models

State snapshotEvent-driven
ShapeOne object rendered into the page at loadA sequence of pushes over the page’s life
SuitsServer-rendered pages, one document per viewSingle-page apps, interactions after load
Reads as”This page is a product page for SKU 4471""A product was viewed”
WeaknessSilent on anything after loadNothing describes the page as a whole

Most real implementations are both: a snapshot at load for page-level context, then events for everything after. That’s fine, as long as the plan says which facts arrive by which route.

The persistence trap

In Google Tag Manager, the data layer is a merged model, not a series of independent messages. Each push is layered over the accumulated state, and keys from earlier pushes persist until something overwrites them.

The distinction to hold: the array is a list of messages; the model GTM reads is their accumulation. Those are different things, and only the first is what you see in the console.

WHAT YOU PUSHED                          WHAT A TAG READS AT EACH POINT

push 1  { page_type: 'product',          page_type    = 'product'
          product_id: 'A1' }             product_id   = 'A1'

push 2  { event: 'view_item' }           page_type    = 'product'
                                         product_id   = 'A1'      ← still there

push 3  { page_type: 'cart' }            page_type    = 'cart'
                                         product_id   = 'A1'      ← STILL there

Push 3 was a cart page. It said nothing about product_id, so product_id survives — and a cart event now reports the last product viewed. A key that isn’t mentioned isn’t cleared; only an explicit new value overwrites it.

Same trap in its most common form:

dataLayer.push({ event: 'view_item', ecommerce: { items: [{ item_id: 'A1' }] } });
dataLayer.push({ event: 'add_to_cart', ecommerce: { items: [{ item_id: 'B2' }] } });

The second push can leave remnants of the first object rather than replacing it cleanly. The fix is to null the key first:

dataLayer.push({ ecommerce: null });
dataLayer.push({ event: 'add_to_cart', ecommerce: { items: [{ item_id: 'B2' }] } });

Same trap, different shape: a value set on a product page — product_id, say — survives into the next pageview in a single-page app, so a checkout event inherits the last product looked at. Anything page-specific must be explicitly cleared on route change, not just repopulated, because a page that has no value for the key won’t overwrite it.

Field names for ecommerce objects are an inherited convention, not yours to choose — see Ecommerce Event Schema [CHECK: current GA4 field list against Google’s spec when that note is written].

Timing

The two failure modes are mirror images:

  • Tag fires before data is present — the tag reads undefined and either drops the hit or sends a zero. Common on client-rendered pages, where the container loads long before the API call resolves
  • Data present, nothing announces it — the state was pushed but no event key accompanied it, so nothing triggered

The rule that avoids both: push the data and the event together, at the moment the fact becomes true. Not on button click — on the successful response. Not on route change start — on render complete. A tag firing on optimistic UI shows conversions for failed payments.

For single-page apps this means the framework’s router has to be an instrumentation surface. There’s no way to sidestep it: history API navigation fires nothing a tag manager can natively hook without a shim — The History API.

PII

The data layer is world-readable. It’s in the page, in the DOM, visible in the console to anyone who opens it, and readable by every third-party script on the page including ones you didn’t audit.

Email addresses, names, order addresses and raw customer IDs put there are disclosed to every tag on the page, whatever its vendor’s data policy says. Hash before pushing if a vendor needs an identifier, and treat “just for the one tag” as false — see PII in Analytics.

The tradeoff nobody says out loud

Tag managers are sold on “deploy tracking without engineering”. A properly maintained data layer is engineering work, forever, on every feature that touches it.

The pitch isn’t wrong, it’s mis-stated. What the data layer actually buys is that the engineering happens once per fact rather than once per vendor. Twelve tools read order_value from the same push. Without it, you’re not avoiding the work — you’re paying it twelve times, in DOM selectors, and rediscovering it every redesign.