Tags: analytics commerce concept

Attribution Models

Date: 2026-08-16


An attribution model is a rule for dividing a fixed amount of credit, not a measurement of influence. Every model splits the same £200 differently, all of them are defensible, and none of them establishes that any channel caused anything.


The setup

  • Touchpoint — an interaction with a channel that got recorded: an ad click, an email open-through, an organic landing, a direct visit
  • Conversion — the outcome being credited, and its value
  • Lookback window — how far back before the conversion touchpoints are eligible for credit. Anything older is discarded, so the window silently decides how much of the path exists at all

Every model takes the same eligible path and answers one question: what fraction goes to each touchpoint. The total is always 100% of one conversion. Credit is conserved — giving paid search more necessarily takes it from something else, which is why these arguments never resolve on evidence.

One path, five answers

The input is always one table — the eligible touchpoints for one conversion:

conversion_id  order_value  touch  channel         days_before
c_9001         £200.00      1      Paid Search     14
c_9001         £200.00      2      Email            6
c_9001         £200.00      3      Organic Search   2
c_9001         £200.00      4      Direct           0

Every model is just a different credit column added to that table, and the column always sums to £200. That’s the whole mechanism — the rest is which rule fills the column.

Last click

All credit to the final touchpoint. Direct gets £200.

In practice almost every tool uses last non-direct click, on the reasoning that direct isn’t a channel that acquired anyone — it’s the absence of a recorded source. So Organic Search gets £200 and Direct gets nothing. Worth knowing, because it means direct traffic quietly donates its credit to whatever preceded it.

First click

All credit to the first eligible touchpoint. Paid Search gets £200.

Note eligible — with a 7-day lookback window, Paid Search at 14 days doesn’t exist, and “first click” becomes Email. The window changed the answer without anyone changing the model.

Linear

Equal split across all four: £50 each.

Time decay

Weight each touchpoint by recency, using a half-life — the interval over which a touchpoint’s weight halves. With a 7-day half-life:

TouchpointDaysWeightShareCredit
Paid Search140.5² = 0.2509.5%£19.07
Email60.5^0.857 = 0.55221.1%£42.11
Organic Search20.5^0.286 = 0.82031.3%£62.56
Direct00.5⁰ = 1.00038.1%£76.26
Sum 2.622100%£200.00

Working for one row: Email’s weight is 0.552; divide by the total weight 2.622 to get its share, 0.211; multiply by £200 to get £42.11.

Position-based (40/20/40)

40% to first, 40% to last, the remaining 20% split among the middle:

  • Paid Search £80 · Email £20 · Organic Search £20 · Direct £80

Side by side

ModelPaid SearchEmailOrganicDirect
Last non-direct click£0£0£200£0
First click£200£0£0£0
Linear£50£50£50£50
Time decay (7d)£19.07£42.11£62.56£76.26
Position 40/20/40£80£20£20£80

Paid search earned somewhere between £0 and £200 from this order. Nothing about the customer or the campaign is unknown — the entire spread comes from choosing a rule. If paid search is being judged on return on ad spend, that decision is being made by the model selection, not by the campaign’s performance.

Data-driven attribution

The algorithmic models try to derive the split from your data rather than assert it. Two mechanisms dominate.

Removal effect (Markov chains): build a graph of observed paths, compute the overall probability of conversion, then delete one channel — every path through it now terminates without converting — and recompute. The drop is that channel’s removal effect.

Baseline conversion probability across all paths: 5.0%. With Email removed: 3.5%.

Removal effects don’t sum to 1, because paths overlap — each channel is individually load-bearing. So they’re normalised. Say the four come out at Paid 0.30, Email 0.30, Organic 0.45, Direct 0.55, totalling 1.60:

  • Paid: 0.30 ÷ 1.60 = 18.8% → £37.50
  • Email: 0.30 ÷ 1.60 = 18.8% → £37.50
  • Organic: 0.45 ÷ 1.60 = 28.1% → £56.25
  • Direct: 0.55 ÷ 1.60 = 34.4% → £68.75

In plain terms: the model asks “how much worse would conversions look if this channel had never appeared in any path?” — and then, because those answers add up to more than the whole, scales them all down proportionally so they fit into 100%. That final scaling is a convenience, not a finding.

Shapley value: average a channel’s marginal contribution across every possible ordering of the channels present. Fairer in a specific game-theoretic sense, and combinatorially expensive, which is why it’s usually approximated.

Both share one limitation, and it’s the important one: they learn from paths that converted and paths that didn’t, in observational data. A channel that appears on converting paths because it attracts people already intending to buy — branded search, retargeting, email to existing customers — gets credited for the intent it selected on. No amount of algorithmic sophistication distinguishes selection from causation without an intervention.

What none of them do

No attribution model measures causation. They allocate observed credit under an assumption. The counterfactual question — would this order have happened anyway — is not in the data being modelled, because the data only contains what did happen.

The intuition that fails here: it feels as though a better model, given more data, would converge on the truth. It won’t, because the truth being sought was never recorded. Establishing it requires withholding exposure from a comparable group and measuring the difference — see Incrementality Testing and Marketing Mix Modelling. A cheaper cross-check, biased differently, is asking customers — Self-Reported Attribution.

The reliable finding when brands run that test against their attribution reports is that heavily-credited bottom-funnel channels are the most over-credited, because they’re the ones best at capturing intent that already existed.

Failure modes

  • Identity, not the model, is usually the problem. Pre-conversion touchpoints only receive credit if the sessions were joined to the converting user. Where Identity Stitching fails, the path collapses to the last session and the output is a plausible, systematically last-touch-biased picture. Check stitch rates before debating models
  • Direct absorbs credit it didn’t earn. Missing referrer data — app webviews, some redirects, https → http transitions — lands traffic in direct. Under last-non-direct rules it hands that credit backwards to a possibly unrelated touchpoint. See Direct Traffic and Lost Referrers
  • Walled gardens sum past 100%. Each ad platform attributes in its own model, on its own view-through and click data, with its own window, and won’t share the path. Add up their claimed conversions and you’ll exceed your actual orders. This is expected behaviour, not a bug to reconcile. See Walled Garden Reporting
  • The window is a model parameter dressed as a setting. Lengthening it manufactures touchpoints and shifts credit upper-funnel; shortening does the reverse. Chosen once, then never revisited, then reported as fact. See Attribution Windows
  • Switching models rewrites every channel’s performance at once. Budget shifts follow, and nothing in the market changed. If you switch, restate history under both models before anyone sees a number, or the transition reads as performance
  • Consent and blocking remove touchpoints, not conversions. The conversion still fires server-side while the earlier touchpoints don’t, which biases attribution towards whatever survived. Modelled Conversions is the vendors’ patch for this, with its own opacity

The one thing they’re genuinely good for

Comparing models diagnostically, rather than picking one and believing it.

Run first-click and last-click side by side. A channel that scores far higher on first click than last click is an opener — it starts paths it doesn’t finish. Reverse, and it’s a closer, harvesting demand created elsewhere.

That comparison is robust precisely because it doesn’t depend on either model being right — it uses the disagreement between them, which is real information about where in the path a channel lives. It’s also the analysis that most often reveals that a channel being cut for poor last-click return is doing most of the opening.