Path Analysis
Date: 2026-08-17
Looking at the actual routes people take through a site rather than a funnel you defined in advance. The output is almost always unreadable — real behaviour branches so hard that the top path is a rounding error — and the useful version answers a narrow question instead of drawing the whole graph.
Path analysis is reconstructing the observed sequences of pages or events users actually took, without a predefined sequence to measure against.
Why the full graph is useless
paths from product page to purchase, one week
path sessions % of purchases
PDP → basket → checkout → confirm 1,204 18.2%
PDP → PDP → basket → checkout → confirm 612 9.3%
PDP → search → PDP → basket → checkout → confirm 288 4.4%
PDP → basket → PDP → basket → checkout → confirm 201 3.0%
PDP → category → PDP → basket → checkout → confirm 187 2.8%
… 4,100 further distinct paths 4,118 62.3%
top 5 paths account for 37.7% of purchases.
the remaining 62% is spread across four thousand routes,
none above 1%.
This is the normal result, not a sign of a broken site. With five page types and a typical session length, the number of possible orderings is enormous, and real users backtrack, compare, open tabs and return days later. Any visualisation of it is a hairball — the Sankey diagrams that path tools produce look impressive and support no decision.
The four questions that are actually answerable
Narrow the question and path data becomes useful.
1. What comes immediately before or after one page?
next page after /delivery-information
/checkout 34% ← they were reassured, came back to buy
/basket 18%
exit 22% ← they weren't
/returns-policy 11% ← delivery worry is really a returns worry
/contact 8%
One step, one page, ranked. Actionable: the returns link suggests what the page should also answer.
2. What do people do instead of the thing you wanted?
on the basket page, before exiting without purchase
opened the delivery calculator 41% ← cost surprise
returned to a PDP 28% ← second-guessing the product
opened the voucher field 19% ← went looking for a code
and didn't come back
The voucher field row is the classic finding, and it’s one you’d never see in a defined funnel — Form Analytics.
3. What are the loops? Repeated cycling between two pages is a signal that something isn’t being answered — PDP ↔ size guide, basket ↔ delivery info, search ↔ results.
4. Which entry points lead where? Landing page to next page, which is a smaller graph and genuinely readable.
Where the data betrays you
The technical caveats that invalidate more path analyses than the readability problem does:
- Page views aren’t the journey on a single-page app. Virtual page views need explicit instrumentation, and the order they fire in isn’t guaranteed to match navigation — Page Views vs Events
- Sessions cut paths arbitrarily. A 30-minute timeout splits a lunchtime browse and an evening purchase into two sessions, so the path to purchase looks like it started at the confirmation step — Sessionisation
- Cross-device journeys are invisible unless the user is logged in throughout. Research on mobile, buy on desktop is one of the commonest patterns in commerce and appears as two unrelated sessions — Cross-Device Tracking
- Back-button navigation often doesn’t register, so the observed path is smoother than the real one
- Tabs. Comparison shopping across four open tabs serialises into a nonsense sequence
- Bots produce clean, fast, unusual paths that rise to the top of any ranking — Bot and Internal Traffic, Sample Pollution
Path grouping, which is what makes it work
The single technique that turns unreadable into readable: analyse page types, not URLs.
BEFORE AFTER
/product/blue-widget-large PDP → PDP → BASKET → CHECKOUT
/product/red-widget-small PDP → SEARCH → PDP → BASKET → CHECKOUT
/category/widgets?colour=blue&sort=…
/basket top 5 groups now cover 71% of purchases
/checkout/delivery
and the remaining tail is
4,100 distinct paths interpretable
Collapse by template, strip query parameters, and merge consecutive repeats of the same type into one step (PDP → PDP → PDP becomes PDP ×3, which preserves the signal without exploding the path count). This is nearly always the difference between a useful path analysis and an abandoned one — URL Structure, Cardinality.
Reading it honestly
Path analysis is descriptive. It tells you what people did, never why, and never what would happen if you changed something.
The standard trap: “users who visit the size guide convert 2.4× better, so we should promote the size guide.” Users who check a size guide are further along in deciding — the page didn’t cause the intent, it accompanied it. Promoting it to everyone doesn’t transfer the conversion rate, and the only way to find out is to test it — Correlation and Causation, A-B Tests.
Use it to generate hypotheses, then test them. That’s the whole role, and it’s a valuable one — path data is unusually good at surfacing behaviour nobody anticipated, which is exactly what’s hard to get from a funnel you designed.
Where it interacts
- Funnel Analysis — the defined-path counterpart; funnels answer “how many got through”, paths answer “what else did they do”
- Session Replay — the qualitative version of the same question, and much faster for understanding why a loop exists
- Sessionisation — the definition that determines where paths start and stop, and therefore what any of this means
- Search and Findability — search-heavy paths are usually a navigation finding rather than a search one