Tags: analytics concept

Path Analysis

Date: 2026-08-17


Looking at the actual routes people take through a site rather than a funnel you defined in advance. The output is almost always unreadable — real behaviour branches so hard that the top path is a rounding error — and the useful version answers a narrow question instead of drawing the whole graph.


Path analysis is reconstructing the observed sequences of pages or events users actually took, without a predefined sequence to measure against.

Why the full graph is useless

paths from product page to purchase, one week

path                                              sessions    % of purchases
PDP → basket → checkout → confirm                    1,204          18.2%
PDP → PDP → basket → checkout → confirm                612           9.3%
PDP → search → PDP → basket → checkout → confirm       288           4.4%
PDP → basket → PDP → basket → checkout → confirm       201           3.0%
PDP → category → PDP → basket → checkout → confirm     187           2.8%
… 4,100 further distinct paths                       4,118          62.3%

top 5 paths account for 37.7% of purchases.
the remaining 62% is spread across four thousand routes,
none above 1%.

This is the normal result, not a sign of a broken site. With five page types and a typical session length, the number of possible orderings is enormous, and real users backtrack, compare, open tabs and return days later. Any visualisation of it is a hairball — the Sankey diagrams that path tools produce look impressive and support no decision.

The four questions that are actually answerable

Narrow the question and path data becomes useful.

1. What comes immediately before or after one page?

next page after /delivery-information

/checkout               34%   ← they were reassured, came back to buy
/basket                 18%
exit                    22%   ← they weren't
/returns-policy         11%   ← delivery worry is really a returns worry
/contact                 8%

One step, one page, ranked. Actionable: the returns link suggests what the page should also answer.

2. What do people do instead of the thing you wanted?

on the basket page, before exiting without purchase

opened the delivery calculator     41%   ← cost surprise
returned to a PDP                  28%   ← second-guessing the product
opened the voucher field           19%   ← went looking for a code
                                          and didn't come back

The voucher field row is the classic finding, and it’s one you’d never see in a defined funnel — Form Analytics.

3. What are the loops? Repeated cycling between two pages is a signal that something isn’t being answered — PDP ↔ size guide, basket ↔ delivery info, search ↔ results.

4. Which entry points lead where? Landing page to next page, which is a smaller graph and genuinely readable.

Where the data betrays you

The technical caveats that invalidate more path analyses than the readability problem does:

  • Page views aren’t the journey on a single-page app. Virtual page views need explicit instrumentation, and the order they fire in isn’t guaranteed to match navigation — Page Views vs Events
  • Sessions cut paths arbitrarily. A 30-minute timeout splits a lunchtime browse and an evening purchase into two sessions, so the path to purchase looks like it started at the confirmation step — Sessionisation
  • Cross-device journeys are invisible unless the user is logged in throughout. Research on mobile, buy on desktop is one of the commonest patterns in commerce and appears as two unrelated sessions — Cross-Device Tracking
  • Back-button navigation often doesn’t register, so the observed path is smoother than the real one
  • Tabs. Comparison shopping across four open tabs serialises into a nonsense sequence
  • Bots produce clean, fast, unusual paths that rise to the top of any ranking — Bot and Internal Traffic, Sample Pollution

Path grouping, which is what makes it work

The single technique that turns unreadable into readable: analyse page types, not URLs.

BEFORE                                AFTER

/product/blue-widget-large            PDP → PDP → BASKET → CHECKOUT
/product/red-widget-small             PDP → SEARCH → PDP → BASKET → CHECKOUT
/category/widgets?colour=blue&sort=…
/basket                               top 5 groups now cover 71% of purchases
/checkout/delivery
                                      and the remaining tail is
4,100 distinct paths                  interpretable

Collapse by template, strip query parameters, and merge consecutive repeats of the same type into one step (PDP → PDP → PDP becomes PDP ×3, which preserves the signal without exploding the path count). This is nearly always the difference between a useful path analysis and an abandoned one — URL Structure, Cardinality.

Reading it honestly

Path analysis is descriptive. It tells you what people did, never why, and never what would happen if you changed something.

The standard trap: “users who visit the size guide convert 2.4× better, so we should promote the size guide.” Users who check a size guide are further along in deciding — the page didn’t cause the intent, it accompanied it. Promoting it to everyone doesn’t transfer the conversion rate, and the only way to find out is to test it — Correlation and Causation, A-B Tests.

Use it to generate hypotheses, then test them. That’s the whole role, and it’s a valuable one — path data is unusually good at surfacing behaviour nobody anticipated, which is exactly what’s hard to get from a funnel you designed.

Where it interacts

  • Funnel Analysis — the defined-path counterpart; funnels answer “how many got through”, paths answer “what else did they do”
  • Session Replay — the qualitative version of the same question, and much faster for understanding why a loop exists
  • Sessionisation — the definition that determines where paths start and stop, and therefore what any of this means
  • Search and Findability — search-heavy paths are usually a navigation finding rather than a search one