Bot and Internal Traffic
Date: 2026-08-16
Traffic that isn’t a customer. It inflates sessions, crushes conversion rate, and distorts every per-session average — and the filtering you apply to remove it is itself a definition change that moves your metrics.
What it is
Bot traffic is automated requests: search crawlers, AI crawlers, monitoring tools, scrapers, security scanners, bad actors. Internal traffic is your own organisation — staff, agencies, developers, QA.
Both look like users to an analytics tool and neither behaves like one.
What it does to the numbers
real +bots effect
sessions 40,000 52,000 +30%
orders 400 400 unchanged
conversion rate 1.00% 0.77% −23%
Bots never convert, so they only ever enter the denominator. Conversion rate falls, per-session averages fall, and none of it is behaviour.
Worse, bot traffic isn’t evenly distributed. It concentrates on particular templates and particular times, so a single page can appear to have collapsed when a crawler discovered it.
Bots
Three groups, and they need different handling:
- Declared crawlers — search engines and AI crawlers with published user agents. Most tools filter the known list automatically. Note that AI crawler volume has grown substantially and the list changes — Crawling and Indexing
- Monitoring and testing — uptime checks, synthetic performance runs, your own CI. These hit the same URL on a schedule, which is a recognisable pattern
- Undeclared — scrapers and fraud tools spoofing a real user agent. These are the hard ones, and no vendor list catches them
Signatures for the third group: implausible session rates from one address, no cookie persistence across requests, perfectly regular timing, impossible viewport dimensions, and traffic patterns with no day-night cycle.
Internal traffic
Smaller in volume, disproportionate in effect — because staff behave nothing like customers. Repeated checkout tests, browsing the same product forty times, never converting, or converting with a test card.
Filtering options, in order of reliability:
- IP-based — simple, and breaks with remote working
- A cookie set by visiting a flagged URL — survives location changes, doesn’t survive clearing
- A staff account flag on identified users — reliable where they log in
- Test orders excluded by a marker on the order itself — the one that matters for revenue
Test orders are the important case. A single £0.01 test order in a revenue figure is obvious; a full-price test order isn’t, and it reconciles against nothing.
The filtering is itself a definition change
The part people miss. Turning on a bot filter changes your metrics on a specific date, in a way that looks like performance.
sessions drop 12%, conversion rate rises 14%
→ reads as a traffic problem and a conversion win
→ is one filter being enabled
Annotate it, and expect the step. Applying a filter retroactively where the tool allows it is better, because it avoids the seam entirely — Metric Drift, Annotation and Change Logs.
And there’s a subtler problem in experiments: filtering after assignment can create sample ratio mismatch. Bots split evenly at assignment and get removed unevenly, because they behave differently in each arm. The filter has to be blind to arm, or it invalidates the test — Sample Ratio Mismatch, Why Randomisation Works.
Practical position
- Enable the known-bot filter, and record the date
- Exclude internal traffic, using a cookie rather than IP
- Mark test orders at the source, in the order system, so they’re excludable everywhere rather than in each tool separately
- Watch the ratio of sessions to identified users. A rise usually means bots, not growth
- Segment out anything with no day-night cycle when a template’s traffic jumps inexplicably
- Reconcile to the order system, which bots never reach — the gap between analytics sessions and real demand is only visible from that side — Guide - Auditing a Tracking Plan