Faceted Navigation and Crawl Budget
Date: 2026-08-16
Filters multiply. Five facets with a handful of values each generate more URL combinations than you have products, all of them crawlable, almost none of them worth indexing — and a crawler will happily spend its entire budget on them.
What it is
Faceted navigation is filtering a listing by attributes — size, colour, brand, price, material. Crawl budget is how much crawling an engine will do on your site.
The collision: every facet combination is a distinct URL, and the count is combinatorial.
The arithmetic
A category with five facets:
size 6 values
colour 8 values
brand 12 values
material 5 values
price 6 bands
combinations = 7 × 9 × 13 × 6 × 7 (each facet + "not applied")
= 34,398 URLs
× sort options (4) = 137,592
× pagination (avg 3 pages) = 412,776
412,776 crawlable URLs for one category containing perhaps 400 products. Add ten categories and a crawler could spend months without reaching your new arrivals.
Worse, the URLs are largely duplicative — ?colour=navy&size=8 and ?size=8&colour=navy are different URLs with identical content unless parameter order is normalised.
What it costs
- Crawl budget consumed on combinations nobody searches for, so genuinely new products are discovered late — Crawling and Indexing
- Index bloat — thousands of thin, near-identical pages, which dilutes site quality signals
- Duplicate content competing with the clean category page
- Server load from bots requesting hundreds of thousands of dynamically-generated listings
- Analytics fragmentation, since each combination is a separate page path
Handling it
There’s no single answer — it’s a per-facet decision, and the useful framing is: would anyone search for this combination?
| Facet type | Example | Treatment |
|---|---|---|
| Genuinely searched | ”womens wool socks” | Index it. Give it a real path, its own title and copy |
| Plausible, low volume | ”navy wool socks” | Crawlable, noindex, follow — links pass, page doesn’t compete |
| Combinatorial noise | size + colour + price band | Block in robots.txt, or don’t generate a URL at all |
| Sort and view options | ?sort=price | Never index. Canonical to the unsorted URL |
| Pagination | ?page=2 | Self-canonical, indexable. Real distinct content |
The mechanisms, and the order to reach for them:
- Don’t create the URL. Apply filters client-side without changing the URL, or use a
#fragment, which crawlers ignore. The cleanest solution and it costs shareability of filtered views robots.txtdisallow on the parameter patterns. Saves crawl budget outright. Remember this blocks crawling, not indexing — a blocked URL with external links can still appearnoindex, followon generated combinations you want crawled for discovery but not indexed. Costs budget, preserves link flow- Canonical to the parent for sorted and view variants — Canonicalisation
nofollowon facet links — a partial measure; it discourages rather than prevents- Promote the valuable few. Take the handful of genuinely searched combinations and give them real static URLs with unique content
The combination that actually works on a large catalogue: block the noise in robots.txt, noindex the plausible-but-low-value tier, and hand-build static landing pages for the ten or twenty combinations with real search demand.
Getting the list right
Don’t guess which combinations matter. The evidence is available:
- Search Console — which filtered URLs already receive impressions
- Internal site search — what people type, which is the closest signal to demand
- Keyword research — external search volume for attribute combinations
- Log files — which URLs bots actually spend their budget on, which is the only source that shows the waste directly
Checking whether you have a problem
- Search Console → Pages: a large “Crawled — currently not indexed” or “Discovered — currently not indexed” count is the signature
- Crawl stats: requests far exceeding your real page count
site:search returning far more URLs than you have products- Log analysis: the proportion of bot requests hitting parameterised URLs. On an unmanaged site this is routinely the majority
Related: URL Structure for keeping parameters out of indexable paths in the first place, and Technical SEO for where this sits.