Tags: web-dev concept

Faceted Navigation and Crawl Budget

Date: 2026-08-16


Filters multiply. Five facets with a handful of values each generate more URL combinations than you have products, all of them crawlable, almost none of them worth indexing — and a crawler will happily spend its entire budget on them.


What it is

Faceted navigation is filtering a listing by attributes — size, colour, brand, price, material. Crawl budget is how much crawling an engine will do on your site.

The collision: every facet combination is a distinct URL, and the count is combinatorial.

The arithmetic

A category with five facets:

size      6 values
colour    8 values
brand    12 values
material  5 values
price     6 bands

combinations = 7 × 9 × 13 × 6 × 7          (each facet + "not applied")
             = 34,398 URLs

× sort options (4)          = 137,592
× pagination (avg 3 pages)  = 412,776

412,776 crawlable URLs for one category containing perhaps 400 products. Add ten categories and a crawler could spend months without reaching your new arrivals.

Worse, the URLs are largely duplicative — ?colour=navy&size=8 and ?size=8&colour=navy are different URLs with identical content unless parameter order is normalised.

What it costs

  • Crawl budget consumed on combinations nobody searches for, so genuinely new products are discovered late — Crawling and Indexing
  • Index bloat — thousands of thin, near-identical pages, which dilutes site quality signals
  • Duplicate content competing with the clean category page
  • Server load from bots requesting hundreds of thousands of dynamically-generated listings
  • Analytics fragmentation, since each combination is a separate page path

Handling it

There’s no single answer — it’s a per-facet decision, and the useful framing is: would anyone search for this combination?

Facet typeExampleTreatment
Genuinely searched”womens wool socks”Index it. Give it a real path, its own title and copy
Plausible, low volume”navy wool socks”Crawlable, noindex, follow — links pass, page doesn’t compete
Combinatorial noisesize + colour + price bandBlock in robots.txt, or don’t generate a URL at all
Sort and view options?sort=priceNever index. Canonical to the unsorted URL
Pagination?page=2Self-canonical, indexable. Real distinct content

The mechanisms, and the order to reach for them:

  1. Don’t create the URL. Apply filters client-side without changing the URL, or use a #fragment, which crawlers ignore. The cleanest solution and it costs shareability of filtered views
  2. robots.txt disallow on the parameter patterns. Saves crawl budget outright. Remember this blocks crawling, not indexing — a blocked URL with external links can still appear
  3. noindex, follow on generated combinations you want crawled for discovery but not indexed. Costs budget, preserves link flow
  4. Canonical to the parent for sorted and view variants — Canonicalisation
  5. nofollow on facet links — a partial measure; it discourages rather than prevents
  6. Promote the valuable few. Take the handful of genuinely searched combinations and give them real static URLs with unique content

The combination that actually works on a large catalogue: block the noise in robots.txt, noindex the plausible-but-low-value tier, and hand-build static landing pages for the ten or twenty combinations with real search demand.

Getting the list right

Don’t guess which combinations matter. The evidence is available:

  • Search Console — which filtered URLs already receive impressions
  • Internal site search — what people type, which is the closest signal to demand
  • Keyword research — external search volume for attribute combinations
  • Log files — which URLs bots actually spend their budget on, which is the only source that shows the waste directly

Checking whether you have a problem

  • Search Console → Pages: a large “Crawled — currently not indexed” or “Discovered — currently not indexed” count is the signature
  • Crawl stats: requests far exceeding your real page count
  • site: search returning far more URLs than you have products
  • Log analysis: the proportion of bot requests hitting parameterised URLs. On an unmanaged site this is routinely the majority

Related: URL Structure for keeping parameters out of indexable paths in the first place, and Technical SEO for where this sits.