Crawling and Indexing
Date: 2026-08-16
Two different verbs, and confusing them wastes a lot of effort. Crawling is fetching a page; indexing is deciding to store it. A page can be crawled and not indexed, and blocking a crawl is not the same as preventing an index.
What it is
- Crawling — a bot requests a URL and receives a response
- Rendering — for some crawlers, executing the JavaScript to produce final content — Rendering and SEO
- Indexing — deciding the result is worth storing and serving
Each step can fail independently, and the diagnosis differs entirely.
discovered → crawled → rendered → indexed → ranked
│ │ │ │
│ │ │ └── "Crawled, not indexed"
│ │ └───────────── content invisible to non-rendering bots
│ └─────────────────────── blocked, 404, 5xx, timeout
└────────────────────────────────── no links point to it
Discovery
A crawler finds URLs three ways, in descending order of reliability:
- Links from pages it already knows. Real
<a href>elements. This is the primary mechanism and the one most affected by architecture - XML sitemaps. A list you provide. Useful for large catalogues and pages with few internal links; it’s a hint, not a guarantee
- External links from other sites
Orphan pages — reachable by search or filtering but linked from nowhere — are the common failure. A product only reachable through a client-rendered facet is invisible to discovery, regardless of it being in the sitemap.
robots.txt blocks crawling, not indexing
The most consequential misunderstanding in the field.
User-agent: *
Disallow: /checkout/
This says don’t fetch this. It does not say don’t index this. A blocked URL with external links pointing at it can still appear in results — with no description, because the crawler was never allowed to look.
To keep something out of the index, it must be crawlable and carry a noindex:
<meta name="robots" content="noindex, follow">Or the header equivalent, X-Robots-Tag: noindex, which works for non-HTML responses like PDFs.
Blocking in robots.txt and adding noindex together is self-defeating — the crawler can’t fetch the page, so it never sees the directive. This combination is common and produces exactly the outcome it was meant to prevent.
| Goal | Method |
|---|---|
| Save crawl budget | robots.txt disallow |
| Keep out of the index | noindex on a crawlable page |
| Keep out of the index and save budget | noindex first, wait for it to be processed, then disallow |
| Keep private | Authentication. Not either of the above |
”Crawled — currently not indexed”
The Search Console status that generates the most confusion. It means the page was fetched successfully and the engine chose not to store it. Usual reasons:
- Thin or duplicated content — near-identical product variants, or a category page with two items
- Duplicate of a page already indexed, where your canonical signals disagree — Canonicalisation
- Low perceived value relative to the rest of the site
- Crawl budget exhausted on low-value URLs before reaching this one
It is not a technical error and there’s nothing to fix in the code. The answer is either to improve the page, consolidate it, or accept it shouldn’t be indexed.
Crawl budget
The volume of crawling an engine will do on your site, roughly a function of site authority and server responsiveness.
It only matters at scale — a 500-page site will be crawled completely regardless. It matters enormously on a large catalogue with faceted navigation, where budget can be consumed almost entirely by URL combinations nobody should ever see. See Faceted Navigation and Crawl Budget.
Protecting it:
- Disallow parameter-generated URLs that add nothing
- Fix redirect chains — each hop is a crawl — Redirects and Link Equity
- Return proper 404s so dead URLs stop being retried
- Keep the server fast; slow responses reduce crawl rate directly — Time to First Byte
- Keep sitemaps accurate; a sitemap full of 404s wastes budget and erodes trust in the file
Diagnosing
- Search Console → Pages for the coverage breakdown by status
- URL Inspection for a single URL: crawled, rendered HTML, canonical chosen, index status
- Log file analysis for what bots actually request — the only source that shows where crawl budget goes
- A crawler against staging to catch orphans and broken links pre-launch