Canonicalisation
Date: 2026-08-16
Telling a search engine which of several near-identical URLs is the real one. It’s a hint, not a directive — and when your signals contradict each other, the engine picks, usually not the one you wanted.
What it is
Canonicalisation is declaring the authoritative URL for a piece of content, so duplicates consolidate onto it rather than competing with each other.
<link rel="canonical" href="https://shop.example.com/products/merino-wool-socks">Two things it does: consolidates ranking signals onto one URL, and stops near-duplicates being indexed separately.
Why duplicates happen without anyone deciding
Almost none of these are deliberate:
https://example.com/socks
https://www.example.com/socks www vs apex
http://www.example.com/socks scheme
https://www.example.com/socks/ trailing slash
https://www.example.com/Socks case
https://www.example.com/socks?sort=price sorting
https://www.example.com/socks?utm_source=x tracking
https://www.example.com/collections/wool/products/socks category path
Eight URLs, one page. Each is a separate URL to a crawler, each may accumulate its own links, and none of them individually looks as authoritative as the one page should.
Most of this is fixed at the URL Structure level with redirects. Canonicals handle what’s left — parameters and legitimately duplicated paths.
It’s a hint
The critical property. Search engines treat rel=canonical as one signal among several and will override it when the evidence disagrees — when internal links point elsewhere, when the sitemap lists a different URL, when the redirect target differs, or when the pages aren’t actually similar.
In plain terms: you’re making a suggestion, and it’s weighed against everything else your site does. If the rest of the site behaves as though a different URL is the real one, that’s what gets believed.
So the signals have to agree:
| Signal | Must say |
|---|---|
rel=canonical | the canonical URL |
| Internal links | the canonical URL |
| XML sitemap | the canonical URL, and only it |
| Redirects | resolve to the canonical URL |
hreflang | reference canonical URLs |
The most common failure is a correct canonical tag undermined by internal links pointing at the parameterised or category-path version.
Rules
- Absolute URLs, always. Relative canonicals resolve unpredictably
- Self-referencing canonicals on every page. A page whose canonical points at itself is unambiguous, and it protects against parameters someone appends later
- One canonical per page. Two tags means both are ignored
- Point at a 200. A canonical to a redirect, a 404 or a
noindexpage is a contradiction and gets discarded - Canonical to the version you’d want ranked, not to the shortest or prettiest
- Paginated pages self-canonicalise. Page 2 is not a duplicate of page 1 — its content is different, and pointing it at page 1 de-indexes real content
Canonical or noindex?
Different tools for different situations, and using the wrong one is common:
| Situation | Use |
|---|---|
| Near-duplicate that should consolidate signals | rel=canonical |
| Page that should exist for users but never be indexed | noindex, follow |
| Page that has permanently moved | 301 redirect |
| Filter combination generating infinite URLs | robots.txt disallow — Faceted Navigation and Crawl Budget |
A canonical passes signals; noindex doesn’t. For a duplicate that has attracted links, canonical is right. For an internal search results page nobody links to, noindex is right.
Cross-domain
rel=canonical works across domains — useful for syndicated content and for a marketplace listing that duplicates your product page. It’s also how a manufacturer’s copy on your site and twenty competitors’ sites gets resolved, usually not in your favour, which is the argument for writing your own product descriptions.
Diagnosing
- Search Console URL Inspection shows the “user-declared canonical” and the “Google-selected canonical” separately. When those differ, your signals are contradicting each other — that’s the single most useful diagnostic here
- A site crawl listing canonical tags will surface pages canonicalising to redirects or 404s
- Check internal links against the canonical, since that’s the usual source of disagreement