Tags: web-dev concept

Canonicalisation

Date: 2026-08-16


Telling a search engine which of several near-identical URLs is the real one. It’s a hint, not a directive — and when your signals contradict each other, the engine picks, usually not the one you wanted.


What it is

Canonicalisation is declaring the authoritative URL for a piece of content, so duplicates consolidate onto it rather than competing with each other.

<link rel="canonical" href="https://shop.example.com/products/merino-wool-socks">

Two things it does: consolidates ranking signals onto one URL, and stops near-duplicates being indexed separately.

Why duplicates happen without anyone deciding

Almost none of these are deliberate:

https://example.com/socks
https://www.example.com/socks              www vs apex
http://www.example.com/socks               scheme
https://www.example.com/socks/             trailing slash
https://www.example.com/Socks              case
https://www.example.com/socks?sort=price   sorting
https://www.example.com/socks?utm_source=x tracking
https://www.example.com/collections/wool/products/socks   category path

Eight URLs, one page. Each is a separate URL to a crawler, each may accumulate its own links, and none of them individually looks as authoritative as the one page should.

Most of this is fixed at the URL Structure level with redirects. Canonicals handle what’s left — parameters and legitimately duplicated paths.

It’s a hint

The critical property. Search engines treat rel=canonical as one signal among several and will override it when the evidence disagrees — when internal links point elsewhere, when the sitemap lists a different URL, when the redirect target differs, or when the pages aren’t actually similar.

In plain terms: you’re making a suggestion, and it’s weighed against everything else your site does. If the rest of the site behaves as though a different URL is the real one, that’s what gets believed.

So the signals have to agree:

SignalMust say
rel=canonicalthe canonical URL
Internal linksthe canonical URL
XML sitemapthe canonical URL, and only it
Redirectsresolve to the canonical URL
hreflangreference canonical URLs

The most common failure is a correct canonical tag undermined by internal links pointing at the parameterised or category-path version.

Rules

  • Absolute URLs, always. Relative canonicals resolve unpredictably
  • Self-referencing canonicals on every page. A page whose canonical points at itself is unambiguous, and it protects against parameters someone appends later
  • One canonical per page. Two tags means both are ignored
  • Point at a 200. A canonical to a redirect, a 404 or a noindex page is a contradiction and gets discarded
  • Canonical to the version you’d want ranked, not to the shortest or prettiest
  • Paginated pages self-canonicalise. Page 2 is not a duplicate of page 1 — its content is different, and pointing it at page 1 de-indexes real content

Canonical or noindex?

Different tools for different situations, and using the wrong one is common:

SituationUse
Near-duplicate that should consolidate signalsrel=canonical
Page that should exist for users but never be indexednoindex, follow
Page that has permanently moved301 redirect
Filter combination generating infinite URLsrobots.txt disallow — Faceted Navigation and Crawl Budget

A canonical passes signals; noindex doesn’t. For a duplicate that has attracted links, canonical is right. For an internal search results page nobody links to, noindex is right.

Cross-domain

rel=canonical works across domains — useful for syndicated content and for a marketplace listing that duplicates your product page. It’s also how a manufacturer’s copy on your site and twenty competitors’ sites gets resolved, usually not in your favour, which is the argument for writing your own product descriptions.

Diagnosing

  • Search Console URL Inspection shows the “user-declared canonical” and the “Google-selected canonical” separately. When those differ, your signals are contradicting each other — that’s the single most useful diagnostic here
  • A site crawl listing canonical tags will surface pages canonicalising to redirects or 404s
  • Check internal links against the canonical, since that’s the usual source of disagreement