Every site has duplicate URLs. Sorting and filtering parameters, tracking codes, HTTP and HTTPS versions, trailing slashes, uppercase paths: the same page often answers at a dozen addresses. The canonical tag is how you tell search engines which one is the real page.
When canonicals are right, nobody notices. When they’re wrong, pages drop out of the index or the wrong version ranks, and it’s hard to spot because nothing looks broken in the browser. This guide covers the common canonical errors and how to audit them across a whole site.
What does a canonical tag do?
A canonical tag is a <link> in the page’s <head>:
<link rel="canonical" href="https://example.com/shoes/" />
It tells search engines: of all the URLs showing this content, index this one. Google’s guide to consolidating duplicate URLs lists the methods in order of strength:
| Method | Signal strength |
|---|---|
| Redirect | Strong: the target should become canonical |
rel="canonical" link |
Strong: the specified URL should become canonical |
| Listed in the XML sitemap | Weak: helps sitemap URLs become canonical |
Strong isn’t absolute. Google’s canonicalization overview says you can indicate your preference, but “Google may choose a different page as canonical than you do.” Most canonical errors are really conflicting signals.
Duplicate URLs on their own aren’t a problem to panic about. The same page says “some duplicate content on a site is normal and it’s not a violation of Google’s spam policies.” The goal is to make the preferred version obvious.
What are the most common canonical errors?
1. Canonical points to a redirect or an error
The canonical names a URL that redirects, returns 4xx or 5xx, or is blocked by robots.txt. Search engines usually ignore it and pick a canonical themselves. This often happens after a migration: pages still point their canonical at the old URL structure.
Fix: point the canonical at the final, 200-status URL.
2. Canonical points to a noindex page
Page A says “index page B instead”, and page B says “don’t index me”. The signals cancel out, and both pages can drop out of the index.
Fix: canonicalise to an indexable page, or remove the noindex from the target. Google also recommends against using noindex to steer canonical selection within a site.
3. More than one canonical
Two <link rel="canonical"> tags in the HTML, often one from the theme and one from an SEO plugin, or one in the HTML and another in an HTTP Link header. When they disagree, search engines may ignore them all.
Fix: keep exactly one. Google supports both the HTML tag and the HTTP header, but recommends choosing one, because using both is more error-prone.
4. No canonical at all
Without a canonical, Google chooses among duplicate and parameterised URLs itself. That usually works, but not always, especially with tracking parameters or faceted navigation.
Fix: add a self-referencing canonical to every indexable page, using an absolute URL.
5. Canonicals that disagree with other signals
The canonical points to one URL while the sitemap lists another, internal links point to a third, and hreflang names a fourth. Google’s guidance is explicit: don’t specify different URLs as canonical for the same page using different methods.
Fix: make every signal agree. Canonicals, sitemaps, internal links, hreflang and redirects should all use the same final URL.
6. JavaScript that changes the canonical
The raw HTML has one canonical and JavaScript replaces it with another. Google’s JavaScript guidance says not to change the canonical to something different from the original HTML. See JavaScript SEO: raw vs rendered HTML.
What shouldn’t you use for canonicalization?
From the same Google documentation:
- robots.txt: blocking a duplicate stops Google from seeing it’s a duplicate.
- The URL removal tool: it hides all versions of a URL from Search.
- URL fragments (
#section) as the canonical. - Different canonicals via different methods for the same page.
How do you audit canonicals across a site?
A crawler that records each page’s canonical, then crawls the canonical target to check its status, can catch all of these at once. Crawlens reads canonicals from both the HTML and the HTTP Link header and runs these checks:
| Check | What it flags | Severity |
|---|---|---|
| Canonical points to a non-200 URL | Canonicals to redirects, errors or URLs blocked by robots.txt | Critical |
| Canonical points to a noindex page | Canonicals to pages that carry noindex | Critical |
| Multiple canonical tags | More than one canonical, counting HTML and HTTP header | Warning |
| Missing canonical tag | Indexable pages with no canonical | Notice |
| Canonicalised pages | Pages pointing their canonical elsewhere, so they won’t be indexed | Notice |
The last one isn’t an error. It’s a list to review: confirm each canonicalised page really should defer to its target. With JavaScript rendering on, JavaScript changes SEO tags also flags canonicals rewritten by scripts.
Checklist
- One canonical per page, in the HTML or the HTTP header, not both
- Absolute URLs, no fragments
- Every canonical target returns 200 and is indexable
- Indexable pages have a self-referencing canonical
- Canonicals, sitemaps, internal links, hreflang and redirects use the same final URL
- No canonical changes made by JavaScript
- Canonicalised pages reviewed: each should really defer to its target
Frequently asked questions
What is a canonical tag?
It's a link element, rel="canonical", that tells search engines which URL is the main version of a page when the same or very similar content is available at several URLs. Google treats it as a strong signal, but not a directive.
Does Google always follow the canonical tag?
No. Google says you can indicate your preference, but it may choose a different page as canonical. Conflicting signals, like a canonical to a redirect or a sitemap listing a different URL, make that more likely.
Should every page have a self-referencing canonical?
It's good practice for indexable pages. Without one, Google picks the canonical itself among duplicate or parameterised URLs, such as tracking or sorting parameters.
Can I use noindex and a canonical together?
Avoid pointing a canonical at a noindex page; it sends conflicting signals and the page may drop out entirely. Google also doesn't recommend using noindex to steer canonical selection within a site.
Can I set the canonical in an HTTP header?
Yes, with a Link header, which is useful for PDFs and other non-HTML files. Google supports both but recommends choosing one method, because using both at once is more error-prone.