How to find duplicate and near-duplicate content on your site

Exact and near-duplicate pages, duplicate titles and descriptions. How to find them across a site and when to redirect, canonicalise or rewrite.

Most duplicate content isn’t copied from anywhere. It’s generated by the site itself: product variants with the same description, filter and sort parameters, printer-friendly versions, tag archives, location pages with one word changed. Nobody wrote it twice, but search engines see it twice.

This guide covers why duplicates matter, how to find exact and near-duplicate pages across a site, and what to do with each.

Does duplicate content hurt SEO?

Not as a penalty. Google’s canonicalization guide says “some duplicate content on a site is normal and it’s not a violation of Google’s spam policies.”

What it does do:

What causes duplicate content?

Cause Example
URL variants http and https, www and non-www, trailing slash, uppercase paths
Parameters ?sort=price, ?utm_source=…, session IDs, faceted filters
Product variants One page per colour or size with the same description
Templates with little unique text Location or service pages where only the city name changes
Archives Tag, category and author pages listing the same posts
Leftovers Staging or demo copies left accessible

How do you find duplicate content?

Exact duplicates: compare the text

Normalise each page’s visible text (lowercase it, collapse whitespace) and hash it. Pages with the same hash have identical text. This is fast and has no false positives.

Near-duplicates: compare fingerprints

Identical hashes miss pages that differ by a sentence. For those you need a similarity fingerprint. Simhash turns a page’s text into a 64-bit number built from overlapping word sequences. Similar texts produce fingerprints that differ in only a few bits, so you can compare thousands of pages without comparing every word.

Duplicate titles and descriptions

Duplicate <title> tags and meta descriptions are quicker to spot and often point straight at the duplicate templates underneath.

How Crawlens checks

Crawlens fingerprints the visible text of every page (scripts and styles removed) during the crawl, then compares indexable pages with at least 50 words:

Check How it’s detected Severity
Duplicate content Identical normalised text Warning
Near-duplicate content 64-bit simhash over 3-word sequences, at most 3 bits apart Notice
Duplicate page titles Same <title> on several pages Warning
Duplicate meta descriptions Same description on several pages Warning
Thin content Under 200 words Notice

Each flagged URL shows the page it duplicates, so you can review pairs side by side.

One thing to keep in mind when reading results: the comparison uses the whole page’s text, including navigation and footer. On pages with very little unique content, the shared template makes up most of the text, which is exactly why they show up as near-duplicates. The fix is the same either way: more unique content, or fewer pages.

What should you do with each duplicate?

The URL shouldn’t exist: redirect it. HTTP/HTTPS, www and trailing-slash variants, old URLs, staging copies. A 301 to the preferred version consolidates everything.

Both URLs need to work: canonicalise. Sorting and filter parameters, tracking codes, print versions. Keep the URL working for users and point its canonical to the main version. See canonical tag errors for the common mistakes.

The pages should stand alone: rewrite them. Location and service pages, product variants that people search for separately. Give each page content that’s genuinely specific to it: local details, specifications, reviews, FAQs.

The pages add nothing: merge or noindex them. Thin archive pages, near-empty tag pages, variants nobody searches for. Merge them into a stronger page with a redirect, or noindex them.

Checklist

Frequently asked questions

Is duplicate content a Google penalty?

No. Google says some duplicate content on a site is normal and not a violation of its spam policies. The real cost is that duplicates compete with each other and Google chooses which version to show.

What's the difference between duplicate and near-duplicate content?

Duplicate pages have identical text. Near-duplicates are almost the same, for example product variants that differ only by colour or size, or location pages where only the city name changes.

How do you detect near-duplicate pages?

Compare a fingerprint of each page's text instead of the raw text. Simhash builds a 64-bit fingerprint from overlapping word sequences; pages whose fingerprints differ in only a few bits have mostly the same text.

Should I use a canonical or a 301 redirect for duplicates?

Use a 301 when the duplicate URL doesn't need to exist for users. Use a canonical when both URLs need to work, like sorting parameters or tracking codes, but only one should be indexed.

Are duplicate titles and meta descriptions a problem?

They don't stop pages being indexed, but they make different pages look identical in search results and often point to duplicate or thin pages underneath.

· Founder, Crawlens

Dien builds Crawlens, a desktop crawler for technical SEO audits, and writes about the checks it runs: crawling, indexing, JavaScript rendering and Search Console data.

Audit your own site with Crawlens

Crawl, run 70 checks, and see them next to Search Console and Core Web Vitals data.

Download free for Windows