Most duplicate content isn’t copied from anywhere. It’s generated by the site itself: product variants with the same description, filter and sort parameters, printer-friendly versions, tag archives, location pages with one word changed. Nobody wrote it twice, but search engines see it twice.
This guide covers why duplicates matter, how to find exact and near-duplicate pages across a site, and what to do with each.
Does duplicate content hurt SEO?
Not as a penalty. Google’s canonicalization guide says “some duplicate content on a site is normal and it’s not a violation of Google’s spam policies.”
What it does do:
- Pages compete with each other. Several URLs with the same content can split links and relevance between them.
- Google picks the version. Among duplicates, Google chooses one canonical, and it may not be the one you’d choose.
- Crawling goes to copies. On large sites, crawlers spend time on duplicates instead of new or changed pages.
- Near-duplicates look thin. Pages that are almost the same add little value on their own.
What causes duplicate content?
| Cause | Example |
|---|---|
| URL variants | http and https, www and non-www, trailing slash, uppercase paths |
| Parameters | ?sort=price, ?utm_source=…, session IDs, faceted filters |
| Product variants | One page per colour or size with the same description |
| Templates with little unique text | Location or service pages where only the city name changes |
| Archives | Tag, category and author pages listing the same posts |
| Leftovers | Staging or demo copies left accessible |
How do you find duplicate content?
Exact duplicates: compare the text
Normalise each page’s visible text (lowercase it, collapse whitespace) and hash it. Pages with the same hash have identical text. This is fast and has no false positives.
Near-duplicates: compare fingerprints
Identical hashes miss pages that differ by a sentence. For those you need a similarity fingerprint. Simhash turns a page’s text into a 64-bit number built from overlapping word sequences. Similar texts produce fingerprints that differ in only a few bits, so you can compare thousands of pages without comparing every word.
Duplicate titles and descriptions
Duplicate <title> tags and meta descriptions are quicker to spot and often point straight at the duplicate templates underneath.
How Crawlens checks
Crawlens fingerprints the visible text of every page (scripts and styles removed) during the crawl, then compares indexable pages with at least 50 words:
| Check | How it’s detected | Severity |
|---|---|---|
| Duplicate content | Identical normalised text | Warning |
| Near-duplicate content | 64-bit simhash over 3-word sequences, at most 3 bits apart | Notice |
| Duplicate page titles | Same <title> on several pages |
Warning |
| Duplicate meta descriptions | Same description on several pages | Warning |
| Thin content | Under 200 words | Notice |
Each flagged URL shows the page it duplicates, so you can review pairs side by side.
One thing to keep in mind when reading results: the comparison uses the whole page’s text, including navigation and footer. On pages with very little unique content, the shared template makes up most of the text, which is exactly why they show up as near-duplicates. The fix is the same either way: more unique content, or fewer pages.
What should you do with each duplicate?
The URL shouldn’t exist: redirect it. HTTP/HTTPS, www and trailing-slash variants, old URLs, staging copies. A 301 to the preferred version consolidates everything.
Both URLs need to work: canonicalise. Sorting and filter parameters, tracking codes, print versions. Keep the URL working for users and point its canonical to the main version. See canonical tag errors for the common mistakes.
The pages should stand alone: rewrite them. Location and service pages, product variants that people search for separately. Give each page content that’s genuinely specific to it: local details, specifications, reviews, FAQs.
The pages add nothing: merge or noindex them. Thin archive pages, near-empty tag pages, variants nobody searches for. Merge them into a stronger page with a redirect, or noindex them.
Checklist
- HTTP,
www, trailing-slash and case variants redirect to one version - Parameter URLs canonicalise to the clean URL
- Exact duplicates redirected or canonicalised
- Near-duplicates reviewed: rewrite, merge or canonicalise
- Unique titles and descriptions on every indexable page
- Thin pages expanded, merged or noindexed
- No staging or demo copies accessible to crawlers
Frequently asked questions
Is duplicate content a Google penalty?
No. Google says some duplicate content on a site is normal and not a violation of its spam policies. The real cost is that duplicates compete with each other and Google chooses which version to show.
What's the difference between duplicate and near-duplicate content?
Duplicate pages have identical text. Near-duplicates are almost the same, for example product variants that differ only by colour or size, or location pages where only the city name changes.
How do you detect near-duplicate pages?
Compare a fingerprint of each page's text instead of the raw text. Simhash builds a 64-bit fingerprint from overlapping word sequences; pages whose fingerprints differ in only a few bits have mostly the same text.
Should I use a canonical or a 301 redirect for duplicates?
Use a 301 when the duplicate URL doesn't need to exist for users. Use a canonical when both URLs need to work, like sorting parameters or tracking codes, but only one should be indexed.
Are duplicate titles and meta descriptions a problem?
They don't stop pages being indexed, but they make different pages look identical in search results and often point to duplicate or thin pages underneath.