Orphan pages: how to find them and what to do with them

Orphan pages have no internal links pointing to them, so users and crawlers struggle to find them. How to find them with sitemaps, Search Console and logs.

Most pages on a site are reachable from the homepage in a few clicks. Orphan pages aren’t. They exist and may even be indexed, but no other page links to them. Visitors can’t navigate to them, and search engines only find them through a sitemap, an old external link, or memory of a previous crawl.

This guide covers why orphan pages matter, why a normal crawl misses them, and how to find and fix them.

What is an orphan page?

An orphan page is a page with no internal links pointing to it from other pages on the same site. Common causes:

Why do orphan pages matter?

Google’s guide to crawlable links is direct about it: “Google uses links as a signal when determining the relevancy of pages and to find new pages to crawl,” and “every page you care about should have a link from at least one other page on your site.”

An orphan page gets none of that. Listing it in your XML sitemap helps Google discover it, but Google’s sitemap documentation is clear that a sitemap “doesn’t guarantee that all the items in your sitemap will be crawled and indexed.” A sitemap entry also passes no internal link signals.

Orphans also cut the other way: outdated pages nobody maintains can still rank, get traffic, and show visitors old prices or broken features.

Why can’t a crawler find orphan pages?

A crawler discovers pages by following links from the start URL. A page with no internal links is never reached, by definition. To find orphans you need a second list of URLs and a comparison: what’s in that list but missing from the crawl?

Source What it tells you Crawlens check
XML sitemap Pages you declared, but nothing links to Orphan pages (in sitemap, not linked)
Google Search Console Pages Google shows in search, but the crawl didn’t reach URLs with impressions the crawl did not find
Google Analytics 4 Pages people land on (3+ sessions), but the crawl didn’t reach Landing pages with visits the crawl did not find
Server logs Pages Googlebot requests that the crawl didn’t find URLs Googlebot crawls that the crawl did not find

Each source catches different orphans. The sitemap finds pages you still consider live. Search Console and Analytics find pages that matter because they bring traffic. Server logs find what Googlebot spends time on, including old URLs you’d forgotten.

In Crawlens, start a crawl in Sitemap mode with Also follow links turned on, so the crawl has both the sitemap and the internal link graph to compare. Then connect Search Console and GA4 for the traffic checks, and import server logs for the Googlebot check. All four checks need a crawl that followed links. The homepage is never flagged, since it doesn’t need a link to itself.

What should you do with each orphan page?

Go through the list and put each URL in one of four groups.

Add links from two or three relevant, well-linked pages, with anchor text that describes the page. If a whole section went missing from navigation, restore it there.

It’s outdated: redirect it

If a better page now covers the same topic, 301 redirect the orphan there. If there’s no replacement and the page has no value, remove it and return 404 or 410.

It’s a duplicate or parameter URL: consolidate

Orphans that Googlebot finds in your logs are often parameter variants, like sort orders or tracking codes. Point them to the main version with a canonical tag, and stop generating links to them.

It’s unlinked on purpose: decide on indexing

Campaign landing pages and thank-you pages are often meant to be reached only from ads or email. That’s fine. Decide whether they should appear in search at all, and add noindex if not.

Some pages look like orphans to a raw-HTML crawl because their links are added by JavaScript. Crawl with JavaScript rendering on before deciding. If a page is linked only after rendering, see JavaScript SEO: raw vs rendered HTML for why that’s a problem of its own.

Checklist

Frequently asked questions

What is an orphan page in SEO?

A page that exists and can be indexed, but has no internal links pointing to it from other pages on the same site. Visitors can't navigate to it, and crawlers only find it through a sitemap, external links or old data.

Why can't a website crawler find orphan pages?

Crawlers discover pages by following links. A page with no internal links is never reached, so you have to compare the crawl with another list of URLs, such as the XML sitemap, Search Console, Analytics or server logs.

Are orphan pages bad for SEO?

Pages you want to rank should not be orphans. Google uses links to find pages and judge their relevance, and recommends that every page you care about has a link from at least one other page on your site.

Is a page in the XML sitemap still an orphan?

Yes, if nothing on the site links to it. A sitemap helps discovery, but Google says it doesn't guarantee crawling or indexing, and a sitemap entry passes no internal link signals.

What should I do with orphan pages?

Link to the ones that should rank from relevant pages. Redirect retired pages to their closest replacement, or remove them. Campaign landing pages can stay unlinked on purpose, but decide whether they should be indexed.

· Founder, Crawlens

Dien builds Crawlens, a desktop crawler for technical SEO audits, and writes about the checks it runs: crawling, indexing, JavaScript rendering and Search Console data.

Audit your own site with Crawlens

Crawl, run 70 checks, and see them next to Search Console and Core Web Vitals data.

Download free for Windows