Most pages on a site are reachable from the homepage in a few clicks. Orphan pages aren’t. They exist and may even be indexed, but no other page links to them. Visitors can’t navigate to them, and search engines only find them through a sitemap, an old external link, or memory of a previous crawl.
This guide covers why orphan pages matter, why a normal crawl misses them, and how to find and fix them.
What is an orphan page?
An orphan page is a page with no internal links pointing to it from other pages on the same site. Common causes:
- a page removed from navigation during a redesign, but never deleted
- blog posts that dropped off paginated archives and were never linked again
- campaign landing pages built for ads or email
- product pages that left every category when they went out of stock
- old URLs from a previous version of the site that still return 200
Why do orphan pages matter?
Google’s guide to crawlable links is direct about it: “Google uses links as a signal when determining the relevancy of pages and to find new pages to crawl,” and “every page you care about should have a link from at least one other page on your site.”
An orphan page gets none of that. Listing it in your XML sitemap helps Google discover it, but Google’s sitemap documentation is clear that a sitemap “doesn’t guarantee that all the items in your sitemap will be crawled and indexed.” A sitemap entry also passes no internal link signals.
Orphans also cut the other way: outdated pages nobody maintains can still rank, get traffic, and show visitors old prices or broken features.
Why can’t a crawler find orphan pages?
A crawler discovers pages by following links from the start URL. A page with no internal links is never reached, by definition. To find orphans you need a second list of URLs and a comparison: what’s in that list but missing from the crawl?
| Source | What it tells you | Crawlens check |
|---|---|---|
| XML sitemap | Pages you declared, but nothing links to | Orphan pages (in sitemap, not linked) |
| Google Search Console | Pages Google shows in search, but the crawl didn’t reach | URLs with impressions the crawl did not find |
| Google Analytics 4 | Pages people land on (3+ sessions), but the crawl didn’t reach | Landing pages with visits the crawl did not find |
| Server logs | Pages Googlebot requests that the crawl didn’t find | URLs Googlebot crawls that the crawl did not find |
Each source catches different orphans. The sitemap finds pages you still consider live. Search Console and Analytics find pages that matter because they bring traffic. Server logs find what Googlebot spends time on, including old URLs you’d forgotten.
In Crawlens, start a crawl in Sitemap mode with Also follow links turned on, so the crawl has both the sitemap and the internal link graph to compare. Then connect Search Console and GA4 for the traffic checks, and import server logs for the Googlebot check. All four checks need a crawl that followed links. The homepage is never flagged, since it doesn’t need a link to itself.
What should you do with each orphan page?
Go through the list and put each URL in one of four groups.
It should rank: link to it
Add links from two or three relevant, well-linked pages, with anchor text that describes the page. If a whole section went missing from navigation, restore it there.
It’s outdated: redirect it
If a better page now covers the same topic, 301 redirect the orphan there. If there’s no replacement and the page has no value, remove it and return 404 or 410.
It’s a duplicate or parameter URL: consolidate
Orphans that Googlebot finds in your logs are often parameter variants, like sort orders or tracking codes. Point them to the main version with a canonical tag, and stop generating links to them.
It’s unlinked on purpose: decide on indexing
Campaign landing pages and thank-you pages are often meant to be reached only from ads or email. That’s fine. Decide whether they should appear in search at all, and add noindex if not.
Watch for links that only exist in JavaScript
Some pages look like orphans to a raw-HTML crawl because their links are added by JavaScript. Crawl with JavaScript rendering on before deciding. If a page is linked only after rendering, see JavaScript SEO: raw vs rendered HTML for why that’s a problem of its own.
Checklist
- Crawl with the XML sitemap and compare it with internal links
- Compare the crawl with Search Console and Analytics URLs that get traffic
- Compare the crawl with Googlebot requests in your server logs
- Re-check suspected orphans with JavaScript rendering on
- Link to orphans that should rank from relevant, well-linked pages
- Redirect or remove outdated orphans; canonicalise parameter variants
- Keep the sitemap to indexable, linked pages
Frequently asked questions
What is an orphan page in SEO?
A page that exists and can be indexed, but has no internal links pointing to it from other pages on the same site. Visitors can't navigate to it, and crawlers only find it through a sitemap, external links or old data.
Why can't a website crawler find orphan pages?
Crawlers discover pages by following links. A page with no internal links is never reached, so you have to compare the crawl with another list of URLs, such as the XML sitemap, Search Console, Analytics or server logs.
Are orphan pages bad for SEO?
Pages you want to rank should not be orphans. Google uses links to find pages and judge their relevance, and recommends that every page you care about has a link from at least one other page on your site.
Is a page in the XML sitemap still an orphan?
Yes, if nothing on the site links to it. A sitemap helps discovery, but Google says it doesn't guarantee crawling or indexing, and a sitemap entry passes no internal link signals.
What should I do with orphan pages?
Link to the ones that should rank from relevant pages. Redirect retired pages to their closest replacement, or remove them. Campaign landing pages can stay unlinked on purpose, but decide whether they should be indexed.