Noindex vs robots.txt: which to use, and the mistakes that hide pages

Robots.txt controls crawling; noindex controls indexing. How each works, why combining them backfires, and how to find pages hidden or exposed by mistake.

“Hide this page from Google” sounds like one instruction. Technically it’s two different ones, and mixing them up is one of the most common reasons pages either vanish from search by accident or refuse to leave it.

This guide explains what robots.txt and noindex each do, when to use which, and how to find the mistakes across a site.

What does each one do?

robots.txt Disallow noindex
Controls Crawling: which URLs a crawler may request Indexing: whether a page may appear in results
Where One file at the site root Meta tag in the page, or X-Robots-Tag HTTP header
Page can still appear in search? Yes, from links, usually without a description No, once Google has crawled it
Good for Saving crawl capacity on endless or useless URLs Keeping specific pages out of search

Google’s introduction to robots.txt is explicit: robots.txt “is used mainly to avoid overloading your site with requests; it is not a mechanism for keeping a web page out of Google.” A blocked URL “can still appear in search results, but the search result won’t have a description.”

Google’s guide to blocking indexing explains noindex: a <meta name="robots" content="noindex"> tag, or an X-Robots-Tag: noindex response header for non-HTML files like PDFs.

Why doesn’t robots.txt plus noindex work?

It seems belt-and-braces: block the page and noindex it. But Google’s guidance says that for noindex to work, the page “must not be blocked by a robots.txt file”. If the crawler isn’t allowed to fetch the page, it never sees the noindex, and the URL can stay in search from links.

The rule: to keep a page out of search, let it be crawled and use noindex. Use robots.txt for crawl control only.

What are the common mistakes?

1. Blocking pages in robots.txt to hide them

Admin sections, internal search results or thin pages blocked in robots.txt, but still linked from the site. They can appear as bare URLs in results. Fix: allow crawling and noindex them, or stop linking to them.

2. A staging noindex or robots.txt that went live

The most damaging version: Disallow: / or a site-wide noindex copied from staging at launch. Rankings drop within days. See the site migration checklist.

3. Noindex on pages that earn traffic

A template change adds noindex to a section that brings search clicks. Nothing breaks visibly; traffic just fades as Google recrawls.

4. Canonicals pointing at noindexed pages

Page A says “index page B”, page B says “don’t index me”. Both can drop out. See canonical tag errors.

5. Sitemaps listing blocked or noindexed URLs

The sitemap says “this matters”, robots.txt or noindex says the opposite. See the XML sitemap audit.

How do you find these across a site?

Crawl the site with Respect robots.txt on, so blocked URLs are recorded as blocked rather than crawled. Crawlens reads noindex from both the meta robots tag and the X-Robots-Tag header, and runs these checks:

Check What it flags Severity
Linked pages blocked by robots.txt Blocked URLs that your own pages link to, with how many pages link to each Warning
Noindex pages Every noindexed page, to confirm each one is intentional Notice
Canonical points to a noindex page Conflicting canonical and noindex signals Critical
Non-indexable URLs in sitemaps Sitemap URLs that are blocked or noindexed (sitemap mode) Warning
Pages with search traffic are not indexable Pages earning Search Console clicks that are now noindexed, blocked, redirected or canonicalised elsewhere Warning

The last check is the one to watch after releases: it catches the noindex that slipped into a template while the pages still have traffic to lose.

Which should you use?

Checklist

Frequently asked questions

What's the difference between noindex and robots.txt?

Robots.txt controls crawling, which URLs a crawler may request. Noindex controls indexing, whether a page may appear in search results. Google says robots.txt is not a mechanism for keeping a page out of Google.

Can a page blocked by robots.txt still appear in Google?

Yes. If other pages link to it, Google can index the URL without crawling it. The result usually appears without a description, because Google couldn't read the page.

Can I use noindex and a robots.txt disallow together?

Not usefully. If the page is blocked by robots.txt, Google can't crawl it, so it never sees the noindex. Google says the page must not be blocked by robots.txt for noindex to work.

How do I noindex a PDF or image?

Send an X-Robots-Tag noindex HTTP header with the file. It works for non-HTML resources where you can't add a meta tag.

How long does noindex take to remove a page from Google?

It's applied the next time Google crawls the page. Pages crawled rarely can take a while; requesting a recrawl in Search Console's URL Inspection tool can speed it up.

· Founder, Crawlens

Dien builds Crawlens, a desktop crawler for technical SEO audits, and writes about the checks it runs: crawling, indexing, JavaScript rendering and Search Console data.

Audit your own site with Crawlens

Crawl, run 70 checks, and see them next to Search Console and Core Web Vitals data.

Download free for Windows