“Hide this page from Google” sounds like one instruction. Technically it’s two different ones, and mixing them up is one of the most common reasons pages either vanish from search by accident or refuse to leave it.
This guide explains what robots.txt and noindex each do, when to use which, and how to find the mistakes across a site.
What does each one do?
robots.txt Disallow |
noindex |
|
|---|---|---|
| Controls | Crawling: which URLs a crawler may request | Indexing: whether a page may appear in results |
| Where | One file at the site root | Meta tag in the page, or X-Robots-Tag HTTP header |
| Page can still appear in search? | Yes, from links, usually without a description | No, once Google has crawled it |
| Good for | Saving crawl capacity on endless or useless URLs | Keeping specific pages out of search |
Google’s introduction to robots.txt is explicit: robots.txt “is used mainly to avoid overloading your site with requests; it is not a mechanism for keeping a web page out of Google.” A blocked URL “can still appear in search results, but the search result won’t have a description.”
Google’s guide to blocking indexing explains noindex: a <meta name="robots" content="noindex"> tag, or an X-Robots-Tag: noindex response header for non-HTML files like PDFs.
Why doesn’t robots.txt plus noindex work?
It seems belt-and-braces: block the page and noindex it. But Google’s guidance says that for noindex to work, the page “must not be blocked by a robots.txt file”. If the crawler isn’t allowed to fetch the page, it never sees the noindex, and the URL can stay in search from links.
The rule: to keep a page out of search, let it be crawled and use noindex. Use robots.txt for crawl control only.
What are the common mistakes?
1. Blocking pages in robots.txt to hide them
Admin sections, internal search results or thin pages blocked in robots.txt, but still linked from the site. They can appear as bare URLs in results. Fix: allow crawling and noindex them, or stop linking to them.
2. A staging noindex or robots.txt that went live
The most damaging version: Disallow: / or a site-wide noindex copied from staging at launch. Rankings drop within days. See the site migration checklist.
3. Noindex on pages that earn traffic
A template change adds noindex to a section that brings search clicks. Nothing breaks visibly; traffic just fades as Google recrawls.
4. Canonicals pointing at noindexed pages
Page A says “index page B”, page B says “don’t index me”. Both can drop out. See canonical tag errors.
5. Sitemaps listing blocked or noindexed URLs
The sitemap says “this matters”, robots.txt or noindex says the opposite. See the XML sitemap audit.
How do you find these across a site?
Crawl the site with Respect robots.txt on, so blocked URLs are recorded as blocked rather than crawled. Crawlens reads noindex from both the meta robots tag and the X-Robots-Tag header, and runs these checks:
| Check | What it flags | Severity |
|---|---|---|
| Linked pages blocked by robots.txt | Blocked URLs that your own pages link to, with how many pages link to each | Warning |
| Noindex pages | Every noindexed page, to confirm each one is intentional | Notice |
| Canonical points to a noindex page | Conflicting canonical and noindex signals | Critical |
| Non-indexable URLs in sitemaps | Sitemap URLs that are blocked or noindexed (sitemap mode) | Warning |
| Pages with search traffic are not indexable | Pages earning Search Console clicks that are now noindexed, blocked, redirected or canonicalised elsewhere | Warning |
The last check is the one to watch after releases: it catches the noindex that slipped into a template while the pages still have traffic to lose.
Which should you use?
- Keep a page out of search:
noindex, and leave it crawlable. - Keep a file type out of search (PDFs, images):
X-Robots-Tag: noindexheader. - Stop crawlers wasting time on infinite filters, internal search or session URLs: robots.txt
Disallow, and don’t link to them prominently. - Keep something private: password protection. Neither robots.txt nor noindex hides content from people.
Checklist
- Pages you want out of search use noindex and are not blocked in robots.txt
- robots.txt used for crawl control, not for hiding pages
- No internal links to URLs blocked by robots.txt, unless intentional
- Every noindexed page reviewed and intentional
- No canonicals pointing at noindexed pages
- Sitemaps list only crawlable, indexable URLs
- No staging
Disallow: /or site-wide noindex on the live site - Pages with search traffic checked after every release
Frequently asked questions
What's the difference between noindex and robots.txt?
Robots.txt controls crawling, which URLs a crawler may request. Noindex controls indexing, whether a page may appear in search results. Google says robots.txt is not a mechanism for keeping a page out of Google.
Can a page blocked by robots.txt still appear in Google?
Yes. If other pages link to it, Google can index the URL without crawling it. The result usually appears without a description, because Google couldn't read the page.
Can I use noindex and a robots.txt disallow together?
Not usefully. If the page is blocked by robots.txt, Google can't crawl it, so it never sees the noindex. Google says the page must not be blocked by robots.txt for noindex to work.
How do I noindex a PDF or image?
Send an X-Robots-Tag noindex HTTP header with the file. It works for non-HTML resources where you can't add a meta tag.
How long does noindex take to remove a page from Google?
It's applied the next time Google crawls the page. Pages crawled rarely can take a while; requesting a recrawl in Search Console's URL Inspection tool can speed it up.