Site AI Auditod Internet Solutions

Robots.txt vs Noindex: Which One Should You Use?

29 sierpnia 2026Czas czytania: 8 minPodstawy SEO
Robots.txt vs Noindex: Which One Should You Use?

Short answer: Robots.txt and noindex solve different problems. Robots.txt tells crawlers which URLs they may not fetch, which saves crawl effort but does not reliably keep pages out of search results. Noindex tells search engines not to show a page in results, but it only works if the page can be crawled. Use noindex to keep a page out of search, robots.txt to stop crawlers wasting time on large sets of useless URLs, and never both on the same page.

Two different switches: crawling and indexing

Search engines work in stages. First they crawl: a bot requests a URL and downloads the page. Then they index: they process the page and decide whether to store it so it can appear in results. The two tools act on different stages:

Once you see them as two separate switches, most of the confusion disappears. The key consequence is that a crawler that is not allowed to fetch a page can never read what the page says, including a noindex instruction.

Why the two are so often confused

Part of the confusion is historical. For many years, some search engines unofficially supported a Noindex: line inside robots.txt, and older tutorials still recommend it. Google announced in 2019 that it would stop honouring such rules, and the official standard for robots.txt never included them. Advice written before then can therefore be actively misleading today.

The other part is language. People say “block Google” when they mean “keep this page out of search results”, and the word “block” leads straight to robots.txt. Asking the more precise question, “do I want to stop crawling, or stop showing?”, points to the right tool almost every time.

Side-by-side comparison

robots.txt Disallownoindex
Where it livesOne file at the root of the domainMeta tag in each page, or X-Robots-Tag HTTP header
What it controlsWhether crawlers fetch the URLWhether the page appears in search results
Keeps page out of results?Not reliably; URL can be indexed from linksYes, once the page is recrawled
Saves crawl effort?YesNo, the page must still be crawled
Links on the page followed?No, the page is not readYes, unless nofollow is also set
Works for PDFs and images?Yes, by pathYes, through the X-Robots-Tag header
Applies to many URLs at once?Yes, with path patternsPer page, or through templates and server rules
Protects private content?No, the file is publicNo, the page is still reachable

When to use noindex

Noindex is the right tool when a page should exist for visitors but not appear in search results. Typical cases:

In each case, the page must remain crawlable, so search engines can see the noindex and act on it. On most platforms you add noindex through an SEO plugin setting or a “hide from search engines” option per page or per content type.

When to use robots.txt

Robots.txt is the right tool when the goal is to stop crawlers from spending time on URLs that have no value and would never be indexed anyway, usually because there are very many of them. Typical cases:

For most small websites, crawl effort is not a real constraint, so robots.txt often needs very few rules. It becomes important on large sites and shops, where crawl traps can generate thousands or millions of useless URLs.

Why you should not combine them on the same page

A very common pattern goes like this: a site owner wants to remove a group of pages from Google, adds noindex, and to be extra safe also disallows them in robots.txt. The result is the opposite of what they want. Because the pages are disallowed, Google stops fetching them and never sees the noindex. If other pages or sites link to them, the URLs can remain in the index indefinitely, often shown without a description. Search Console lists them as “Indexed, though blocked by robots.txt”.

The correct order is:

  1. Add noindex and keep the pages crawlable.
  2. Wait until Search Console shows them as excluded by noindex, or request recrawling for the important ones.
  3. Only then, if there are many such URLs and crawl effort matters, add a robots.txt rule.

For a small site with a handful of noindex pages, step 3 is not needed at all.

What neither tool does: privacy

Neither robots.txt nor noindex protects content. Robots.txt is a public file anyone can read, so listing a secret path there actually advertises it. Noindex pages remain fully accessible to anyone with the link. Malicious bots ignore both.

For content that must stay private, such as client documents, internal tools, price lists for partners or staging sites, use authentication: a login, HTTP password protection or IP restrictions. If private content was already indexed, protect it first, then use Search Console’s Removals tool to hide it from results quickly while search engines process the change.

Common scenarios and the right choice

Other robots directives worth knowing

The robots meta tag and X-Robots-Tag header accept more values than noindex. The ones you are most likely to meet are:

Like noindex, all of these only work when the page can be crawled. A robots.txt block silences every one of them.

How to check what is applied to a page

  1. Open /robots.txt and check whether any Disallow rule matches the page’s path.
  2. View the page source and search for noindex.
  3. Check the response headers in the browser’s developer tools for X-Robots-Tag.
  4. Use URL Inspection in Search Console, which reports both whether crawling is allowed and whether indexing is allowed.
  5. Crawl the site to see indexing directives for every page at once.

How Site AI Audit helps

Site AI Audit reads your robots.txt and sitemap and crawls your pages, reporting noindex rules and other things that keep pages out of search results in its SEO section. Each finding names the affected pages and explains the fix in plain words. Its own crawler follows robots.txt, so the check also shows what your rules allow. Run a free check to see how the two switches are set on your site.

Related reading

The bottom line

Robots.txt decides what crawlers fetch; noindex decides what appears in results. To keep a page out of search, use noindex and leave it crawlable. To stop crawlers wasting time on large numbers of worthless URLs, use robots.txt. Do not combine them on the same page, and use real authentication for anything private.

FAQ

Can robots.txt remove a page from Google?

Not reliably. It prevents crawling, but a blocked URL can still be indexed if other pages link to it. Use noindex, a 404 or a redirect to remove pages.

Does noindex save crawl budget?

No. Search engines must crawl a page to see its noindex. Over time they may crawl noindexed pages less often, but robots.txt is the tool for saving crawl effort.

What does “Indexed, though blocked by robots.txt” mean?

Google indexed the URL based on links without being allowed to read the page. If you want it out of search, remove the robots.txt block and add noindex.

Can I put noindex in robots.txt?

No. Google stopped supporting noindex rules in robots.txt in 2019. Use a meta robots tag or an X-Robots-Tag header.

Which is better for duplicate pages?

Neither is ideal. For duplicates, use a canonical tag or a 301 redirect, which consolidate signals on the preferred page instead of simply hiding the copy.

#Crawling#Indexing#robots.txt
Sprawdź swoją stronę — za darmo.Co poprawić na Twojej stronie — i od czego zacząć.
Zacznij za darmo
Internet Solutions

Więcej od naszego zespołu

Stworzone przez Internet Solutions. Wypróbuj nasze pozostałe produkty — każdy oszczędza czas na swój sposób.

internet-solutions.net ↗
Site AI Audit
Przegląd prywatności

Ta strona używa plików cookie, abyśmy mogli zapewnić Ci jak najlepsze wrażenia. Informacje z plików cookie są przechowywane w Twojej przeglądarce i pełnią funkcje takie jak rozpoznawanie Cię po powrocie na stronę oraz pomagają naszemu zespołowi zrozumieć, które sekcje strony są dla Ciebie najciekawsze i najbardziej przydatne.