Skip to content

What is robots.txt, and can it keep a page out of Google?

What robots.txt controls, and why it won't hide a page from Google

robots.txt is a file that tells search engine crawlers which URLs on your site they may access, and it can't reliably keep a page out of Google. Google says the file is mainly for avoiding too many requests to your site, and that a blocked URL can still be indexed if other pages link to it. To keep a page out of Google Search, use noindex or password protection instead.

How an engagement works.

Sources read October 11, 2026 · by Heddle

What the file does

robots.txt tells search engine crawlers which URLs on a site they may visit. Google's documentation says its main job is managing crawler traffic so a site doesn't receive more requests than it can handle. For ordinary pages, including HTML and PDF, Google suggests using it when crawler requests are overloading your server, or when you don't want Google crawling unimportant or near duplicate pages.

If your site runs on a CMS such as Wix or Blogger, you may not be able to edit the file at all. Google says these platforms often offer a search settings page or a similar way to tell search engines whether to crawl a page.

Keeping a page out of Google Search

It won't hide a page from Google Search. Google is direct about this: if other pages link to a blocked URL with descriptive text, Google can still index that URL without visiting it. The result can show up in search with no description, because Google never read the page.

To keep a page out of results, Google recommends a noindex meta tag or response header, password protection on the server, or removing the page entirely. Google also warns that mixing crawling rules with indexing rules can cause some rules to conflict with others.

Pages, media and resource files

The file behaves differently by file type. For image, video and audio files, robots.txt can keep them out of Google search results, though other pages and users can still link to them.

For resource files such as unimportant images, scripts or style files, you can block them if the page loads fine without them. Google advises against blocking resources that it needs to understand a page, because it can't properly analyze pages that depend on them. When a page itself is blocked, files embedded in it aren't crawled either, unless another crawlable page references them.

Who obeys it

The rules are requests, not enforcement. Google says Googlebot and other reputable crawlers follow robots.txt, but not every crawler does, and not every search engine supports every rule. Crawlers can also read the same syntax in different ways. For information that must stay private, Google recommends password protecting the files on your server rather than relying on robots.txt.

Sources

Each was read on October 11, 2026.

Questions

Does blocking a page in robots.txt remove it from Google?

No. Google can still index a blocked URL if other pages link to it, and it may show the URL in results without a description. Use noindex or password protection to keep it out.

Do all crawlers follow robots.txt?

No. Google says Googlebot and other reputable crawlers follow it, but obeying the rules is up to each crawler, and some may ignore them or read the syntax differently.

Should I block scripts and style files in robots.txt?

Only if the page works without them. Google advises against blocking resources it needs to understand a page, because it can't properly analyze pages that depend on them.

Work with us

Let the agents bring you the leads.

Tell us about your business and who buys from you. On a short call we'll show you the searches your buyers make where you don't show up yet, and what the agents would do about them in their first month.

On the call, for your site

  1. The searches your buyers makeMeasured, with how many people make each one
  2. Where you show up, and where you don'tOn Google and in ChatGPT's answers
  3. The agents' first monthThe pages, fixes and links they would start with