Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

robots.txt tells compliant web crawlers which parts of a site they should avoid; it does not secure those pages or reliably remove them from Google Search. Use it to guide crawler requests, use a crawlable noindex directive to keep an accessible page out of Google, and use authentication or another access control to protect private content.

What robots.txt does

robots.txt is a plain-text crawler instruction file, normally published at /robots.txt at the root of a site authority. Its rules request that automated clients avoid or access specified URL paths. The Internet Engineering Task Force’s RFC 9309, Robots Exclusion Protocol, published in September 2022, states: “These rules are not a form of access authorization.”

# Preview Product Price
1 Advanced Robots.txt Generator Manual Advanced Robots.txt Generator Manual $32.46

The distinction matters: crawler guidance is not a technical barrier. A crawler may choose not to honor a rule, and a person can still visit a URL unless the server separately restricts access. RFC 9309 also warns that paths named in the file are publicly exposed, so listing a sensitive-looking path can draw attention to it rather than protect it. See the RFC 9309 standard and Google’s overview of robots.txt.

How rules select crawlers and paths

A robots.txt file is UTF-8 text. Each group begins with a User-agent line identifying a crawler, followed by path rules such as Disallow and Allow. A wildcard group applies when no more specific group is selected.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
User-agent: ExampleBot
Disallow: /private-looking/

User-agent: *
Allow: /
  • User-agent: ExampleBot identifies the crawler group this rule applies to.
  • Disallow: /private-looking/ asks that crawler not to fetch paths beginning with that path. The name is only a label; it does not make the content private.
  • User-agent: * is the general group for crawlers without a matching specific group. Allow: / permits paths under the site root for that group.

Under RFC 9309, a crawler applies the most specific matching path rule; if an allow and disallow rule are equivalent, the standard says the allow rule should win. Individual crawlers may differ in their implementation details. Google documents its own rule handling in its robots.txt specification and guidance.

Rules are scoped to the protocol, host, and port where the file is hosted, as Google explains in its robots.txt introduction. A file on https://example.com does not automatically govern a sibling subdomain, the HTTP version of the site, or a different port. Publish rules for each applicable site authority.

What robots.txt can and cannot stop

It can guide compliant crawlers away from paths

A disallow rule can reduce requests to selected paths by crawlers that choose to honor it. This is useful for managing crawler traffic, but it is not a way to block every bot, human visitor, or unauthorized user. RFC 9309 defines the protocol as crawler instructions, not access control; Google likewise notes that crawler behavior is up to each crawler.

It cannot reliably keep a URL out of Google Search

Disallow prevents Google from fetching a blocked path while the rule applies; it does not guarantee that Google will omit the URL from search results. Google may discover the URL through links and show it without having crawled the page content. In that situation, the URL can appear even though Google cannot read the page to build a content-based snippet.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Do not put noindex in robots.txt: Google does not support it there. For an accessible page that should not appear in Google Search, leave it crawlable and place a supported noindex directive on the page or in its HTTP response. Google must be able to fetch the page to see the directive. Its instructions are in Preventing search indexing with noindex.

It cannot protect private content

If a page or file must be inaccessible to people who are not authorized, enforce that at the server with authentication or another valid access-control mechanism. RFC 9309 recommends an application-layer security measure such as HTTP authentication. Do not place secrets at a publicly reachable URL and rely on a robots.txt rule to conceal them.

Choose the mechanism for the outcome you need

Goal Mechanism Key limitation
Reduce requests to selected paths from compliant crawlers robots.txt rules Crawlers can ignore the instructions; the rules do not restrict human or unauthorized access. RFC 9309
Keep an accessible page out of Google Search A crawlable noindex meta tag or X-Robots-Tag HTTP response header Google must fetch the page to read the directive; a robots.txt disallow can hide it. Google’s noindex guidance
Make private content inaccessible to unauthorized visitors Server-side authentication or another valid access-control measure A robots.txt rule is not access control and should not be used to conceal secrets. RFC 9309
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

What happens if robots.txt cannot be fetched?

The protocol’s general behavior and Google’s implementation should not be treated as identical. RFC 9309 distinguishes an unavailable file from an unreachable one: a 4xx unavailable response may allow a crawler to access resources, while server or network errors that make the file unreachable require complete disallow under the standard. It also recommends that crawlers generally not use a cached copy for more than 24 hours unless the file is unreachable.

Google documents its own handling: most 4xx responses other than 429 are treated as if there are no crawl restrictions. With 5xx errors, Google initially stops crawling and retries; if it cannot fetch a fresh file, it may use a cached version for a period. Google generally caches robots.txt for up to 24 hours, but may retain a cached file longer when it cannot refresh it. Those are Google-specific details, not universal crawler behavior. See Google’s robots.txt handling guidance.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Technical limits in the standard

RFC 9309 requires parsers to support a robots.txt file size limit of at least 500 KiB. That is a minimum parser capability, not a suggested target file size. The standard’s general recommendation that cached rules not be used for more than 24 hours has the unreachable-file exception described above. Google separately documents a 500 KiB limit and its own cache behavior; these implementation details apply to Google, not automatically to every crawler.

Quick Recap

SaleBestseller No. 1

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.