Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more

To create a robots.txt file, first check whether your site platform already manages one, then write only the crawler rules needed for a clear purpose, publish the UTF-8 text file at the root of the exact site origin, and test the live result. Use it to guide compliant crawlers away from selected URL paths—not to secure private content or reliably remove pages from Google Search.

What robots.txt does—and what it cannot do

Google describes robots.txt as a file that tells search engine crawlers which URLs they can access. It is a public set of crawler instructions, not an access-control system. The Internet Engineering Task Force’s RFC 9309 makes the distinction explicit: “These rules are not a form of access authorization.” Compliant crawlers are asked to follow the rules; a bot that ignores them can still request the paths.

A Disallow rule also does not guarantee that a URL will disappear from search. A search engine may know a URL from links or other references even if it cannot crawl the page to read its content. Choose the mechanism based on the outcome you need:

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Method What happens to crawler access Effect on search visibility Use it for
robots.txt Disallow Asks compliant crawlers not to fetch matching paths. Does not guarantee exclusion; a discovered URL can remain visible. Managing crawler access to URL patterns or crawl traffic.
noindex The crawler must be able to fetch the page to see the directive. Requests that the page be excluded from search results. Keeping a page accessible to crawlers while requesting that it not appear in search.
Authentication or password protection Prevents unauthorized retrieval. Keeps protected content unavailable to public search crawlers. Private or restricted material.

For a page you want deindexed, leave it crawlable and use an appropriate noindex directive. For content that must remain private, require authentication. Blocking a private URL in robots.txt can both expose its path in a public file and prevent a crawler from seeing a page-level noindex directive.

Where to put robots.txt

Save the file as robots.txt in the top-level path of the relevant service, so it loads at that origin’s /robots.txt. For example, the rules at https://www.example.com/robots.txt do not automatically apply to https://example.com/robots.txt, an HTTP version of the site, or a different subdomain. Protocol, host, and port determine scope.

The file should be plain UTF-8 text served as text/plain, consistent with RFC 9309. Putting it in a subdirectory does not make it the site-wide robots.txt. If a CMS or hosted platform provides a search-visibility or crawler setting, check its official documentation before editing files directly; the platform may generate the file or manage related settings for you.

How to write robots.txt rules

Each group starts with a User-agent line naming the crawler it targets, followed by directives such as Allow and Disallow. Paths are relative to the origin root, and URLs that are not disallowed are allowed by default. Google supports the * and $ wildcards in path values; do not assume every crawler interprets every directive the same way.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Start with a small file rather than copying another site’s rules. This illustrative example blocks paths beginning with /private-preview/ for crawlers that follow the group and advertises a sitemap location; those paths are examples, not universal SEO recommendations:

User-agent: *
Disallow: /private-preview/

Sitemap: https://www.example.com/sitemap.xml

A sitemap record must use a fully qualified URL. It tells crawlers where a sitemap is; it does not allow or block access to listed paths, nor does it guarantee indexing. Keep sitemap discovery separate from crawl restrictions.

For Google, the supported records include user-agent, allow, disallow, and sitemap. Google does not support crawl-delay, so adding it will not control Googlebot. Google also notes that an asterisk user-agent group does not cover AdsBot crawlers; name an AdsBot explicitly if you need rules for one. Check the current documentation for each crawler before using nonstandard records.

How overlapping rules are resolved

When Allow and Disallow rules both match, RFC 9309 says the most specific matching rule should govern. If equivalent Allow and Disallow rules tie, the standard says to resolve the tie in favor of Allow. Because crawler implementations and supported syntax can differ, test overlapping rules with the crawler-specific tool or documentation relevant to your site.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

A safe workflow for creating and optimizing the file

  1. Define the goal. Decide whether you need to manage crawler requests to a URL area, rather than hide a page from search or protect private content. Use noindex for search exclusion requests and access controls for privacy.
  2. Inspect the live origin and platform. Open the exact site origin followed by /robots.txt, and check whether your CMS or host already creates or controls the file. Account for each relevant protocol, hostname, and port.
  3. Inventory paths before blocking. Identify URL patterns that crawlers do not need to request. Check that the rules will not block useful pages or resources needed to render and understand them. Google recommends not blocking CSS or JavaScript when their absence would impair its understanding of a page.
  4. Draft targeted groups. Put each directive on its own line beneath the intended User-agent. Use root-relative paths and review wildcard patterns against actual URLs. Avoid broad rules copied from another site and check the effect of overlapping Allow and Disallow rules.
  5. Add a sitemap record if useful. Include the complete sitemap URL, including protocol and host. Confirm that the sitemap itself is valid; its presence does not change which URLs a crawler may fetch.
  6. Publish and verify. Save the file as UTF-8 plain text and publish it at the root of the exact origin. Confirm that the public /robots.txt loads as text, then test representative URLs that should be allowed and blocked. Google documents testing through Search Console; a compatible local parser can also help check the resulting rules.
  7. Monitor after changes. Review crawl and indexing reports for unexpected effects. A crawler may use a cached copy, so a changed file need not affect requests immediately; recheck the live file and allow for crawler-specific refresh behavior.

Common robots.txt mistakes and how to avoid them

  • Expecting Disallow to remove a result: it controls crawler requests, not search visibility. Leave the page accessible for a crawler to read a noindex directive, or protect it with authentication if it is private.
  • Using it as a secret list: the file is publicly retrievable, and its paths can disclose URL patterns. Do not place sensitive information in it or treat it as a security barrier.
  • Publishing it under the wrong URL: a file in a subdirectory, or on another host or protocol, does not govern the origin you meant to control. Verify the exact root URL.
  • Blocking rendering resources: a rule that prevents Google from fetching important CSS or JavaScript may impair its understanding of a page. Preserve access to resources needed for rendering.
  • Relying on unsupported or crawler-specific syntax: Google does not support crawl-delay, and directives beyond common records may vary by bot. Check the target crawler’s documentation rather than assuming one file works identically everywhere.
  • Using a partial sitemap address: write a fully qualified sitemap URL with protocol and host.
  • Assuming changes apply instantly: crawlers can cache robots.txt. Verify the published version and investigate again after the crawler has had a chance to refresh.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

What the standard says about retrieval and file size

RFC 9309, an IETF Standards Track document published in September 2022, says crawlers should use a parsing limit of at least 500 KiB. It recommends following at least five consecutive redirects when retrieving robots.txt. It also distinguishes a file that is unavailable from one that cannot be reached because of network or server errors: the standard’s handling differs, and crawler implementations document their own details. When diagnosing a Googlebot issue, consult Google’s current guidance rather than assuming every crawler handles retrieval failures the same way.

The standard says crawlers should not use a cached robots.txt for more than 24 hours unless the file is unreachable. That is a standards recommendation, not a promise that every crawler will fetch a change on a fixed schedule. Bing documentation also notes caching, but refresh timing should be treated as crawler-specific rather than a universal guarantee.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.