robots.txt Generator

Generate a robots.txt file with Allow, Disallow, Crawl-delay and Sitemap directives for any user agent. Runs entirely in your browser, no signup.

Loading tool…

Worked examples

  • Block admin areas, allow one public asset, with a sitemap

    The common real-world shape: block a couple of sections wholesale, carve out one specific exception with Allow, and point crawlers at the sitemap.

  • Allow everything, for all crawlers

    An empty Disallow value is the standard way to explicitly tell every crawler nothing on the site is off-limits, rather than omitting the file entirely.

  • A specific bot with multiple blocked paths and a crawl delay

    Targeting a named user agent instead of * is how you apply different rules to one specific crawler, such as asking Bing's crawler to wait 10 seconds between requests with Crawl-delay without affecting others (Google ignores Crawl-delay, so it would be pointless on a Googlebot group).

What this tool does

This tool builds a robots.txt file from the rules you specify: which user agent it applies to, which paths are disallowed, which paths are explicitly allowed as exceptions, an optional crawl delay, and an optional sitemap URL. The result is a complete, correctly formatted file ready to drop at the root of your site.

When you need it

  • Setting up crawl rules for a new site, where getting the syntax exactly right — the directive names and the leading slashes on paths — matters more than it might seem, since a malformed line can be ignored or misread by some crawlers. The format is standardised in RFC 9309 (the Robots Exclusion Protocol), and the file must live at the root of the host, such as https://example.com/robots.txt.
  • Stopping well-behaved crawlers from fetching a staging environment, an admin area, or search/filter URLs that create near-duplicate content, without needing to hand-write the rules each time. Note that blocking crawling is not the same as blocking indexing: a disallowed URL that other sites link to can still appear in search results without a description. To keep a page out of the index, leave it crawlable and serve a noindex robots meta tag or X-Robots-Tag header, or put it behind authentication.
  • Carving out a specific exception to a broader block — allowing one public asset inside an otherwise disallowed directory, for instance — which requires getting the Allow/Disallow precedence right.
  • Auditing an existing robots.txt by rebuilding your intended rules here and comparing the two, to catch a typo'd path or a rule that's blocking more than intended.

How Allow and Disallow interact

Rules are evaluated per user agent, and when both an Allow and a Disallow rule could match the same URL, the more specific rule — the one with the longest matching path — wins, regardless of which directive it uses or which order it appears in. This is why the earlier example of blocking /private entirely but allowing /private/public-asset.png works: the more specific Allow rule takes precedence for that one path, while everything else under /private stays blocked by the broader Disallow rule.

What an empty Disallow line means

A Disallow: line with nothing after the colon is the standard, explicit way to tell a crawler that nothing on the site is off-limits — functionally equivalent to having no restrictions at all, but stated directly rather than by omission. Under RFC 9309 a missing robots.txt (a 404 response) also means the whole site may be crawled, so the line does not change the outcome; it makes your intent explicit and keeps the file from being empty, which is why this tool always emits it when you haven't specified any paths to block. Be careful with the opposite case: a robots.txt that returns a server error (5xx) is treated by crawlers as a full disallow until it recovers. Also be careful with Disallow: /, which blocks the entire site.

Crawl-delay isn't universally respected

Bing and several other crawlers honor the Crawl-delay directive as a request to wait a given number of seconds between requests. Google does not — it ignores Crawl-delay entirely, and crawl rate for Google is instead managed through Search Console's own settings. The sitemap line must be a full URL beginning with http:// or https://; the tool rejects anything else. Include Crawl-delay if you're specifically trying to throttle crawlers that respect it, but don't rely on it as a universal rate-limiting mechanism.

Limits

This tool generates syntactically correct robots.txt content; it doesn't verify the paths you enter actually exist on your site, and it doesn't substitute for testing the live file with each search engine's own robots.txt testing tool once it's deployed, since crawler-specific quirks in interpretation do exist.

Frequently asked questions

Is my site information uploaded anywhere?
No. The file is assembled entirely in your browser from the paths and options you enter; nothing about your site is sent anywhere.
What does an empty Disallow: line mean?
It explicitly means nothing is disallowed — the crawler may access the entire site. Under RFC 9309, a missing robots.txt (a 404) also means the whole site may be crawled, so the line is there to state your intent explicitly rather than to change the outcome.
How many paths can I list?
One path per line in the Disallow and Allow boxes — each is normalized to start with a leading slash automatically if you leave it off.
Do all search engines respect Crawl-delay?
No — Google ignores the Crawl-delay directive entirely (use Search Console's crawl rate settings instead), while Bing and several other crawlers do honor it. Include it for the crawlers that support it.