Robots.txt Tester
Parse, validate, and test your robots.txt rules against any URL and user-agent — and see which rule decided the answer.
Test a URL
A full URL or a path — either works, and the result updates as you type. The query string is kept, because rules like Disallow: /*?sort= exist to match it.
Enter a path to see whether Googlebot may fetch it.
Bulk URL test
Test a whole list at once — one URL or path per line.
Shared with the single-URL test above.
Parsed rule structure
Nothing to parse yet. Paste a robots.txt, fetch one from a domain, or load the example — the structure below updates as you type. A site with no robots.txt at all is treated as allowing everything.
How to use it
- Paste a robots.txt into the editor, or switch to Fetch from domain to pull the live file from any site. The rule tree, the validation list and the test result all update as you type.
- Pick a user-agent and enter a path to test. The verdict names the exact rule that decided it, along with every other rule that matched but lost — click any line number to jump to it in the editor.
- Use the bulk test to check a whole list of URLs against the same rules in one pass, then export the results as CSV.
Getting robots.txt right
- Google ignores Crawl-delay entirely. Bing and Yandex honour it; for Google, set the crawl rate in Search Console instead — or fix the server response times that made you want to slow it down.
- The * wildcard matches any run of characters, and a trailing $ anchors the pattern to the end of the URL. Without the $, Disallow: /*.pdf also blocks /report.pdf?download=1 — with it, only the bare .pdf URL is blocked.
- Robots.txt stops crawling, not indexing. A blocked URL that other sites link to can still appear in results as a bare blue link with no snippet, because Google knows the URL exists but was never allowed to read the page.
- Never pair a robots.txt Disallow with a meta noindex on the same URL. Google has to crawl the page to see the noindex, and the Disallow stops it — the page stays indexed indefinitely.
- Google reads the first 500KB of a robots.txt and ignores everything after it. If your file is generated, check the size before you trust the rules at the bottom of it.
- GPTBot, ChatGPT-User, ClaudeBot, Google-Extended, CCBot and PerplexityBot can each be blocked to keep your content out of AI training and retrieval. Blocking Google-Extended does not affect your Search rankings — it only opts you out of Gemini and AI Overviews grounding.
- Once a crawler finds a group naming it, it ignores the User-agent: * group completely. A Googlebot block with two rules in it does not inherit the twenty rules you wrote under *.
Frequently asked questions
How does robots.txt work?
It is a plain text file at the root of a host — https://example.com/robots.txt — that a crawler fetches before it requests anything else on that host. Inside, rules are gathered into groups: one or more User-agent lines naming crawlers, followed by the Allow and Disallow paths that apply to them. A crawler reads the file, finds the single group that names it, and obeys that group's rules for every URL it considers fetching on that host. The file is advisory, not enforced — well-behaved crawlers follow it, and anything scraping your site in bad faith will not.
Does robots.txt block pages from appearing in Google?
No, and this is the single most expensive misunderstanding in technical SEO. Robots.txt controls crawling — whether Google may fetch the URL — not indexing. If other pages link to a URL you have disallowed, Google can still index it from those links alone and show it in results as a title-only entry with no description, because it was never permitted to read the page. To keep a page out of the index, leave it crawlable and add a meta robots noindex tag or an X-Robots-Tag header; Google has to be able to fetch the page to see that directive.
What is the difference between Allow and Disallow?
Disallow names a path prefix a crawler must not fetch; Allow names one it may, and exists to carve exceptions out of a broader Disallow. Both are matched against the path from the root of the site, and both are case-sensitive. When more than one rule in the applicable group matches a URL, Google does not use the first or the last — it uses the longest matching pattern, so Allow: /admin/public/ beats Disallow: /admin/ for a URL under /admin/public/. Where two matching patterns are exactly the same length, Allow wins the tie.
How do wildcards work in robots.txt?
Two characters are special. An asterisk matches any sequence of characters, including none, so Disallow: /*?sort= blocks any URL with a sort parameter anywhere in the query string. A dollar sign at the very end of a pattern anchors the match to the end of the URL, so Disallow: /*.pdf$ blocks /guide.pdf but not /guide.pdf?v=2. Everywhere else a pattern is a plain prefix match, and there is no escape syntax — a path containing a literal asterisk cannot be expressed. Google, Bing and Yandex all support both characters; smaller crawlers may not.
Can I block AI crawlers like GPTBot and ClaudeBot?
Yes, the same way you block any other crawler: give each one its own group with Disallow: /. The main product tokens are GPTBot and ChatGPT-User for OpenAI, ClaudeBot for Anthropic, Google-Extended for Google's AI training and grounding, CCBot for Common Crawl, PerplexityBot for Perplexity, and Bytespider for ByteDance. Google-Extended is worth understanding separately: it is not a search crawler, so blocking it removes your content from Gemini training and AI Overviews grounding without affecting how Googlebot crawls or ranks you. All of this depends on the crawler choosing to obey the file, which the major ones do and unnamed scrapers do not.