Robots.txt Parser

Paste a robots.txt and read it the way a crawler does, per RFC 9309 (the Robots Exclusion Protocol, standardized 2022): user-agent groups, Allow/Disallow rules, sitemaps and unknown directives, all laid out — plus an audit for the classic mistakes (empty disallow lines, unknown directives, oversized files). The matcher then answers the real question: paste a user agent and a path, and get the allow/deny verdict with the exact rule that decided it — longest match, wildcard *, end anchor $ and group merging, exactly as the standard specifies. Everything runs in your browser.
Nothing parsed yet. Paste a robots.txt above and press Parse file, or load a sample.
The matcher works on the text currently in the box above (no need to press Parse first). The user agent is matched the way RFC 9309 defines: the User-agent value in the file must be a case-insensitive substring of what you type here — so typing ExampleBot/1.0 matches a group declared as ExampleBot. A full URL is accepted as the path.
User agent (a real UA string works) Path, or full URL
No verdict yet. Press Test permission.
How to fetch your file: open https://your-domain.example/robots.txt in a browser, select-all, copy, and paste here — or curl -s https://your-domain.example/robots.txt in a terminal. This page never fetches anything by itself: it only reads what you paste. To inspect the response headers that carried the file, run them through the HTTP header parser.
Self-test
Re-runs the parser and matcher against the RFC 9309 reference vectors: longest-match precedence (including the allow-beats-disallow tie and the deep-path case), group merging, the * group fallback, case sensitivity of paths and user agents, wildcards, the $ anchor, the implicit allow of /robots.txt, percent-encoding normalization, sitemap extraction, and the negative paths. Nothing is sent anywhere.
What is robots.txt?
In one sentence: robots.txt is a plain-text file at the root of a site that tells well-behaved crawlers which paths they may fetch — a voluntary convention, standardized as RFC 9309 in 2022, that every major crawler honors.
The shape of a file

A file is a sequence of records. A group starts with one or more User-agent lines, followed by Allow and Disallow rules that apply to those crawlers. A new User-agent line after rules begins a new group; several User-agent lines in a row share one group (a crawler with two names needs this). Standalone records like Sitemap can appear anywhere. # starts a comment anywhere in a line. Field names are case-insensitive; paths are not.

LineMeans
User-agent: foobotThe rules below apply to crawlers whose name contains foobot (case-insensitive)
User-agent: *The rules below apply to every crawler without a more specific group
Disallow: /private/Do not fetch paths starting with /private/
Allow: /private/public/Exception that beats the disallow above for these paths (longest match)
Disallow: (empty)Nothing is disallowed — the classic "allow all" idiom
Sitemap: https://…/sitemap.xmlWhere your sitemap lives; crawlers may collect several
Crawl-delay: 30Not in RFC 9309 — a legacy extension; major crawlers ignore or interpret it themselves
How a verdict is reached (RFC 9309 §2.2.2)

1. Find the group whose User-agent value is a substring of the crawler's name; merge all matching groups. If none, use the * group; if there is none of those either, everything is allowed. 2. Test the path against every rule: matching is prefix-based, case-sensitive, and */$ work as wildcards. 3. The longest matching rule wins; at equal length, Allow beats Disallow. No matching rule means allowed. And whatever the rules say, /robots.txt itself is implicitly allowed — a crawler could not read the rules otherwise.

Common mistakes
  • Believing robots.txt blocks access. It is a request to polite crawlers, not a wall: rogue scrapers ignore it completely, and it has no effect on what a browser can open. For real protection use authentication; for steering search engines only, robots.txt is the right tool.
  • Expecting Disallow: / to remove a page from search results. It stops crawling, not indexing: if other pages link to yours, a search engine may still list the URL (without a description). The noindex meta tag or response header is the removal tool.
  • Relying on order instead of length. The first matching rule does not win — the most specific (longest) one does. Reordering lines to "fix" behavior is a no-op; make the intended rule longer instead.
  • Blocking CSS and JS. Disallowing /*.css$ and /*.js$ used to be common "optimization"; renderers now need those files to see the page as visitors do, and a half-rendered page can rank worse.
  • Treating one big group as documentation. Comment lines inside a group are fine, but a paragraph of prose before the first User-agent line without # is a parse error that can confuse lenient parsers.
FAQ

Do all crawlers support Allow, * and $? The majors do, and RFC 9309 standardizes all three. Some smaller or older crawlers understand only Disallow and read * literally — one reason to keep rules simple when you can.

What does an empty Disallow: mean? "Disallow nothing" — the group explicitly allows everything. It usually appears when someone edits a template and deletes the path but not the directive; the audit calls it out.

How big may the file be? RFC 9309 requires parsers to handle at least 500 KiB; beyond that a crawler may stop reading and treat the rest as disallowed. Keep it well under, and if it grows huge, question whether crawling control belongs in page headers instead.

Why does the matcher say /robots.txt is allowed even under Disallow: /? Because the standard says so, in so many words: the robots file itself is implicitly allowed. Python's urllib.robotparser, among others, does not implement this special case — one of the small divergences between real-world parsers that this page documents in its self-test.

Is a Disallow path case-sensitive? Yes. /Private/ and /private/ are different paths. The field name disallow may be any case, but the value must match the URL exactly.