AI Crawler Robots.txt Test Suite: 24 RFC 9309 Policy Cases

An AI crawler robots.txt test suite is a small, repeatable safety net: this package parses a populated policy, evaluates exactly 24 synthetic requests, and compares generated CSV output byte-for-byte with a committed expectation. It does not observe a production crawler, verify a bot's identity, or prove that a provider will obey the policy.
The downloadable package uses only Python 3's standard library. It is intentionally a documentation-based engineering asset, not a live-provider benchmark. The ExampleAnswerBot name is synthetic so readers do not mistake a fixture for a current allowlist.
Table of contents
- What does policy testing prove?
- How does RFC 9309 precedence work?
- What are the 24 cases?
- How do I run the package?
- Why is robots.txt not access control?
- How should teams adapt this to real agents?
- What are the limitations?
- FAQ
What does policy testing prove?
The suite proves a narrower and useful claim: given this file, parser, and request fixture, the evaluator produces the expected decision deterministically. That catches regressions when someone edits crawler groups, adds a broad wildcard, or changes an exception path.
It does not prove retrieval, indexing, recommendation, or visibility in ChatGPT, Google, Perplexity, or another answer engine. Crawler accessibility is a prerequisite signal for content retrieval, not evidence that a system selected, understood, cited, or recommended a page. Connect these checks to an AI search audit, rather than treating a green policy test as a visibility score.
The package is also not a crawler simulator. It does not send requests, inspect response headers, follow redirects, or identify a client from IP ranges. Those boundaries matter because a syntactically permissive policy can coexist with authentication, outages, rate limits, or an uncooperative client.
How does RFC 9309 precedence work?
The evaluator follows a practical RFC 9309 subset. Field names and product-token matching are case-insensitive. A crawler uses the most specific matching user-agent group; if no specific token matches, the * group is the fallback. Rules from repeated groups for the same token are merged.
For a matching path, the longest matching rule wins. An Allow wins an equal-length tie with Disallow. * matches a sequence of characters, $ anchors a rule at the end of the path, and percent-encoded octets are compared after decoding for this suite's path cases. An empty Disallow contributes no restriction.
The implementation intentionally keeps the rule boundary visible. It extracts the path from an absolute URL, preserves the raw rule for the report, and records the winning rule rather than only returning a boolean. That makes a failure reviewable: an analyst can see whether /private/public/ won because its exception was longer than /private/.
What are the 24 cases?
The populated cases.csv contains exactly 24 synthetic rows. They cover ordinary allow and disallow outcomes, merged groups, product-token matching with a version suffix, case-insensitive tokens and fields, wildcard fallback, and a specific group taking precedence over *.
The edge cases are the reason to keep fixtures around. /assets/*.json$ distinguishes an exact suffix from a longer filename; /encoded/%E2%9C%93 checks percent encoding; /tie/* and Allow: /tie/abc$ demonstrate the longest-match and terminal-anchor interaction. /private/ versus /private/public/ tests a more specific exception, while /merged/ proves that a second group for the same synthetic token is not silently discarded.
The expected decisions are committed in expected.csv. Every row also records the winning rule, which is useful when a policy review asks “why was this path blocked?” The fixtures are examples, not observations: none of the 24 rows describes a real provider's current fetch behavior.
Editorial image: Pexels photo 574071; it is not test evidence.
How do I run the package?
Download the README.md, robots.txt, cases.csv, expected.csv, and evaluate_policy.py as one directory. Python 3.9 or newer is sufficient; there are no package installs and no API keys.
From that directory, run:
python3 evaluate_policy.py --robots robots.txt --cases cases.csv --output /tmp/robots-expected.csv
cmp expected.csv /tmp/robots-expected.csv
The runner rejects a fixture that is not exactly 24 rows, rejects duplicate case IDs, and exits nonzero if a generated decision differs from its declared expected decision. cmp must produce no output and exit 0. This byte comparison protects against a test that passes because it merely prints a hand-written expected answer.
For a broader measurement workflow, pair this deterministic artifact with an AI visibility metrics dictionary and record policy version, collection date, requested path, user-agent string, and response observations separately. The policy file alone is never a substitute for an evidence ledger.
Why is robots.txt not access control?
robots.txt is a cooperative convention. A compliant crawler can use it to decide which paths to request, but an adversary can ignore it. Do not place secrets in a supposedly disallowed URL, and do not use a passing test as evidence that private content is protected.
Use authentication, authorization, network controls, and server-side data minimization for confidential material. Then use robots.txt to communicate crawl preferences for content that may safely be requested. This distinction is especially important for AI search: “crawlable” does not mean “eligible for citation,” and “blocked” does not prove a provider never saw a URL through another channel.
How should teams adapt this to real agents?
Start by copying the suite and replacing the synthetic policy with a version-controlled snapshot of your own public robots.txt. Add cases for paths that matter to your content system: documentation, pricing, product pages, feeds, staging prefixes, and parameterized URLs. Keep expected output under review so a broad change cannot silently remove an exception.
Use official specifications as the source of parser behavior. RFC 9309 defines the Robots Exclusion Protocol; RFC 3986 supplies URI syntax context. Google's robots.txt documentation explains its documented implementation, while OpenAI's crawler documentation is a dated discovery reference—not proof that this synthetic test suite validates OpenAI agents.
For production follow-through, compare these policy assertions with the AI crawler log analyzer, which frames observed requests separately from declared policy. If your team needs to draft a policy before testing it, the robots.txt generator is a useful starting point; always run the resulting file through review and this regression suite.
Keep provider names out of the fixture unless you are explicitly testing a documented token and recording the date. A current user-agent string can change, and a token match cannot establish that the connection really came from that product.
What are the limitations?
This is a focused regression evaluator, not a complete production robots implementation. It omits HTTP fetch status and caching, redirect chains, the RFC's 500 KiB parse ceiling, non-UTF-8 recovery, authentication, IP verification, and any question of whether a bot complies. Its percent-encoding behavior is limited to the documented path comparison cases; it is not a general URL canonicalizer or public-suffix parser.
The 24 requests are synthetic and deliberately small. They cannot estimate crawler traffic, search demand, indexing probability, citation rate, or brand recommendation. A green result means only that this parser and fixture agree. Record parser version and policy hash when using it in an audit, and independently test the site's HTTP delivery and access controls.
FAQ
What does this robots.txt test suite prove?
It proves that the included evaluator produces the declared result for 24 synthetic policy cases under its documented RFC 9309 subset. It does not prove that a real crawler identifies itself correctly, fetches the file, or complies.
Does it test OpenAI, Google, or another live crawler?
No. It uses the synthetic ExampleAnswerBot token. Official documentation is linked as context, but the fixtures make no claim about current provider behavior.
Why can Allow win over Disallow?
The longest matching rule wins. When matching Allow and Disallow rules have equal length, Allow wins in this RFC 9309 subset.
Can I use robots.txt as access control?
No. Use server-side authentication and authorization for private data. robots.txt communicates crawl policy; it does not enforce secrecy.
FAQ
What does this robots.txt test suite prove?+
It proves that the included evaluator produces the declared result for 24 synthetic policy cases under its documented RFC 9309 subset. It does not prove that a real crawler will identify itself correctly, fetch the file, or comply.
Does the suite test OpenAI, Google, or another live crawler?+
No. It uses the synthetic ExampleAnswerBot token deliberately. Official crawler documentation is linked as context, but the fixtures make no claim about current provider behavior.
Why can Allow win over Disallow?+
When both rules match, the suite selects the longest matching rule. If matching Allow and Disallow rules have equal length, Allow wins, following the RFC 9309 precedence rule.
Can I use this as access control?+
No. robots.txt is a cooperative crawl policy, not authentication or authorization. Protect private data with server-side access controls and treat crawler compliance as a separate operational question.
Sources
Is AI recommending you?
Run a free AI visibility audit and find out in under a minute.