Free tool
robots.txt checker
robots.txt is your published crawl policy, and one wrong line can remove you from AI answers without an error appearing anywhere. Enter a domain and this checker parses the whole file: every user-agent group, what each major AI and search crawler is allowed or forbidden to fetch, and the syntax problems that make the file mean something different from what its author intended.
One file, and it decides who gets to read you
The crawlers behind AI answers each read robots.txt before fetching: GPTBot for OpenAI training, OAI-SearchBot for ChatGPT search, ChatGPT-User for user-triggered browsing, ClaudeBot and Claude-User for Anthropic, PerplexityBot for Perplexity, and the Google-Extended token governing Gemini's use of your content. A blanket Disallow written years ago to slow scrapers now silently covers all of them, which means the file quietly decides whether your brand can appear in the answers your buyers read.
The most misunderstood rule of the format is group selection. A crawler picks the single most specific user-agent group that matches it and obeys only that group; it never merges rules across groups. A file that carefully allows GPTBot in one group while a later wildcard group disallows everything does what the GPTBot group says for GPTBot alone, and many hand-edited files get this backwards in both directions.
How to read your results
The report shows a verdict per crawler: which group it matched and what it may fetch, including the paths that matter, since a bot allowed on the homepage but disallowed from /blog is invisible where your citable content lives. Below the verdicts sit the syntax findings: rules placed before any user-agent line, which are dead text; misspelled bot names that match nothing; wildcard and dollar-sign patterns that cover more or less than intended; and crawl-delay lines, which most modern crawlers ignore.
One boundary to keep in mind: this tool audits the policy as written. Plenty of sites publish a permissive robots.txt while their CDN blocks AI crawlers by user agent at the edge, and the file never shows that. Our bot access checker covers that second layer by sending live requests as each bot, and the two results together give you the full picture.
What to fix first
Start with any finding where a search-answering bot, OAI-SearchBot, PerplexityBot or ChatGPT-User, is blocked from content you want cited, because that is present-tense invisibility. Then decide the training bots, GPTBot and CCBot, deliberately rather than by inherited wildcard. Then clean the dead syntax so the next editor inherits a file that means what it says. Once access is settled, the question becomes whether the engines that now can read you actually name you, and Reachroller measures that question by question, from $29 per month with a free 3-day trial.
Frequently asked questions
How is this different from the bot access checker?+
This tool parses your robots.txt and reports the policy: which crawlers the file allows where, plus syntax problems. The bot access checker performs live fetches carrying each AI bot's user agent, which exposes CDN and firewall blocks the file never mentions. Policy and enforcement disagree on a surprising number of sites, so the tools answer different questions and are worth running together.
Does blocking GPTBot remove my brand from ChatGPT answers?+
GPTBot governs training data, so blocking it keeps your content out of future model training while live ChatGPT search answers depend on OAI-SearchBot and ChatGPT-User. Those are separate user agents with separate rules, and blocking the search pair is what removes your pages from the answers buyers read today. Decide the two categories separately; one Disallow line often catches both by accident.
What is the most common robots.txt mistake you see?+
The stale wildcard: a User-agent: * group with broad Disallow rules written long ago for scrapers or a staging launch, now silently governing every AI crawler that has appeared since. Close behind are rules that assume groups merge, when each crawler obeys only its single best-matching group, and typos in bot names that leave the intended rule matching nothing at all.
My robots.txt allows everyone, so why am I absent from AI answers?+
Permission is the floor. Check enforcement next, since firewalls and CDN bot rules block AI crawlers on sites whose robots.txt says allow. Then look at the content layer: engines cite pages that answer specific buyer questions with evidence, and an accessible site with generic pages still loses citations to a competitor whose pages answer the exact question asked.
More free tools and reading
Tools find the gaps. Tracking closes them.
Reachroller asks the questions your buyers ask and shows which brands ChatGPT names. Three days free, no card.
Check my brand free