Free tool

Can AI crawlers read your site?

Plenty of sites block an AI crawler without knowing it, usually through a CDN firewall rule rather than robots.txt. Test the nine bots that matter in about ten seconds. Free, no signup.

Why access is the first thing to check

Every AI answer starts with retrieval. ChatGPT search reads pages through OAI-SearchBot, Perplexity through PerplexityBot, Claude through ClaudeBot, and Gemini's training respect is governed by the Google-Extended directive. A crawler that cannot fetch your pages cannot cite them, and the engine will quietly build its answer from your competitors' pages instead. Nothing in your analytics tells you this happened.

The failure is usually invisible because it lives in two different places. robots.txt is the visible one, and most teams have checked it once. The silent one is the CDN: bot-management products ship prebuilt rules that block AI crawlers by user agent, and those rules are often enabled by a security team optimizing for scraper defense, with nobody in marketing in the room. This tool tests both layers with real requests, so what you see is what the crawlers get.

Training bots and search bots are different decisions

GPTBot and CCBot feed training corpora: blocking them is a statement about your content's use in future models, and for a publisher selling content that can be rational. OAI-SearchBot, PerplexityBot, ChatGPT-User and Claude-User feed live answers: blocking them removes your brand from the responses buyers are reading this week. Treating those as one decision is the most common mistake we see, because the same robots.txt line often blocks both categories at once.

If you sell something, the practical default is: allow the search and user-triggered bots everywhere, decide on training bots case by case, and re-test after every CDN or security change. The re-test matters more than the first test; access rots when firewall rules update silently.

Frequently asked questions

What does this tool actually test?+

Two things per crawler, both real. First it fetches and parses your robots.txt to see whether each AI bot is allowed or disallowed. Then it requests your homepage carrying each bot's user agent string, which reveals WAF and CDN rules that block bots at the edge even when robots.txt says allow.

A bot shows blocked at the edge. What does that mean?+

Your server or CDN returned an error status to a request identifying itself as that bot. Cloudflare, Akamai and similar services ship AI-bot blocking rules that many sites enable without realizing it. Your robots.txt can say allow while the firewall still turns the crawler away, and the crawler sees the firewall.

Should I allow or block AI crawlers?+

It is a real tradeoff. Blocking training bots like GPTBot protects content from training but removes you from what future models know. Blocking search bots like OAI-SearchBot and PerplexityBot removes you from live AI answers your buyers read today. Most brands that sell something want the search bots allowed.

Is this test 100% accurate?+

The robots.txt reading is exact. The live fetch is a strong signal with one honest limit: our requests come from our servers, and a small number of sites verify crawler IP ranges as well as user agents. A block shown here is real UA-based blocking; a pass here plus IP verification elsewhere is rare but possible.

The bots can read my site. Does that mean AI mentions me?+

No. Access is the entry ticket, and being named in answers is the game. Engines choose which brands to mention based on what they find across the web, which is exactly what Reachroller tracks question by question, with every answer stored so you can audit the score.

Keep reading

Access is step one. Being the answer is the goal.

Reachroller asks the questions your buyers ask and shows you exactly which brands ChatGPT names. Three days free, no card.

Check my brand free