AEO
Should you block GPTBot and PerplexityBot?
This decision gets made badly because the crawlers get lumped together, and they are not the same thing.
Two different jobs
Some crawlers gather data used to train models. Blocking those is a preference about how your content is used, and it costs you very little visibility, because they are not what fetches a page to answer somebody's question today.
Other crawlers fetch pages live to support answers being generated right now. Blocking those removes you from the answers themselves.
The names are similar enough that people block both while intending only the first. That is the single most common error here.
What we would generally recommend
For a local business trying to be found, allow the search crawlers without hesitation. Being present in AI answers is the entire objective, and blocking the crawler that produces them is self-defeating.
Training crawlers are a judgement call with no visibility consequence either way. Businesses with original written work sometimes block them on principle. Businesses that want reach allow them. Both are defensible, and neither changes whether you appear in ChatGPT Search or Google's overviews.
Check what your host is doing
This is the part worth acting on today. Bot protection at the CDN or hosting layer frequently blocks AI crawlers by default, independently of anything in your robots.txt.
The result is a site whose owner believes it is open, which is in fact invisible to every assistant. If you have ever wondered why you never appear in AI answers despite doing everything right, check this before rewriting anything.
Write it down
Whatever you decide, put a comment in the robots.txt explaining the reasoning and the date.
This decision gets revisited, usually by somebody who was not there the first time, and a file full of unexplained rules is how businesses end up accidentally blocking themselves for a year.
Common questions
Should I block AI crawlers in robots.txt?
It depends entirely on which crawler. Training crawlers gather data to improve models, and blocking them is a data-use preference with little visibility cost. Search crawlers fetch pages to answer live questions, and blocking those removes you from the answers themselves.
What is the difference between GPTBot and OAI-SearchBot?
They are separate crawlers with separate purposes. One gathers training data, the other fetches pages to support live search answers. Blocking both is a common mistake made by owners who only intended to opt out of training.
Does blocking AI crawlers protect my content?
Partially, and only from the crawlers that respect the instruction. Robots.txt is a convention rather than an enforcement mechanism, so it stops well-behaved crawlers and nothing else.
Could my host be blocking AI crawlers without my knowledge?
Yes, and this is common. Bot protection at the CDN or hosting level can refuse AI crawlers regardless of what robots.txt says, so a site can be invisible to assistants while its owner believes it is open.