llms.txt, robots.txt and AI Crawlers: How to Stop Blocking the Engines
If your robots.txt blocks OAI-SearchBot or PerplexityBot, ChatGPT search and Perplexity can't cite you. The fix takes ten minutes.
Most businesses that block AI crawlers in their robots.txt are doing it by accident, and the fix takes under ten minutes.
What I want to cover here is the more urgent issue upstream of llms.txt. Before your content can be cited in AI answers, the engines have to be able to reach it, and that holds for generative engine optimisation (GEO) and answer engine optimisation (AEO) alike. For a significant proportion of the clients who come to Lore, that first step has never been taken.
The technical arm that nobody touches
There are three arms to effective visibility work: content, technical (crawlability, indexing, speed, internal linking, metadata), and backlinks plus brand mentions. Content gets most of the attention, because it's visible and it's something most people feel confident doing. The other two arms sit quietly and get left alone.
Technical work is unglamorous. It's invisible when it's working correctly, and most business owners have no reason to open their robots.txt file unless something has clearly broken. The result is that the technical arm gets neglected, and the content work built on top of it underperforms, sometimes dramatically.
I see this every time a new client arrives on a GEO plan. The first thing I check is their robots.txt. A significant number have AI crawlers blocked by default. Usually nobody chose that: their CMS or hosting platform shipped with the blocks in place and nobody reviewed them. It's a cheap fix with a severe cost if it goes untouched. The major AI crawlers say they honour robots.txt. Block them there, and all the content in the world does nothing for your share of voice in AI answers.
What your robots.txt is doing right now
Open a browser and go to yourdomain.com/robots.txt. This plain-text file is the first document every well-behaved crawler reads before it indexes a single page. If it says `Disallow: /` for a particular user-agent, that crawler stops at the door.
Look for OAI-SearchBot, Claude-SearchBot and PerplexityBot, and for a `User-agent: *` group. If any of them sits under a `Disallow` directive covering your whole site, those engines can't read you. Your content isn't being read. It can't be cited, regardless of how strong it is or how well-structured your pages are.
The time between correcting your robots.txt and seeing the change reflected in ChatGPT search is approximately 24 hours for OAI-SearchBot, per OpenAI's own documentation. That's how quickly you can begin to recover ground.
Three OpenAI user agents and what each one does
OpenAI documents several user agents, and three of them matter here. The difference between them is where most of the confusion lies.
GPTBot crawls content that may be used in training future models. If you disallow GPTBot, your content won't be incorporated into OpenAI's training data. This is a legitimate choice for some publishers, particularly those with concerns about content rights. Blocking GPTBot alone, however, doesn't remove you from ChatGPT search results.
OAI-SearchBot is the crawler that powers ChatGPT search. Disallow it, and per OpenAI's own documentation, your site "will not be shown in ChatGPT search answers, though can still appear as navigational links." This is the directive that directly affects your GEO and AEO visibility in ChatGPT, and the one most commonly blocked inadvertently.
ChatGPT-User is used when a ChatGPT user's question leads ChatGPT to visit a page. It's not an automatic indexing crawler and isn't used to determine inclusion in search answers. Your standard robots.txt directives may not apply to it in the usual way. It behaves more like a browser visiting your page than a crawler building an index.
The critical point: the GPTBot and OAI-SearchBot settings are independent. You can allow OAI-SearchBot while disallowing GPTBot, which means you can appear in ChatGPT search answers without contributing to model training. A blanket AI block switches both off.
Beyond OpenAI: ClaudeBot, PerplexityBot and the rest
The same logic applies to the other major AI engines, each with their own crawler.
Anthropic splits the same jobs. ClaudeBot collects training data; Claude-SearchBot indexes pages for Claude's search answers, so that's the one to allow if you want Claude to cite you. PerplexityBot is Perplexity's indexing agent, and it needs access to your pages before Perplexity can cite them.
Google works differently. AI Overviews are part of Google Search and use the regular Googlebot crawl, so a page that's indexed in Google Search and eligible for a snippet can appear. Google's separate Google-Extended token only controls whether your content is used to train and ground Gemini. Blocking it doesn't remove you from AI Overviews or affect your rankings.
Each of these is configurable independently. A properly structured robots.txt allows the crawlers that serve citation-based answers while giving you full control over which training data pipelines your content feeds into. The change itself takes minutes.
Where llms.txt fits in
llms.txt is still only a proposed standard. Know that before you prioritise it.
The idea works like a sitemap aimed specifically at LLMs: a structured markdown file at your domain root that points to your most useful, citation-ready pages in clean format. Rather than leaving an AI crawler to infer what matters from your full site, llms.txt makes the hierarchy explicit: here is our product page, here is our pricing FAQ, here is the evergreen guide on this topic.
It's worth building. But it's not the urgent priority if your AI crawlers are still blocked. llms.txt isn't yet widely adopted, and the major AI platforms don't currently require it to index your content. robots.txt is the file the crawlers actually read today. llms.txt is a forward-looking complement, worth adding once the foundation is in order.
Ahrefs covers both files in detail in their AEO course: how they interact, what each one changes, and how to set them up correctly.
<YouTube id="RyJYGpVyl0o" />
One action, two channels
Allowing AI crawlers serves both GEO and traditional search. This is worth stating clearly, because the most common mistake I see is treating the two as competing priorities, so resources go to one channel while the other sits unaddressed. Treated as a continuum, the same robots.txt update serves both sides.
Allowing OAI-SearchBot and PerplexityBot doesn't change how Googlebot behaves. You aren't redirecting crawl budget or touching anything that affects your standard search rankings. You're extending crawlability to the engines driving a growing share of referral traffic, and that traffic converts. AI referrals convert at three to six times the rate of traditional organic.
Understanding how the engines choose what to cite is useful context here, because crawlability is a prerequisite for everything else. Authority signals, content structure and citation density only count once the technical layer is correct.
Lore's case studies show what the full programme produces. Tides Mental Health grew organic traffic 4.7 times in five months, with AI platforms recommending them for anxiety treatment searches. Ambiance Creations went from 939 to over 31,000 clicks in twelve months. BizScout built 450-plus AI mentions, with ChatGPT citing them as the go-to source for business buyers.
Start with the five-minute check
Go to yourdomain.com/robots.txt. Look for OAI-SearchBot, Claude-SearchBot and PerplexityBot. If any appear under `Disallow: /`, update the file to allow them or remove the directive, then give it 24 hours to propagate. GPTBot and ClaudeBot only feed training, so allow or block them as you prefer.
From there, the next layers are llms.txt, structured metadata, internal linking, and making sure your strongest pages are genuinely citation-ready. Technical work is rarely the exciting part of a visibility strategy. It is, however, what makes everything else possible.
The free 46-point AI visibility checklist walks through every layer, from crawler access to brand signals, so you can see where your site stands.
Frequently asked questions
What is llms.txt, and do I need it?
llms.txt is a proposed standard: a markdown file at your domain root that points AI systems to your most useful pages. It's worth adding, but the major AI platforms don't require it to read your site. Fix robots.txt first.
Does blocking GPTBot stop ChatGPT recommending my brand?
Not for ChatGPT search. GPTBot collects training data. OAI-SearchBot is the crawler behind ChatGPT search answers, and OpenAI lets you allow one while blocking the other.
Which AI crawlers should I allow?
For citations, allow OAI-SearchBot for ChatGPT search, Claude-SearchBot for Claude and PerplexityBot for Perplexity. Google's AI Overviews use regular Googlebot. Training crawlers such as GPTBot and ClaudeBot are your call.