robots.txt

A plain-text file at /robots.txt that tells web crawlers (search engines and AI bots) which pages they may access. Mis-configuring it is the #1 way brands accidentally block themselves from AI engines.

By 2025, the major AI crawlers each have their own user agent: GPTBot (OpenAI), ClaudeBot (Anthropic), PerplexityBot (Perplexity), Google-Extended (Gemini training), Bytespider (ByteDance/Doubao). Many sites block these by default after a 2024 wave of “protect our content” advice — which has the unintended effect of removing the brand from AI answers entirely.

For most B2B brands, you want these bots to crawl you. Add explicit User-agent: GPTBot / Allow: / rules to your robots.txt. VibecodeAEO audits this automatically.

Real-world example

A SaaS company's DevOps team added 'Disallow: /' to all bots after a 2024 scraping scare. Six months later, the brand disappears from Perplexity answers for their core category. The fix: add explicit Allow rules for GPTBot, ClaudeBot, and PerplexityBot. Citation rates recover within 3 weeks.

Frequently asked questions

Which AI crawlers should I allow in robots.txt?+
At minimum, allow: GPTBot (OpenAI/ChatGPT), ClaudeBot (Anthropic/Claude), PerplexityBot (Perplexity), Google-Extended (Gemini), and Bingbot (Bing Copilot). Each needs an explicit 'User-agent: [name] / Allow: /' rule if you have a blanket Disallow or wildcard block in your robots.txt.
Will blocking AI crawlers hurt my SEO?+
Blocking training crawlers (GPTBot, ClaudeBot) does not directly affect Google rankings today, but it removes your content from future AI model training. Blocking retrieval crawlers (PerplexityBot, Bing's crawler) directly kills your visibility in those engines' live answers immediately.

Related terms

Audit your brand against this concept

VibecodeAEO scans your site for all AEO factors weekly and tells you exactly what to fix.