1. Global & General Settings
Universal Crawlers2. AI Crawler Permission Matrix
Toggle permission for each major AI engine individually.
AI Crawlers Reference & Robots.txt Cheat-Sheet
Standardized syntax and recommendations for the major search engines and foundation model crawlers.
Powers real-time search results and source links in ChatGPT Search. Recommended: Always Allow.
User-agent: OAI-SearchBot Allow: /
Indexes content for Perplexity AI answers and footnotes. Recommended: Always Allow.
User-agent: PerplexityBot Allow: /
Collects training data for OpenAI models. Disallowing does not affect ChatGPT Search citations.
User-agent: GPTBot Disallow: /
Anthropic model training crawler. Use Claude-Web to control interactive browsing in Claude chat.
User-agent: ClaudeBot Disallow: /
Controls Gemini and Vertex AI training. Disallowing does not affect Googlebot Search indexation.
User-agent: Google-Extended Disallow: /
Controls content ingestion for Apple Intelligence and Siri generative search features.
User-agent: Applebot-Extended Allow: /
Step-by-Step Platform Setup & Troubleshooting Guides
Is Your robots.txt Working on Your Live Website?
Robots directives can be easily overridden by CDN caching, Cloudflare WAF, or server header misconfigurations.