# ai-bot.txt - welcome list for AI crawlers, search bots and agent frameworks # https://apisphere.us.ci/ai-bot.txt # # APISphere is built for autonomous agents, so every crawler below is # explicitly allowed to crawl, index and quote this site. Nothing here is off # limits, and every block is a robots.txt group, so any robots.txt parser can # read this file. # # Crawl policy (canonical): https://apisphere.us.ci/robots.txt # AI usage preferences: https://apisphere.us.ci/.well-known/ai.txt # (same file at https://apisphere.us.ci/ai.txt) # Sitemap: https://apisphere.us.ci/sitemap.xml # LLM documentation index: https://apisphere.us.ci/llms.txt # API documentation: https://apisphere.us.ci/openapi.json # Global policy: everything public may be crawled, indexed and trained on. User-Agent: * Crawl-delay: 1 Allow: / # OpenAI - model training: Collects content that may be used to train OpenAI models. User-Agent: GPTBot Crawl-delay: 1 Allow: / # OpenAI - search index and citations: Builds the results and citations shown in ChatGPT search. User-Agent: OAI-SearchBot Crawl-delay: 1 Allow: / # OpenAI - user-triggered answers: User-triggered page fetch inside ChatGPT when someone asks. User-Agent: ChatGPT-User Crawl-delay: 1 Allow: / # Anthropic - model training: Collects content that may be used to train Claude models. User-Agent: ClaudeBot Crawl-delay: 1 Allow: / # Anthropic - search index and citations: Builds the search results cited inside Claude. User-Agent: Claude-SearchBot Crawl-delay: 1 Allow: / # Anthropic - user-triggered answers: User-triggered page fetch inside Claude when someone asks. User-Agent: Claude-User Crawl-delay: 1 Allow: / # Google - model training: Controls Gemini training and grounding. User-Agent: Google-Extended Crawl-delay: 1 Allow: / # Google - search index and citations: Google Search index, including AI Overviews and AI Mode. User-Agent: Googlebot Crawl-delay: 1 Allow: / # Microsoft - search index and citations: Bing index, which also grounds Copilot answers. User-Agent: Bingbot Crawl-delay: 1 Allow: / # Apple - search index and citations: Apple search surfaces (Siri, Spotlight). User-Agent: Applebot Crawl-delay: 1 Allow: / # Apple - model training: Controls Apple Intelligence training and grounding. User-Agent: Applebot-Extended Crawl-delay: 1 Allow: / # Perplexity - search index and citations: Builds the index behind Perplexity answers. User-Agent: PerplexityBot Crawl-delay: 1 Allow: / # Perplexity - user-triggered answers: User-triggered fetch from Perplexity. User-Agent: Perplexity-User Crawl-delay: 1 Allow: / # Amazon - search index and citations: Powers Amazon assistant surfaces. User-Agent: Amazonbot Crawl-delay: 1 Allow: / # Meta - model training: Meta AI training and indexing. User-Agent: meta-externalagent Crawl-delay: 1 Allow: / # Meta - search index and citations: Link previews and other Meta surfaces. User-Agent: FacebookBot Crawl-delay: 1 Allow: / # ByteDance - model training: ByteDance / TikTok AI training. User-Agent: Bytespider Crawl-delay: 1 Allow: / # DeepSeek - model training: Collects content that may train DeepSeek models. User-Agent: DeepSeekBot Crawl-delay: 1 Allow: / # Cohere - model training: Training and retrieval for Cohere systems. User-Agent: cohere-ai Crawl-delay: 1 Allow: / # Mistral - user-triggered answers: User-triggered page fetch in Le Chat. User-Agent: MistralAI-User Crawl-delay: 1 Allow: / # DuckDuckGo - user-triggered answers: Fetches pages to answer questions in DuckDuckGo. User-Agent: DuckAssistBot Crawl-delay: 1 Allow: / # You.com - search index and citations: Builds the index behind You.com search. User-Agent: YouBot Crawl-delay: 1 Allow: / # Diffbot - model training: Extracts pages into a knowledge graph. User-Agent: Diffbot Crawl-delay: 1 Allow: / # Timpi - model training: Collects training data for Timpi search. User-Agent: Timpibot Crawl-delay: 1 Allow: / # Imagesift - model training: Collects image data from public pages. User-Agent: ImagesiftBot Crawl-delay: 1 Allow: / # Webz.io - search index and citations: Feeds Webz.io web-data products. User-Agent: Omgilibot Crawl-delay: 1 Allow: / # Common Crawl - open research dataset: Open crawl archive reused as training data by many labs. User-Agent: CCBot Crawl-delay: 1 Allow: / # Product endpoints under /api/v1/* are priced, not closed: a call without # payment returns HTTP 402 with x402 payment requirements, and an agent that # pays in USDC on Base gets the data. See https://apisphere.us.ci/llms.txt