# AI Girlfriend App Reviews — public crawl policy, 2026-10-07 # Public HTML, Markdown, JSON and editorial feeds are available to crawlers. # Search indexing, AI discovery, retrieval and training crawls are allowed. # Crawling preferences do not grant ownership of third-party material. # User-triggered fetchers and product-control tokens may follow different policies. # Canonical content and source dates are listed in llms.txt and data/site.json. User-agent: * Allow: / # Potential foundation-model training crawl # Independent from search permission; operator publishes IP ranges. Allow indicates crawl preference, not an indexing guarantee. User-agent: GPTBot Allow: / # ChatGPT search discovery # Search choice is independent of GPTBot. Allow and reachable origin help eligibility; results are not guaranteed. User-agent: OAI-SearchBot Allow: / # User-requested retrieval # Not automatic crawling or search inclusion control; robots.txt may not apply to these user actions. User-agent: ChatGPT-User Allow: / # Potential model-training collection # Anthropic says its bots honor robots.txt; separate from user retrieval and search. User-agent: ClaudeBot Allow: / # Claude search relevance and indexing # Distinct search permission; disabling can reduce search visibility. No ranking promise from Allow. User-agent: Claude-SearchBot Allow: / # User-directed retrieval # Anthropic states that its user-directed retrieval crawler honors robots.txt. User-agent: Claude-User Allow: / # Google Search crawling # Allow permits crawling but does not ensure indexing, ranking or rich results. User-agent: Googlebot Allow: / # Gemini training and specified grounding controls # A robots.txt product token, not a separate HTTP crawler. Does not affect Google Search inclusion or ranking. User-agent: Google-Extended Allow: / # Bing search crawling # Official user-agent includes lowercase bingbot; token matching is case-insensitive. Honors REP; Allow is not an indexing promise. User-agent: bingbot Allow: / # Perplexity search surfacing # Operator says not foundation-model training; requires reachable content and compatible WAF policy, with no appearance guarantee. User-agent: PerplexityBot Allow: / # User-requested retrieval # Not automatic training crawl; generally ignores robots.txt for user-requested fetches. User-agent: Perplexity-User Allow: / # Apple search discovery and content collection # Used by Spotlight, Siri and Safari. Collection may feed generative features; Extended controls training use separately. User-agent: Applebot Allow: / # Apple foundation-model training use control # Does not crawl pages; controls use of Applebot-collected data. Does not affect Search ranking. User-agent: Applebot-Extended Allow: / # DuckDuckGo AI-assisted answer retrieval # Not model training. Its permission does not affect ordinary organic search rankings; changes may take 72 hours. User-agent: DuckAssistBot Allow: / # Amazon product improvement and potential AI training # Automated crawls honor REP. Some robots metadata is honored; crawl-delay is unsupported. User-agent: Amazonbot Allow: / # Amazon search experiences such as Alexa # Not generative-model training. Separate user-agent setting; may follow other search-bot rules when unnamed. User-agent: Amzn-SearchBot Allow: / # User-requested live Amazon retrieval # Not model-training crawl. User-initiated requests may not follow all robots.txt directives. User-agent: Amzn-User Allow: / # Common Crawl open web archive # Public crawl repository can support downstream research and training. Operator warns that user-agent strings can be spoofed. User-agent: CCBot Allow: / # Shared-link previews in Meta products # May bypass robots.txt for security or integrity checks; preview retrieval is distinct from model-training collection. User-agent: FacebookExternalHit Allow: / # Meta AI search citations and linking # Documented search indexing crawler; Allow records crawl preference but does not guarantee inclusion or citations. User-agent: Meta-WebIndexer Allow: / # Ads and business-product improvement # Distinct product-improvement crawler; its purpose differs from AI search and user-requested retrieval. User-agent: Meta-ExternalAds Allow: / # AI-model training and direct indexing for product improvement # Separate from user-requested fetching; an Allow group records public-content crawl permission. User-agent: Meta-ExternalAgent Allow: / # User-requested links and agentic product features # Meta says it may bypass robots.txt because the fetch was requested by a user. User-agent: Meta-ExternalFetcher Allow: / Sitemap: https://ai-girlfriend-app-2026.com/sitemap.xml