Sitemap: https://igrownews.com/sitemap.xml Sitemap: https://igrownews.com/news-sitemap.xml User-agent: * Disallow: /wp-admin/ Disallow: /month/ Allow: /wp-admin/admin-ajax.php # --------------------------------------------------------------- # AI training crawlers — blocked. # The archive is licensed through mcp.igrownews.com, not scraped. # See https://mcp.igrownews.com/terms # --------------------------------------------------------------- User-agent: GPTBot Disallow: / User-agent: ClaudeBot Disallow: / User-agent: anthropic-ai Disallow: / User-agent: Claude-Web Disallow: / User-agent: CCBot Disallow: / User-agent: Google-Extended Disallow: / User-agent: Applebot-Extended Disallow: / User-agent: Meta-ExternalAgent Disallow: / User-agent: Meta-ExternalFetcher Disallow: / User-agent: FacebookBot Disallow: / User-agent: Amazonbot Disallow: / User-agent: Bytespider Disallow: / User-agent: Diffbot Disallow: / User-agent: omgili Disallow: / User-agent: omgilibot Disallow: / User-agent: Webzio-Extended Disallow: / User-agent: ImagesiftBot Disallow: / User-agent: AI2Bot Disallow: / User-agent: Timpibot Disallow: / User-agent: PanguBot Disallow: / User-agent: Kangaroo Bot Disallow: / User-agent: cohere-ai Disallow: / User-agent: cohere-training-data-crawler Disallow: / # Added 2026-08-15 (first pass, from server-log evidence). Both are # training-style crawlers doing systematic archive reads; neither is a # search engine. User-agent: IbouBot Disallow: / User-agent: Intelligensia Disallow: / # Added 2026-08-15 (second pass, competitive-scraping review). # ShapBot reads each page exactly once -- 153 requests, 153 distinct URLs -- # with no contact URL, and it has NEVER fetched robots.txt. This rule records # intent but will NOT stop it; blocking ShapBot needs a WAF/CDN rule. User-agent: ShapBot Disallow: / # Declares a fake contact URL (example.com is an IANA-reserved placeholder) # and spent 32 of its 37 requests enumerating sitemaps -- pre-harvest # reconnaissance, not research. It does read robots.txt, so this one lands. User-agent: AcademicResearchBot Disallow: / # --------------------------------------------------------------- # AI search & answer engines — ALLOWED, deliberately. # These cite and link, so they are referral traffic, not extraction. # Free discovery, paid depth. # --------------------------------------------------------------- User-agent: OAI-SearchBot Allow: / User-agent: ChatGPT-User Allow: / User-agent: PerplexityBot Allow: / User-agent: Perplexity-User Allow: / User-agent: Claude-User Allow: / User-agent: Claude-SearchBot Allow: / # Added 2026-08-15 (second pass). Already permitted by User-agent: * -- # named explicitly so the permission reads as deliberate and nobody later # "tidies up" by blocking them. YouBot fetched robots.txt 47 times in the # sample, the most compliant behaviour of any crawler observed. User-agent: YouBot Allow: / User-agent: AzureAI-SearchBot Allow: / User-agent: LinkupBot Allow: / User-agent: trendictionbot Allow: /