# ============================ # AI / LLM crawlers # ============================ User-agent: GPTBot Disallow: / User-agent: ChatGPT-User Disallow: / User-agent: CCBot Disallow: / User-agent: anthropic-ai Disallow: / User-agent: ClaudeBot Disallow: / User-agent: Claude-Web Disallow: / User-agent: Google-Extended Disallow: / User-agent: Bytespider Disallow: / User-agent: Amazonbot Disallow: / User-agent: Applebot-Extended Disallow: / User-agent: FacebookBot Disallow: / # ============================ # AI / LLM crawlers - Entraînement (bloqués) # ============================ User-agent: GPTBot Disallow: / User-agent: CCBot Disallow: / User-agent: anthropic-ai Disallow: / User-agent: ClaudeBot Disallow: / User-agent: Google-Extended Disallow: / User-agent: Bytespider Disallow: / User-agent: Amazonbot Disallow: / User-agent: Applebot-Extended Disallow: / User-agent: FacebookBot Disallow: / User-agent: Meta-ExternalAgent Disallow: / User-agent: YouBot Disallow: / User-agent: Diffbot Disallow: / User-agent: Omgilibot Disallow: / User-agent: Omgili Disallow: / User-agent: cohere-ai Disallow: / User-agent: AI2Bot Disallow: / User-agent: Timpibot Disallow: / # ============================ # AI / LLM crawlers - Récupération temps réel (autorisés) # Ces agents ne font pas d'entraînement, ils répondent à une # requête utilisateur explicite (recherche IA, citation) # ============================ User-agent: ChatGPT-User Allow: / User-agent: Claude-Web Allow: / User-agent: PerplexityBot Allow: / # ============================ # SEO / marketing crawlers # ============================ User-agent: Baiduspider Disallow: / User-agent: AhrefsBot Disallow: / User-agent: SemrushBot Disallow: / User-agent: SemrushBot-SA Disallow: / User-agent: MJ12bot Disallow: / User-agent: DotBot Disallow: / User-agent: BLEXBot Disallow: / User-agent: SeekportBot Disallow: / User-agent: Barkrowler Disallow: / User-agent: serpstatbot Disallow: / User-agent: MegaIndex Disallow: / User-agent: ZoominfoBot Disallow: / # ============================ # Generic scraping tools # ============================ User-agent: PetalBot Disallow: / User-agent: SeznamBot Disallow: / User-agent: DataForSeoBot Disallow: / User-agent: HeadlessChrome Disallow: / User-agent: python-requests Disallow: / User-agent: Scrapy Disallow: / User-agent: Wget Disallow: / # ============================ # Règle par défaut pour tous les autres (incl. Googlebot, Bingbot...) # ============================ User-agent: * Allow: / Disallow: /fileadmin/_recycler_/ Disallow: /fileadmin/_temp_/ Disallow: /fileadmin/user_upload/_temp_/ Disallow: /typo3/ Disallow: /*?id=* Disallow: /*&id=* Disallow: /*?L=0* Disallow: /*&L=0* Disallow: /typo3temp/* Allow: /typo3temp/*.css Allow: /typo3temp/*.css.*.gzip Allow: /typo3temp/*.js Allow: /typo3temp/*.js.*.gzip Allow: /typo3temp/*.jpg Allow: /typo3temp/*.jpeg Allow: /typo3temp/*.gif Allow: /typo3temp/*.webp Allow: /typo3temp/*.png Disallow: *.sql Disallow: *.sql.gz Sitemap: https://fim-moto.com/fr/?type=1533906435 Sitemap: https://www.fim-rmm.com/sitemap.xml Crawl-delay: 5