# === Default policy === User-agent: * Disallow: /pages/ # === Major search engines === User-agent: Googlebot User-agent: Bingbot User-agent: Applebot User-agent: DuckDuckBot User-agent: Slurp User-agent: Yandex User-agent: Baiduspider Disallow: /pages/ # === AI answer / citation crawlers === User-agent: OAI-SearchBot User-agent: ChatGPT-User User-agent: PerplexityBot User-agent: Perplexity-User User-agent: Claude-SearchBotP User-agent: Claude-User Disallow: /pages/ # === AI training / model-use crawlers === User-agent: GPTBot User-agent: ClaudeBot User-agent: Google-Extended User-agent: Bytespider Disallow: /pages/ # === SEO / analytics / research crawlers === User-agent: AhrefsBot User-agent: SemrushBot User-agent: DotBot User-agent: MJ12bot User-agent: CCBot User-agent: archive.org_bot User-agent: ia_archiver Disallow: /pages/ # === Optional: block a few legacy high-noise downloaders/scrapers === # Note: malicious scrapers often ignore robots.txt; use CDN/WAF/rate limiting too. User-agent: HTTrack User-agent: wget User-agent: WebCopier User-agent: WebZIP User-agent: Offline Explorer Disallow: / # === Sitemap === Sitemap: https://www.nvhousehunt.com/sitemap_index.xml