# Barrhaven Bugle crawl policy # # Two different machines read this site and they are not the same thing. # # 1. ASSISTANTS and SEARCH/CITATION crawlers are WELCOME. When a reader asks an assistant a # question about Barrhaven, these fetch the page so the answer can name and link a local # source. They send readers back, and they credit the businesses we write about. # 2. MODEL-TRAINING crawlers are REFUSED. They take the text into a training set and return # nothing to the reporting, the advertiser, or the reader. Cited, not consumed. # --- Welcome: user-triggered assistants (a person asked a question right now) --- User-agent: ChatGPT-User Allow: / User-agent: Claude-User Allow: / User-agent: Perplexity-User Allow: / User-agent: DuckAssistBot Allow: / User-agent: MistralAI-User Allow: / User-agent: Meta-ExternalFetcher Allow: / # --- Welcome: search and citation indexes (how an assistant finds us at all) --- User-agent: OAI-SearchBot Allow: / User-agent: PerplexityBot Allow: / User-agent: Claude-SearchBot Allow: / User-agent: Applebot Allow: / # --- Refused: model-training harvesters --- User-agent: GPTBot Disallow: / User-agent: ClaudeBot Disallow: / User-agent: Google-Extended Disallow: / User-agent: Applebot-Extended Disallow: / User-agent: Meta-ExternalAgent Disallow: / User-agent: Amazonbot Disallow: / User-agent: Bytespider Disallow: / User-agent: CCBot Disallow: / User-agent: Google-CloudVertexBot Disallow: / User-agent: Diffbot Disallow: / User-agent: omgili Disallow: / User-agent: * Allow: / # Personal interview-form links (unlisted, tokened, noindex) — never index Disallow: /i/ # Internal lead-panel audit pages (unlisted, tokened, noindex) — never index Disallow: /picks/ # Reporter field notebook app (key-gated, noindex) — never index Disallow: /reporter/ # Tracked advertiser links. These are counters, not content: there is nothing here to index, and a # crawler following one books a click against a paying advertiser. Half of every click ever logged # (544 of 1,073 on 2026-08-19) was automated traffic reaching these paths, which we then had to # filter back out. Keeping compliant crawlers out at the source is better than discarding them after. Disallow: /c/ Disallow: /go/ Sitemap: https://barrhavenbugle.ca/sitemap.xml Sitemap: https://barrhavenbugle.ca/news-sitemap.xml