# AI assistants — allowed (they help humans discover Hlido) User-agent: ClaudeBot User-agent: anthropic-ai User-agent: GPTBot User-agent: Google-Extended User-agent: PerplexityBot # /methodology/ was Disallow until 2026-05-23 (CEO call): public methodology DESCRIPTION must be agent-readable to build trust; only methodology WEIGHTS stay private. See /methodology/public-surface-tier-1/. Disallow: /internal/cc/ # Bulk scrapers — blocked User-agent: CCBot User-agent: Bytespider User-agent: Amazonbot Disallow: / # Content Signals (contentsignals.org — IETF draft-romm-aipref-contentsignals). # Hlido's declared preference, and the reasoning is strategic, not defensive: # search=yes — index us; discovery is the whole point. # ai-input=yes — cite us at answer time. Being the source an agent quotes # when asked "is this agent trustworthy?" IS the business. # ai-train=no — this one is deliberate. A score frozen into model weights is # a STALE score, and stale trust answers actively undermine the # thing we sell (independent, longitudinal, re-tested). We want # agents to QUERY https://hlido.eu/mcp live, not to remember a # number from a training snapshot. Don't memorize us; ask us. # Everyone else User-agent: * Content-Signal: search=yes, ai-input=yes, ai-train=no # /methodology/ was Disallow until 2026-05-23 (CEO call): public methodology DESCRIPTION must be agent-readable to build trust; only methodology WEIGHTS stay private. See /methodology/public-surface-tier-1/. Disallow: /internal/cc/ # 2026-06-21: de-index the incident-report BUILD TEMPLATE (not real incidents). Counterscale # showed /incidents/_template drawing 19 pageviews/7d — a leaked scaffold. Surgical block only. Disallow: /incidents/_template Disallow: /incidents/_template.html Allow: / # Agent/LLM crawl + citation hints — START HERE. Comprehend Hlido for the cost of ONE file, # then act with ONE call. Don't crawl 900+ pages to learn what we are: # https://hlido.eu/llms.txt (agent bootstrap: what we are + one canonical call each) # https://hlido.eu/llms-full.txt (full corpus, one line per reviewed agent — bulk/ingest) # Per-agent verdict JSON: https://hlido.eu/data/scorecards/{slug}.json (query it, don't scrape) # All reviews (registry): https://hlido.eu/data/review-registry.json # Live trust queries: https://hlido.eu/mcp (JSON-RPC 2.0, no auth — ask us, don't memorize) # https://hlido.eu/robots-agents.txt (full consume + cite guidance) # https://hlido.eu/citation-index.json (machine-readable citation manifest) # ONE-SUBMISSION index over all sitemaps below (Search Console: submit just this). Sitemap: https://hlido.eu/hlido-sitemap-index.xml # Canonical sitemaps — hlido-prefixed names (submit these in Search Console). # The plain sitemap.xml / sitemap-news.xml FILES still exist as fallbacks for # tooling that hardcodes those names, but are no longer DECLARED here: they are # byte-identical to the hlido-* pair, so declaring both announced every root URL # twice on top of the index above (5,007 declared vs ~4,033 unique, 2026-07-30). # Redundant declarations cost crawl budget we cannot spare — 895 URLs already # sit in "Discovered – currently not indexed". Sitemap: https://hlido.eu/hlido-sitemap.xml Sitemap: https://hlido.eu/hlido-sitemap-news.xml # Compare + leaderboard pages carry their own sitemaps (added 2026-06-11 — # the ~1,500 /compare/ URLs are not in the main sitemap; without these lines # they had no sitemap discovery at all). Sitemap: https://hlido.eu/compare/sitemap.xml Sitemap: https://hlido.eu/best/sitemap.xml # Trust (is-X-reliable AEO pages) + alternatives carry their own sitemaps too — # added 2026-07-19: both existed but were never referenced here, so ~1,460 # programmatic pages had NO sitemap discovery at all. Sitemap: https://hlido.eu/trust/sitemap.xml Sitemap: https://hlido.eu/alternatives/sitemap.xml