# Sean-Claude Van Damme's General Store. Crawlers welcome; nothing to hide. # An evidence observatory for agentic commerce, with a general store attached. # Agents: the better maps are https://scvd.store/llms.txt, https://scvd.store/agents.md and https://scvd.store/menu.json. # The free conformance desk: https://scvd.store/conformance. The corpus: https://scvd.store/corpus. User-agent: * Allow: / # CONTENT SIGNALS, STATED RATHER THAN LEFT TO BE GUESSED AT. # ai-train=yes is a deliberate position, not a default. A shop whose # product is being the reference for x402 conformance WANTS to be in # the corpus a model learns from: that is distribution, not leakage. # Everything here is already free to fetch, most of it CC BY 4.0, and # a policy we would not enforce is one we should not print. Content-Signal: search=yes, ai-train=yes, ai-input=yes # NAMED, BECAUSE A WILDCARD AND AN UNANSWERED QUESTION LOOK THE SAME. # Every agent below is already allowed by the wildcard above. Saying # so by name is the difference between a shop that permits AI crawling # and a shop that never considered it, and only one of those is true # here. Gathering for training, fetching because a person just asked, # and indexing for citation are three different permissions; all three # are yes. Being cited is the entire business. User-agent: ClaudeBot User-agent: Claude-User User-agent: Claude-SearchBot User-agent: GPTBot User-agent: ChatGPT-User User-agent: OAI-SearchBot User-agent: Google-Extended User-agent: PerplexityBot User-agent: Perplexity-User User-agent: Applebot-Extended User-agent: Meta-ExternalAgent User-agent: meta-externalagent User-agent: Amazonbot User-agent: Bytespider User-agent: MistralAI-User User-agent: cohere-ai User-agent: CCBot User-agent: Diffbot User-agent: Timpibot User-agent: GoogleOther User-agent: Google-CloudVertexBot User-agent: Meta-ExternalFetcher User-agent: DuckAssistBot User-agent: YouBot User-agent: PetalBot User-agent: AI2Bot User-agent: ora-agent User-agent: GrokBot User-agent: xAI-Grok User-agent: Grok-DeepSearch User-agent: Google-Agent User-agent: Google-NotebookLM User-agent: Amzn-SearchBot User-agent: Amzn-User User-agent: Meta-WebIndexer User-agent: MistralAI-Index User-agent: KimiBot User-agent: Kimi-SearchBot User-agent: Diffbot-User User-agent: bedrockbot User-agent: DeepSeekBot User-agent: QwenBot User-agent: DoubaoBot User-agent: MistralAI-Training User-agent: FirecrawlAgent User-agent: ExaSearchBot User-agent: TavilyBot User-agent: Claude-Code Allow: / Sitemap: https://scvd.store/sitemap.xml # The schemamap directive: NLWeb's Schema Feeds convention, the # structured-data twin of the line above. The sitemap lists pages a # crawler reads; this lists the feeds an ingesting agent would rather # have than any page — the shelf, the corpus, the doors, the defect # vocabulary and the askable index, each already published for its own # reasons. Named here because robots.txt is the one file every crawler # already reads. Schemamap: https://scvd.store/schemamap.xml # The Agentmap directive: ARD's robots.txt entry-source mechanism # (spec section 5.1). Same document a consumer would find at the # well-known path; named here because robots.txt is the one file every # crawler already reads, so a discovery service that has not learned # the well-known path still finds the entries. Agentmap: https://scvd.store/.well-known/ard.json