Short answer: Check three places. Open yourstore.com/robots.txt and look for rules that block search crawlers such as OAI-SearchBot, PerplexityBot, Claude-SearchBot, Googlebot or Bingbot. If your domain runs through Cloudflare, review its AI bot settings. Then request your homepage with curl using a crawler's user agent and see whether you get a 200 or a 403. Shopify's default robots.txt does not block AI crawlers, so problems usually come from rules someone added or a CDN placed in front of the store.

Allow the search crawlers; the training crawlers are your call

A crawler is a program that fetches web pages for its owner. Every major AI company separates crawlers that fetch pages to answer questions (search) from crawlers that collect content to train models (training). To show up in answers, let the search crawlers in. Per each company's documentation, blocking the training crawlers does not keep you out of answers.

AssistantAllow (search)Can block (training)Watch out for
ChatGPT searchOAI-SearchBotGPTBotBlock OAI-SearchBot and your pages won't be shown in ChatGPT search answers, apart from plain navigational links. The two settings are independent.
PerplexityPerplexityBot(PerplexityBot is not used for training)Perplexity-User fetches pages when a user asks; Perplexity says it generally ignores robots.txt.
Google AI Overviews and AI ModeGooglebotGoogle-ExtendedThese use Google's search index. There is no separate AI crawler.
Gemini appGooglebot, and do not block Google-ExtendedNoneGoogle says Google-Extended also controls whether Gemini Apps can use your content for grounding.
CopilotBingbot(no separate training token)Don't put noarchive or nocache on pages you want cited.
ClaudeClaude-SearchBot, Claude-UserClaudeBotAnthropic says it relies on robots.txt; blocking its IP addresses may not work reliably.
Apple (Siri, Spotlight)ApplebotApplebot-ExtendedIf robots.txt names Googlebot but not Applebot, Applebot follows the Googlebot rules.

Every cell comes from the company's own crawler documentation, linked under Sources.

Check robots.txt for rules that name these crawlers

robots.txt is a plain text file at the root of your site that tells crawlers where not to go. Open yourstore.com/robots.txt in a browser. You are looking for something like this (illustration):

User-agent: OAI-SearchBot
Disallow: /

User-agent: *
Disallow: /

The first group shuts out OAI-SearchBot. The second shuts out every crawler, Googlebot included. Any search crawler followed by Disallow: / means that door is closed.

On Shopify: every store has a default robots.txt. Shopify's help center lists what it blocks: admin, cart, checkout, account and order pages, plus filtered and sorted collection URLs. Products, collections, pages, blog posts and policies are open. On October 8, 2026 we opened the robots.txt of Shopify's own Dawn theme demo store: it has only a User-agent: * group and an adsbot-google group, and blocks no AI crawler.

If your file looks different, someone probably added a robots.txt.liquid template. You'll find it under Online Store > Themes > ⋯ > Edit code, in the templates folder. Shopify recommends editing it with Liquid rather than replacing it with plain text, and deleting the file restores the default. Shopify notes that changes are instant, but crawlers don't always react immediately (Shopify's documentation).

On WordPress and WooCommerce: the robots.txt that WordPress generates by default only blocks /wp-admin/. Plugins, hosts and CDNs can change it, so trust what you actually see in the browser.

Check whether Cloudflare is blocking them for you

Many stores sit behind Cloudflare, a CDN (a layer in front of your site that speeds it up and filters traffic). It can turn crawlers away no matter what robots.txt says.

  • Since July 1, 2025, Cloudflare asks every new domain at sign-up whether to allow AI crawlers, and describes this as blocking AI crawlers by default (Cloudflare's press release).
  • Cloudflare now sorts AI bots into three groups: Search (building an index to answer questions later), Agent (acting in real time for a person, such as a chat assistant fetching a page) and Training. From September 15, 2026, the default for new domains blocks Training and Agent bots on pages that show ads, and leaves Search allowed (Cloudflare's documentation). Crawlers that do both search and training are blocked by any setting that blocks training.
  • Its managed robots.txt option adds rules blocking AI crawlers to the top of your file. The example in Cloudflare's docs blocks Amazonbot, Applebot-Extended, Bytespider, CCBot, ClaudeBot, Google-Extended, GPTBot and meta-externalagent. The Google-Extended line affects the Gemini app. If you want to be cited there, leave this option off and write your own robots.txt.

Where to look: in the Cloudflare dashboard, go to Security > Settings, filter by Bot traffic, and check whether each of the three AI bot policies is set to Block or Allow, and whether Bot Fight Mode is on. Then open AI Crawl Control. The Crawlers tab shows Allow or Block for each crawler; the Metrics tab breaks requests down by status code, for the last 24 hours on the free plan (Cloudflare's documentation).

On Shopify: Shopify says bot management at the network layer is handled for you and you don't need to take any action. It does not recommend putting a proxy in front of Shopify and won't support that setup. So for a Shopify store, the question is whether you added Cloudflare in front of it yourself.

Example: a store owner on the Shopify Community asked how Cloudflare's default AI crawler blocking, effective July 1, 2025, might affect Shopify stores. Per Shopify's documentation, if you didn't add Cloudflare yourself, Shopify handles that layer.

Request a page as a crawler with curl: a rough test only

curl is a command-line tool that ships with macOS and Linux. The text after -A is the user agent, the name a crawler sends with each request. These two are copied from OpenAI's and Perplexity's documentation:

$ curl -s -o /dev/null -w "%{http_code}\n" \
    -A "Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/131.0.0.0 Safari/537.36; compatible; OAI-SearchBot/1.4; +https://openai.com/searchbot" \
    https://theme-dawn-demo.myshopify.com/
200
$ curl -s -o /dev/null -w "%{http_code}\n" \
    -A "Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; PerplexityBot/1.0; +https://perplexity.ai/perplexitybot)" \
    https://theme-dawn-demo.myshopify.com/
200

Real output from Shopify's Dawn demo store, October 8, 2026. Swap in your own domain.

  • 200: the request got through.
  • 403, 503 or a verification page: it was blocked. Replace -o /dev/null -w … with -sI to see the response headers. Cloudflare's docs say a cf-mitigated: challenge header means you were served its challenge page.

Why this is only a rough test: Cloudflare verifies real crawlers by cryptographic signature, published IP ranges or reverse DNS, not by user agent alone, and your laptop's IP is not on OpenAI's list. So the result can be wrong either way: you may be blocked while the real crawler gets in, or get a 200 while an IP-based rule still stops the real one. The final word is the status code real crawlers get in your server or CDN logs. OpenAI, Perplexity and Anthropic publish IP lists for checking that a logged request is genuine.

The quick way: one command that checks all of it

geo-score is our free, open-source checker (MIT license; all it needs is Python 3). It reads your robots.txt, requests sample pages as ten different search crawlers to see whether any is treated differently, and checks whether the page text is in the HTML or only appears after JavaScript runs:

$ curl -sL https://raw.githubusercontent.com/jianruntech/geo-score/v1.8.0/cli/geo_score.py \
    | python3 - yourstore.com --brief

The first group in the report, Reachable, covers exactly this. That group for the same demo store:

  Reachable 13/15
   ✓ Crawlers allowed in robots.txt   5/5
   ✓ Reachable to retrieval agents    5/5
   ◐ Main content server-rendered     3/5

Real output, Reachable group only: Shopify's Dawn demo store, geo-score 1.8.0, October 8, 2026. The first two lines are full marks, so the door is open. The third is partial: only some of the sampled pages had their main text in the HTML the server sent. geo-score also runs from your own machine, so the same caveat as curl applies.

Found a block? What to change

  1. robots.txt: remove Disallow: / under any search crawler. On Shopify, edit robots.txt.liquid, or delete it to go back to the default.
  2. Cloudflare: set the Search policy to Allow, and set the search crawlers from the table to Allow in AI Crawl Control. If Bot Fight Mode is challenging them, note that Cloudflare says WAF custom rules cannot bypass it. Turn it off, or switch to Super Bot Fight Mode and set Verified bots to Allow.
  3. Training crawlers: your choice. The one exception is Google-Extended, which affects the Gemini app.
  4. Re-test the same day: run geo-score and curl again. OpenAI and Perplexity say robots.txt changes reach their systems within about a day. When they next crawl your pages has no official answer; see how long until a new page shows up.

Common questions

If I block GPTBot, will my store disappear from ChatGPT?

No. OpenAI says GPTBot is for training and OAI-SearchBot is for search, and each setting is independent. You can block GPTBot and allow OAI-SearchBot.

I'm on Shopify. Should I add Cloudflare to manage bots?

No. Shopify says it handles network-level bot management and does not recommend a proxy in front of Shopify. What you control is robots.txt.liquid and the settings under Sales channels > Agentic in your admin, covered in Shopify Agentic Storefronts: what to check.

curl returned 200. Am I fine?

Not necessarily. It was a test from your own machine, and your logs are the real record. Also, an open door is not enough if the page text only loads through JavaScript; a crawler may receive a page with little text. geo-score checks this separately.

Want to know where your store is stuck? Run geo-score yourself, or send us your store URL and get a free quick review within 48 hours.

Sources

  1. Overview of OpenAI Crawlers (OpenAI), checked 2026-10-08
  2. Perplexity Crawlers (Perplexity docs), checked 2026-10-08
  3. Google's common crawlers (Google), checked 2026-10-08
  4. AI features and your website (Google Search Central), checked 2026-10-08
  5. Bing Webmaster Guidelines, checked 2026-10-08
  6. Does Anthropic crawl data from the web? (Claude Help Center), checked 2026-10-08
  7. About Applebot (Apple Support), checked 2026-10-08
  8. Editing robots.txt.liquid (Shopify Help Center), checked 2026-10-08
  9. robots.txt of Shopify's Dawn theme demo store, checked 2026-10-08
  10. do_robots() (WordPress Developer Resources), checked 2026-10-08
  11. Cloudflare press release, July 1, 2025, checked 2026-10-08
  12. Block AI Bots (Cloudflare docs), checked 2026-10-08
  13. Managed robots.txt (Cloudflare docs), checked 2026-10-08
  14. Get started with AI Crawl Control (Cloudflare docs), checked 2026-10-08
  15. Verified bots (Cloudflare docs), checked 2026-10-08
  16. Bot Fight Mode (Cloudflare docs), checked 2026-10-08
  17. Super Bot Fight Mode (Cloudflare docs), checked 2026-10-08
  18. Detect a Challenge Page response (Cloudflare docs), checked 2026-10-08
  19. Cloudflare's new default AI-crawler blocking policy (Shopify Community), checked 2026-10-08
  20. geo-score (GitHub, MIT license), checked 2026-10-08