GEO for WooCommerce and WordPress stores: what to check and fix
GEO for a WooCommerce or WordPress store starts with access, then content. First make sure AI crawlers can reach and read your pages: the robots.txt they actually receive, the "Discourage search engines" and WooCommerce "Coming soon" switches, Cloudflare's AI crawler settings, and whether your text needs JavaScript to appear. Then publish one post per real shopper question. Access problems usually take days to fix; answer pages usually take weeks to show up in AI answers.
GEO (generative engine optimization) means making sure assistants like ChatGPT, Google AI Overviews and Perplexity can find your pages, read them and cite them when shoppers ask questions. These assistants search the web while they answer, then send a crawler (a program that fetches pages on their behalf) to read what they found. The catch on WordPress is that every layer can change what a crawler sees: WordPress core, WooCommerce, your plugins, your host, a caching plugin, a CDN. The setting you see in the dashboard is not always what the crawler gets. Below is what we check at each layer and what we change.
How we checked: anything marked "real output" comes from three fresh test stores we ran locally with WordPress Playground on 2026-10-08 (WordPress 7.1.3, WooCommerce 11.2.0, default Twenty Twenty-Five theme, one product priced at $399). One kept the defaults, one had "Discourage search engines" ticked, one was in Coming soon mode for the whole site. Everything else links to official WordPress, WooCommerce, Google, Cloudflare and AI company documentation.
Which robots.txt do crawlers actually get?
robots.txt is the plain text file at the root of your domain that tells crawlers where not to go. A WordPress site usually has no such file on disk. When someone requests /robots.txt, WordPress builds a virtual one with do_robots(): it blocks /wp-admin/, allows admin-ajax.php, and blocks no AI crawler. Since WordPress 5.5 it also ends with a line pointing to the built-in sitemap (do_robots() reference, WordPress 5.5 sitemaps announcement). WooCommerce adds a few lines of its own. This is what our default test store served:
User-agent: * Disallow: /wp-content/uploads/wc-logs/ Disallow: /wp-content/uploads/woocommerce_transient_files/ Disallow: /wp-content/uploads/woocommerce_uploads/ Disallow: /*?add-to-cart= Disallow: /*?*add-to-cart= Disallow: /wp-admin/ Allow: /wp-admin/admin-ajax.php Sitemap: http://127.0.0.1:9400/wp-sitemap.xml
Real output, fresh local test store, 2026-10-08. The first five Disallow lines come from WooCommerce and cover log, temporary and download folders plus add-to-cart URLs (WooCommerce source, class-woocommerce.php). The rest is WordPress. Nothing here blocks an AI crawler.
Four layers can change that file, and the crawler sees the combined result:
| Layer | What it does | Watch out for |
|---|---|---|
| WordPress and WooCommercevirtual file | Generated on each request. Blocks the admin area and a few WooCommerce folders. | Blocks no AI crawler. Usually fine. |
| Other pluginsSEO and security plugins | Can rewrite the virtual file through the robots_txt filter, or write a physical file to the site root. | When several plugins touch it, trust what the browser shows, not any one settings screen. |
| A physical file in the rootsomeone uploaded robots.txt | The web server sends that file and WordPress never builds its own. | Robots settings inside plugins stop working. You can change them all day and nothing happens. |
| Cloudflaremanaged robots.txt | Puts its own block of rules in front of yours, blocking a list of AI training crawlers. WordPress's virtual file returns a 200, so it gets the same treatment (Cloudflare's documentation). | That block disallows Google-Extended, which keeps your content out of Gemini app answers. |
The basis for row three: WordPress documentation notes that the old robots.txt method of hiding a site "only works if WordPress is installed in the site root and no robots.txt exists", and the standard Apache rewrite rules WordPress ships only hand a request to WordPress when no real file matches it (Settings Reading Screen, WordPress .htaccess rules). Row four: Cloudflare says that when your origin's robots.txt answers with HTTP 200, it prepends its managed rules; when there is none, it creates one for you. Our test store's virtual robots.txt answered with 200 (real output, 2026-10-08) (Cloudflare managed robots.txt).
So skip the plugin screens and open yourstore.com/robots.txt in a browser. Look for Disallow: / under the search crawlers: OAI-SearchBot, PerplexityBot, Claude-SearchBot, Googlebot and Bingbot. Which ones to allow and which training crawlers you can safely block is in our guide to checking AI crawler access.
Two switches that hide your store from crawlers
"Discourage search engines from indexing this site". It lives under Settings → Reading. Since WordPress 5.3, ticking it adds a noindex, nofollow robots meta tag to every page, and it turns off the built-in sitemap (Settings Reading Screen, WordPress 5.5 sitemaps announcement). noindex means "don't keep this page in the index". Google AI Overviews only use indexed pages (AI features and your website), and Bing says noindex keeps a URL out of Copilot too (Bing Webmaster Guidelines).
You won't spot it from robots.txt. On our test store with the box ticked, robots.txt only lost its last Sitemap: line; every page gained <meta name='robots' content='noindex, nofollow' />; and /wp-sitemap.xml returned 404 (real output, 2026-10-08). If your store was built on a staging site and then moved to the live domain, the setting may have moved with it. To check, view the source of any page and search for noindex.
WooCommerce "Coming soon" mode. It lives under WooCommerce → Settings → Site visibility. According to WooCommerce, stores that install WooCommerce and finish the setup wizard start in Coming soon mode: visitors who aren't logged in see a "coming soon" page instead of your store, and crawlers are visitors who aren't logged in. On a brand-new WordPress site (WooCommerce checks the fresh_site flag and whether the first admin account is less than a month old) it covers the whole site. Otherwise, for example when WooCommerce is added to a WordPress site you have run for years, it covers only the store pages (shop, cart, checkout and a few others) and your blog stays visible. Stores that are already live default to Live and are not affected. To turn it off, switch to Live. WooCommerce also warns that server caching can keep the old page up for a while after launch, so clear the cache (WooCommerce Coming soon mode).
This one is harder to catch. With Coming soon on, our test store's product page still returned 200 to a logged-out visitor and the page title was still the product name. The body was one line, "Demo Store is coming soon": no description, no price, no product structured data, and no noindex tag either (real output, 2026-10-08). The status code and title look fine. You have to read the body.
Is Cloudflare, your host or a security plugin blocking AI crawlers?
Plenty of WordPress stores sit behind Cloudflare, which can stop a crawler before robots.txt even comes into play. Per Cloudflare's documentation, from September 15, 2026 new domains default to blocking Training and Agent bots (Agent means real-time fetches for a person, such as a chat assistant opening your page) on pages that show ads, while Search stays allowed. What we set: allow Search. Allow Agent too; Anthropic says blocking its Claude-User agent "may reduce your site's visibility for user-directed web search" (Claude Help Center). Training is your call, except Google-Extended: block it and the Gemini apps can't use your content when they answer with live search results (Google's common crawlers). Where these settings live in the Cloudflare dashboard and how to test them yourself: our guide to checking AI crawler access.
Your host's firewall or a security plugin can also mistake an unfamiliar crawler for an attack. Perplexity's crawler documentation includes step-by-step allow rules for Cloudflare WAF and AWS WAF, and OpenAI publishes IP ranges for its crawlers and asks site owners to allow them (Perplexity Crawlers, Overview of OpenAI Crawlers). Every firewall has its own rules, so the answer is in your server or CDN logs: what status code did the real crawler get? After any change, purge your page cache and CDN cache, or crawlers may keep getting the old version.
Is your text in the HTML the server sends?
WordPress builds each page on the server with PHP when it is requested (WordPress documentation, Create pages), so themes and most pages arrive as finished HTML. That puts WordPress ahead of sites that render everything in the browser. The exceptions are page builder widgets, "load more" buttons and some blocks that fetch their content with JavaScript after the page loads. Google runs JavaScript, but still recommends server-side rendering because "not all bots can run JavaScript" (JavaScript SEO basics). The crawler documentation from OpenAI, Perplexity and Anthropic doesn't say whether their crawlers run JavaScript at all.
To check a page yourself, copy a sentence from it and fetch the raw HTML. A result of 0 means that sentence is not in what the server sends:
$ curl -sL https://yourstore.com/your-answer-page/ | grep -c "a sentence copied from the page"
Structured data: WooCommerce already outputs it, so check for duplicates
Structured data is code on the page that describes your product to search engines. WooCommerce core outputs it without any plugin: Product with an Offer on product pages (AggregateOffer when variations have different prices), Review and AggregateRating once there are reviews, BreadcrumbList for breadcrumbs, and WebSite when the shop page is your home page. It all goes into one JSON-LD block in the footer, rendered by the server (WooCommerce source, class-wc-structured-data.php). Our test product page had exactly one JSON-LD block with BreadcrumbList, Product and an Offer at 399.00 USD, matching the $399.00 shown on the page (real output, 2026-10-08).
The usual trouble is stacking. An SEO plugin or theme adds its own Product markup, so one product is described twice, sometimes with different prices or stock. The opposite also happens: WooCommerce attaches its product markup to the woocommerce_single_product_summary hook in the product template, so a theme or page builder that replaces the whole product template can drop it. Check both:
$ curl -sL https://yourstore.com/product/your-product/ | grep -o '"@type": *"Product"' | wc -l
Our test store returned 1, which is what you want. 2 or more means duplicate markup: keep one, and make sure its price matches the page. 0 means run Google's Rich Results Test, because the markup may be added by JavaScript. Keep expectations in proportion: Google says no special structured data is needed to appear in AI Overviews (AI features and your website), and Microsoft says it "may support clearer grounding" in Copilot without any promise of visibility (Bing Webmaster Guidelines, section 14). More in does structured data help with AI answers?
One sitemap, one IndexNow sender
A sitemap is a file listing the URLs on your site. Since WordPress 5.5 there is a built-in one at /wp-sitemap.xml, with up to 2,000 URLs per sub-sitemap by default. Plugins can switch it off with the wp_sitemaps_enabled filter and serve their own at a different address (WordPress 5.5 sitemaps announcement). We make sure exactly one sitemap is live, that the Sitemap: line in robots.txt points to it, and that it is submitted in both Google Search Console and Bing Webmaster Tools.
IndexNow is a protocol for telling search engines a URL was added, changed or deleted. Participating engines include Bing, Yandex and Naver; Google is not on the list (IndexNow participating engines). Copilot runs on Bing's index (Bing Webmaster Guidelines). On WordPress you can use the IndexNow Plugin published by the Bing Webmaster team, which submits URLs when posts are published, updated or deleted. Cloudflare's Crawler Hints also speaks IndexNow, and the IndexNow FAQ lists WordPress among the systems that support it natively or through SEO plugins. One sender is enough. Whether IndexNow helps ChatGPT see your page sooner, OpenAI has never said.
Where answer pages live: usually posts
WordPress has two kinds of content. Posts are dated, appear in your RSS feed and take categories and tags. Pages are undated, stay out of feeds and taxonomies by default, can be nested, and suit content that rarely changes, like an About page (WordPress documentation, Create pages).
We publish answer pages as posts by default. A post carries a publish date, an author and a category out of the box, and we require a named author and an accurate updated date on every answer page. Evergreen help that belongs in your menu, like a sizing guide or warranty terms, can be a page instead. Don't bury answers in product descriptions: the product page carries specs, the answer post handles one specific question, and the two link to each other. One real question gets one page; different phrasings of it don't get pages of their own. How to write one: how to write a page AI will quote.
Speed: only the part that affects crawling
Google crawls a site as fast as the server comfortably allows. Steady responses mean more crawling; slow responses, 5xx errors or 429 (too many requests) mean less. Google also says crawl budget mainly matters for sites with 1 million+ pages that change weekly, or 10,000+ pages that change daily, and calls those numbers a rough estimate (Optimize your crawl budget). For stores below that size, we check two things: whether slow database queries make crawler requests time out or error, and whether filter and sort URLs have multiplied out of control.
WooCommerce filtering and sorting use parameters such as orderby, anything starting with filter_, and min_price / max_price (WooCommerce source, class-wc-query.php). With a few attributes, the combinations quickly outnumber your real pages. Google's advice: if you don't need those URLs indexed, disallow them in robots.txt (Managing crawling of faceted navigation URLs). An illustration; adjust to the parameters your store really uses:
User-agent: * Disallow: /*?*orderby= Disallow: /*?*filter_ Disallow: /*?*min_price=
Page speed scores on their own are not a GEO deliverable for us.
What we change, and what you get at each step
The order is fixed: access first, content second, measurement throughout. If an earlier step is broken, nothing after it shows results.
Baseline and a layer-by-layer check
We run geo-score, then go through each layer above: the robots.txt crawlers really get, noindex and Coming soon, blocking at Cloudflare and the host, whether text is in the HTML, how many Product blocks each product page has, and how many sitemaps are live.
You get: an issue list naming the layer and the fix for each item, sorted into days, weeks and months of work.
Fix access
Allow the search crawlers, turn off any switch left on by mistake, keep one sitemap and submit it to Search Console and Bing, set up IndexNow, remove duplicate product markup and match it to the page, and block filter URLs if needed.
You get: a change log (which setting or file changed and how to undo it) and geo-score results before and after. Usually a few days. When AI crawlers come back to recrawl is up to them.
One question, one answer page
We pick specific questions shoppers really ask, from support tickets, product reviews and Search Console queries, and publish them as posts on your WordPress site with a named author and date, linked to and from the matching product pages.
You get: the pages, plus a fact sheet for each one. Prices, specs and warranty terms go live only after you confirm them.
Monthly measurement
We ask the same set of questions across the major assistants on a fixed schedule and record whether you are mentioned and which page is cited. geo-score runs again every month.
You get: a monthly report. Acceptance is the AIV score only (0 to 100 on geo-score's public rubric, which you can rerun yourself).
What access we need from you
- A WordPress administrator account, or your developer's help. Installing plugins, editing themes and fixing robots.txt need admin rights. Create a separate account for us under Users → Add User and delete it when we're done. Please don't share your own password.
- Less access if we only write answer pages. WooCommerce's Shop Manager role can change all WooCommerce settings and create posts and pages, with fewer rights than an administrator. WordPress's Editor role can publish and manage everyone's posts (WooCommerce roles, WordPress roles).
- A Cloudflare member account, if you use Cloudflare. We need to see the AI crawler settings, AI Crawl Control and status codes.
- Hosting panel or SFTP, only when needed. Only for a physical robots.txt in the root or a server cache purge. If you'd rather not share it, your developer or host can follow our checklist.
- View access to Google Search Console and Bing Webmaster Tools. To submit the sitemap and see which pages are indexed.
- A backup first. Before any change, you or your host take a full backup. If you have a staging site, we make changes there first.
Pricing: a diagnosis starts at US$1,190, one time. The quarterly service starts at US$2,990 per quarter and covers the four steps above. The number of answer pages depends on your category and is written into the contract. What each tier includes is on the pricing page.
Is a plugin enough? Your options compared
| Option | Covers | Doesn't cover | Pick it when |
|---|---|---|---|
| An SEO plugin plus an IndexNow plugindo it yourself | Sitemap, titles and descriptions, basic structured data, pinging Bing. | Won't check whether Cloudflare or your host blocks crawlers, or whether text is in the HTML. Doesn't write answer pages. | The three gate checks in geo-score's public rubric already pass and someone on your team writes content. The plugins are enough; you don't need us. |
| Your agency or developer works from a checklistpay per job | Every access problem in the layers above. | Usually not choosing questions, writing answer pages or measuring over time. | Your problems are access only and you have a developer. |
| Diagnosisfrom US$1,190, one time | Full report, list of shopper questions, how each assistant answers them today, prioritized fix list. | No hands-on changes. | You want the full picture first, then have your own team do the work. |
| Quarterly servicefrom US$2,990 per quarter | Access fixes, a batch of answer pages each quarter, monthly measurement and report. | No promises of rankings, traffic or placement. No head terms like "best running shoes". | You want to cover the long-tail questions in one product category. |
Service scope as of 2026-10-08. The pricing page is the current reference.
When you shouldn't hire us
- Access is already fine and you have a writer. If the three geo-score gate checks pass (criteria in the public rubric) and your team can write answer pages the way this guide describes, install the plugins, follow our free guides and keep the money.
- The store isn't live yet, or products and prices are still changing. Answer pages need stable facts. Launch first.
- You want to be "in ChatGPT in a few days", a ranking promise, or only head terms like "best running shoes". We don't sell those outcomes. How timing really works: how long until a new page shows up in Google and AI answers.
- You want hundreds of near-identical keyword pages, or bought links. We don't do that; see what we won't do.
- No dashboard access and no developer. Then access can't be fixed and answer pages can't go live. The diagnosis is the only thing worth buying; take the report to someone who can make the changes.
Common questions
I already have an SEO plugin. Do I still need GEO?
It depends on whether the access checks pass. An SEO plugin handles sitemaps, titles and markup. It can't see Cloudflare or your host's firewall, and it won't write answer pages. Run geo-score once: if the gate checks pass and someone writes your content, the plugin is enough.
Should I install an llms.txt plugin?
No rush. As of 2026-10-08, no major AI search product has said it reads a site's llms.txt when answering, and Google says its search ignores these files, so they neither help nor hurt there (Google's guide to optimizing for generative AI). It does no harm, but don't expect citations from it. Details in does llms.txt matter for an online store?
We built the store on a staging site and moved it to our real domain. What should we check?
Three things. Under Settings → Reading, make sure "Discourage search engines from indexing this site" didn't come across still ticked. Under WooCommerce → Settings → Site visibility, make sure it says Live. Then open yourstore.com/robots.txt and look for a Disallow: / carried over from staging. Purge your page and CDN caches, then rerun geo-score.
Can I put the answers in my product descriptions?
We don't recommend it. Keep specs and selling points in the description, and give a specific question its own post, for example "Which portable power station under $500 can run a 12V fridge overnight?" (an illustration). Link the two both ways. One page per question is easier for an assistant to quote as a whole.
Should I add FAQ structured data to product pages?
Not for its own sake. Google stopped showing FAQ rich results on May 7, 2026 (Google Search Central documentation updates). Write FAQs for readers. Keeping WooCommerce's own Product markup single and accurate matters more. Other markup types: does structured data help with AI answers?
Do you need my WordPress admin password?
No. Create a separate account for us under Users → Add User and delete it when we're done. If we're only writing answer pages, Shop Manager or Editor is enough.
How soon will I see results?
Access problems take days to fix, and you can rerun geo-score the same day. An answer page for a specific question usually takes weeks to get indexed and start appearing in AI answers. Head terms take months, and we don't sell that outcome. Why it works on that timescale: how long until a new page shows up.
Sources
- do_robots() (WordPress Developer Resources), checked 2026-10-08
- Settings Reading Screen (WordPress documentation), checked 2026-10-08
- Apache HTTPD / .htaccess (WordPress Advanced Administration Handbook), checked 2026-10-08
- New XML Sitemaps Functionality in WordPress 5.5 (Make WordPress Core), checked 2026-10-08
- Create pages (WordPress documentation), checked 2026-10-08
- Roles and Capabilities (WordPress documentation), checked 2026-10-08
- Coming soon mode (WooCommerce documentation), checked 2026-10-08
- Roles and Capabilities (WooCommerce documentation), checked 2026-10-08
- class-woocommerce.php (WooCommerce source, GitHub), checked 2026-10-08
- class-wc-structured-data.php (WooCommerce source, GitHub), checked 2026-10-08
- class-wc-query.php (WooCommerce source, GitHub), checked 2026-10-08
- Block AI Bots (Cloudflare documentation), checked 2026-10-08
- Managed robots.txt (Cloudflare documentation), checked 2026-10-08
- Crawler Hints (Cloudflare documentation), checked 2026-10-08
- Google's common crawlers (Google), checked 2026-10-08
- AI features and your website (Google Search Central), checked 2026-10-08
- Google's Guide to Optimizing for Generative AI Features on Google Search (Google Search Central), checked 2026-10-08
- Understand JavaScript SEO basics (Google Search Central), checked 2026-10-08
- Optimize your crawl budget (Google), checked 2026-10-08
- Managing crawling of faceted navigation URLs (Google), checked 2026-10-08
- Latest Google Search documentation updates (FAQ rich results not shown from 2026-05-07), checked 2026-10-08
- Rich Results Test (Google), checked 2026-10-08
- Bing Webmaster Guidelines, checked 2026-10-08
- IndexNow Plugin (Bing Webmaster team, WordPress plugin directory), checked 2026-10-08
- IndexNow FAQ (indexnow.org), checked 2026-10-08
- IndexNow participating search engines (indexnow.org), checked 2026-10-08
- Overview of OpenAI Crawlers (OpenAI), checked 2026-10-08
- Perplexity Crawlers (Perplexity documentation), checked 2026-10-08
- Does Anthropic crawl data from the web? (Claude Help Center), checked 2026-10-08
- geo-score (GitHub, MIT license), checked 2026-10-08
Find out which layer your WooCommerce store is stuck on
Send your store URL and get a free quick review within two days: robots.txt, noindex, Coming soon, Cloudflare and server-rendered text, each checked, plus what to fix first.