Short answer: Pick 20 questions your buyers really ask. Every week, on the same day and with the exact same wording, ask each one in ChatGPT, Perplexity, Google (AI Mode and AI Overviews), Gemini and Copilot, logged out where you can. Note the date and your location, then three things: did the answer mention you, did it cite your site, and which page. Judge the share over several weeks, never a single answer.
Pick 20 questions and freeze them
These questions are your exam. Change the exam every week and the scores can't be compared. Take them from buyers:
- Support emails, live chat logs and product page Q&A, in the buyer's own words.
- Long queries in the Google Search Console Performance report that already bring you impressions.
- Reddit threads, forums and YouTube comments where people ask about your category.
Write them the way a shopper types, and mix the types: what to buy, comparisons, price, and plain descriptions of a situation. Three examples (illustrative, not from a real log):
- Example: best portable power station for van life under $500
- Example: quietest cat water fountain for a small apartment
- Example: is a LiFePO4 battery worth it for weekend camping
Keep questions that contain your brand name ("yourstore vs rival") in a separate group. With your name in the question, the answer cites you almost every time, so counting them only flatters the number. geo-score sets them apart too.
Ask every assistant under the same conditions
Answers shift with the account, chat memory and location. If those change, you are measuring them, not your store.
- Stay logged out where you can. OpenAI says people who are not signed in can use ChatGPT's web search, and that with memory turned on, ChatGPT may use saved memories when it rewrites your question into searches. If an assistant requires a login, use a clean account with no chat history and memory off.
- Same device, a private window, a new chat for every question. Paste the question exactly. No follow-ups, no edits.
- Write down the date and your location. ChatGPT estimates your rough location from your IP address, and a VPN changes that estimate. If your team works outside the US and sells to US shoppers, pick one US location and keep it.
- Check on the same weekday.
Check Google by hand. Open Google in a normal browser, read the AI Mode answer, then look for an AI Overview above the regular results and whether you are in it. Do not script it or use automation to pull Google results: Google's spam policies call sending automated queries to Google, including scraping results to check rankings, machine-generated traffic. Not every question gets an AI Overview, so log that too.
Log three things for every answer
A mention and a citation are different. A mention is your brand name or domain in the answer text. A citation is a link to your site in the sources or footnotes. An assistant can name you without linking, or list your page as a source without naming you in the text. Log them separately.
| Date | Location | Assistant | Question | Mentioned | Cited | Which page | Who else is cited | Notes |
|---|---|---|---|---|---|---|---|---|
| Oct 12 | US (VPN) | ChatGPT | Q03 | Yes | No | None | reddit.com, a review site | Quoted an old price |
| Oct 12 | US (VPN) | Google AI Overview | Q03 | No | No | None | No Overview shown | Checked by hand |
| Oct 12 | US (VPN) | Perplexity | Q07 | Yes | Yes | /pages/van-life-power | A van-life forum thread |
The three rows are illustrative, not a real log.
Don't skip "who else is cited". After a few weeks it shows which review sites, forums and videos the assistants lean on in your category: the places worth your attention next. How to show up there honestly is in our Reddit and YouTube guide.
Twenty questions on five assistants is 100 answers a week. One spreadsheet tab per week is enough. After a few weeks, count how many answers mention you, how many cite you, and which page gets cited most.
Why one screenshot proves nothing
Ask the same question a few minutes later and the cited sources can be quite different. That is how these systems work, not a sign that something broke on your site. The geo-score documentation says it plainly: asked again, the same question gets a largely different set of sources. OpenAI's own help page warns that search results and citations can be incomplete, outdated or incorrect.
So the founder who asks once on a phone and doesn't see the store has drawn one ticket from a lottery. A screenshot of "AI recommends you" sent as proof of results is one ticket too. One or two more mentions than last week is not a change either. The log turns "I have a feeling" into a share you can follow across weeks.
Save time with geo-score's --ask and watch, on your own API keys
geo-score is our open-source tool (MIT license, see the README). Level 1 scores your site with no key. Level 2, --ask, puts one question to the assistants now and shows who they cite. Level 3, watch, asks a fixed set every week, saves the answers and compares runs. Levels 2 and 3 use your own OpenAI, Perplexity, Gemini, Anthropic or OpenRouter keys: you pay those providers directly, each run logs its requests, tokens and cost, and geo-score reads keys from the environment and never writes them anywhere.
$ git clone --branch v1.8.0 --depth 1 https://github.com/jianruntech/geo-score && cd geo-score/cli $ export OPENAI_API_KEY=… PERPLEXITY_API_KEY=… $ python3 geo_score.py yourstore.com --ask 'best portable power station for van life under $500' $ python3 geo_score.py watch init --brand YourStore --domain yourstore.com --competitor "Rival=rival.com" # put your 20 questions in queries.csv, then: $ python3 geo_score.py watch run --dry-run $ python3 geo_score.py watch run $ python3 geo_score.py watch diff
--dry-run makes no calls and costs nothing. watch diff only calls something a change when it beats normal run-to-run noise. An assistant with no key is reported as not measured, never as zero.
Why don't API answers match the apps? The APIs are each provider's search-enabled developer interface; the consumer apps use other models, prompts and personalization. Compare API runs with API runs, never with answers collected by hand. And Google AI Overviews and AI Mode can't be reached through these APIs at all, so the Google column stays a manual check.
What your analytics can and can't tell you
The weekly log is your main record. Analytics help, but they won't split traffic cleanly by source.
- Google Analytics 4 has an "AI Assistant" channel for visits referred by ChatGPT, Gemini, DeepSeek, Copilot, Grok and similar. It excludes Google's AI Overviews and AI Mode.
- Direct, in GA4's words, is a visit from a saved link or a typed URL. A shopper who reads your name in an AI answer and types your URL lands here, as does a click with no referrer.
- utm_source=chatgpt.com. OpenAI says ChatGPT adds
utm_source=chatgpt.comto links clicked in its search results. Filter GA4 by session source to see them. It only covers clicks. - Shopify. If you use Agentic Storefronts, orders from AI channels show channel or referrer attribution in your Shopify admin.
- Search Console has a Generative AI performance report, open to all sites since August 31, 2026: how often links to your site were shown in AI Overviews and AI Mode, by page, country and device. Impressions, not clicks; it appears once you have enough of them.
- Bing Webmaster Tools has an AI Performance report: how often Copilot and Bing's AI summaries cite you, which pages, and the phrases the AI searched with.
- Ask at checkout. The cheapest way to catch what analytics miss is to ask the buyer. Put one multiple-choice question on the thank-you page or in the confirmation email ("Where did you first hear about us?") and include "ChatGPT or another AI assistant" as an option. B2B sites can add it to the inquiry form as an optional field, or ask in the first reply. People skip it and misremember, so watch how the share moves, not the exact count.
Example: a Shopify owner posted on r/shopify that a customer said they first found the product through ChatGPT, yet the order showed up as direct traffic. That gap is normal. The weekly log tells you whether AI mentions you; the checkout question tells you how many orders came that way. Read them together.
Common questions
Do I have to check every week?
No, every two weeks works. What matters is that nothing else changes: same questions, same weekday, same conditions. A longer gap just means a trend takes longer to show.
Are 20 questions enough?
Enough to start. With too few questions you can't tell a real change from noise. watch run --dry-run shows how big a change your set can detect, and when two runs share fewer than 6 questions, watch diff says there are too few to tell (geo-score's documentation).
AI mentions us but quotes the wrong price. What now?
Log it and check which page was cited. A real case from r/shopify: a crystal jewelry store had a shopper ask it to match the $24 ChatGPT quoted; the set sells for $34 (the original thread). The owner found ChatGPT was pulling launch-sale pricing from an old third-party blog post. Keep prices and specs on your own pages current, fix your own outdated pages, and ask third-party authors to update theirs.
How is your Monitoring plan different from doing this myself?
Same job, run by us: every week we ask the same 20 questions through four AI APIs, log whether you are mentioned, and send a one-page report each month, for US$590 per quarter. We don't touch your site. Because it runs on APIs, it can't see Google AI Overviews or AI Mode; that column still needs a person checking by hand. See pricing.
Want to know where your store gets stuck? Test it yourself with geo-score, or send us your store's URL and get a free quick review within 48 hours.
Sources
- Searching the web with ChatGPT (OpenAI Help Center), checked 2026-10-08
- Publishers and Developers - FAQ (OpenAI Help Center), checked 2026-10-08
- Spam policies for Google web search (Google Search Central, last updated 2026-08-28), checked 2026-10-08
- Generative AI performance report (Search Console Help), checked 2026-10-08
- Default channel group (Google Analytics Help), checked 2026-10-08
- Shopify agentic storefronts (Shopify Help Center), checked 2026-10-08
- Introducing AI Performance in Bing Webmaster Tools Public Preview (Bing Webmaster Blog, February 10, 2026), checked 2026-10-08
- geo-score README: Three levels, one tool; Levels 2 and 3; Scope, checked 2026-10-08
- geo-score cli/README: watch commands and queries.csv, checked 2026-10-08
- r/shopify: Customer said they found us through ChatGPT but Shopify shows direct traffic, read through the Arctic Shift public archive, checked 2026-10-08
- r/shopify: Anyone know why ChatGPT would quote us a price we have not charged since we first opened, over a year ago, read through the Arctic Shift public archive, checked 2026-10-08