Guide

How to track whether ChatGPT recommends your business

By Asier Mugica, Founder · Published · Last updated

The short answer

You cannot answer this by asking ChatGPT once. AI answers change between runs and get shaped by your memory, your account and your location, so a single screenshot is a coin flip, not a measurement. The method: freeze 15 buying prompts in writing, pick your engines, run every prompt three times in a clean session, screenshot each run, and log whether you were named, who else was, and which sites got cited. Repeat monthly with the same method.

Why can't you just ask ChatGPT and screenshot the answer?

Because the answer you got is one sample from a distribution, and you have no idea how wide that distribution is.

Start with the machinery. Model output is not deterministic in practice, even when you ask it to be. Horace He and the team at Thinking Machines Lab put it plainly in their September 2025 write-up on inference nondeterminism: "even when we adjust the temperature down to 0 (thus making the sampling theoretically deterministic), LLM APIs are still not deterministic in practice". Google says the same about its own products: "AI Mode and AI Overviews may use different models and techniques, so the set of responses and links they show will vary".

The effect is bigger than most owners expect. Research published by SparkToro and Gumshoe in January 2026 ran buying prompts through ChatGPT, Claude and Google AI Overviews across 2,961 responses and concluded there is "<1 in 100 chance that ChatGPT or Google's AI, if asked 100X, will give you the same list". It comes from a company selling research software, so read the exact figure as directional. The direction matches our own logs.

Here is ours. In our own tracking, July 2026 (90 runs across 15 buying prompts), a Google AI Overview fired in 42 of 45 Google runs. Sounds stable. But for one prompt, "best marketing agency for med spas", the AI Overview did not fire in any of its 3 runs. Check that prompt alone and you conclude AI Overviews barely exist. Check any of the other fourteen and you conclude they are everywhere. Both conclusions come from real screenshots, and both are wrong.

The third source of noise is you. Memory personalizes what ChatGPT says: OpenAI describes it as remembering "useful context from your chats, files, and connected apps to personalize your experience", switchable in Settings, Personalization, Memory. Spend a year talking to ChatGPT about your own business and of course it names your business. You trained that answer. Location does the same thing more quietly on Google.

So the unit of measurement is not an answer. It is a rate across a fixed set of prompts, run often enough to compare to last month.

Step 1: Freeze 15 prompts real buyers type

Fifteen is deliberate: wide enough to cover how people shop for your service, small enough to rerun by hand every month without quietly giving up in month three. Pull them from evidence, not imagination. What people ask on the phone before booking, the wording in your reviews, the queries already bringing you clicks in Search Console. Then spread them across the four shapes a buying question takes:

  • Category plus city. "Best dentist for implants in Austin". The obvious one, and the most contested.
  • Problem first. "My crown fell out on a Sunday, who do I call in Austin". How people type when they are in trouble, and usually wide open.
  • Comparison and price. "Invisalign vs braces cost in Austin". These answers name providers as examples.
  • Qualifier prompts. "Gym in Austin with childcare and 5am hours". Narrow, high intent, easiest to win first.

Write the city into the prompt rather than relying on the location of whoever runs the test, and use the buyer's words, not keyword shorthand. Then freeze them: paste all 15 into a document, date it, and do not touch them again. Editing wording between months is the easiest way to fake improvement, and just as easy to do by accident. If you need a new prompt later, add it with its own start date and keep the original 15 reporting separately.

Step 2: Decide which engines count, and write that down too

"AI search" is not one place. Each engine reads different sources and answers differently, so the engine list is part of the method and gets frozen alongside the prompts.

Pick two to start. We use Google AI Overviews and Perplexity for the run panel, for one practical reason: both show their sources on screen, so every run hands you a citation list you can log. ChatGPT is the engine everybody asks about and is worth adding once your process is steady, but it is harder to log cleanly because so much depends on account state. Copilot and Gemini are reasonable third and fourth additions.

Two details people skip. Record the mode, because "ChatGPT with web search" and "ChatGPT answering from memory" are effectively different products, and only the first is driven by what OpenAI's search crawler OAI-SearchBot can reach. And if you ever add or drop an engine, note the date, because your panel-wide numbers before and after are not comparable.

Step 3: Run every prompt three times in a clean session

Three runs per prompt, per engine, per month. That is the floor. Fifteen prompts across two engines at three runs each is 90 runs, which takes us around two hours including screenshots.

Clean session means the answer is not shaped by anything about you:

  • A new chat for every single run. Never ask prompt 2 in the thread where you asked prompt 1. Earlier turns steer later ones, and you end up measuring your own conversation.
  • Memory off. Turn it off in Settings, Personalization, Memory, or use a Temporary Chat, which OpenAI states "won't access or create memories for personalization".
  • Custom instructions off too. The trap: OpenAI's own documentation notes a temporary chat "will still follow your custom instructions if they're enabled", so it alone does not give you a clean room.
  • Logged out, or a fresh browser profile. A signed-out window in a profile you do not use daily, not a private window on top of a signed-in session you forgot about.
  • One sitting, timestamps recorded. Run the panel in one block, and write down date and time. Engines ship changes without telling anyone, and the timestamp is what lets you explain a strange month later.

Three is a floor, not a target. If two of three runs disagree on whether you are named, run five for that prompt and record all five. Disagreement is information: you are on the edge of the answer, which is the most winnable position there is.

Step 4: Screenshot every run and log the same five fields

Screenshot every run including the sources panel, and name the file so it sorts: prompt number, engine, run number, date. It feels excessive on day one. In month three, when someone questions a number, the screenshots are the only reason the number survives. Then log these five fields, identically, every run:

FieldWhat you recordWhy it matters
Named or not Yes or no. Your business name in the body of the answer. The headline number. Everything else explains it.
Position First, second, third, or "mentioned in passing" if you appear outside the recommendation list. People act on the first name far more than the fourth. Third to first is real progress a yes-or-no column hides.
Competitors named Every other business named, in the order they appear. Your only benchmark. Names that repeat across runs are the incumbents to displace.
Domains cited Every source link the engine shows, by domain. The source diet: the column that turns measurement into a plan (step 6).
Run notes Did an AI answer fire at all? Did it refuse, or answer generically without naming anyone? Any wrong facts about you? "No AI answer appeared" is a result, not a blank cell. So is an answer with your hours wrong.

Five columns, one row per run. Ninety rows a month is a spreadsheet, not a data problem.

One distinction to keep straight from day one: being cited as a source (your URL in the citation panel) and being named as a recommendation (your business in the sentence) are different outcomes. Log them separately. Merging them is how a link gets sold to you as a win.

Step 5: Turn 90 runs into two numbers

Two numbers carry almost all the signal.

Presence rate, per prompt. Runs where you were named, divided by runs. Named in 4 of 6 runs is 67%. Calculate it per prompt before the panel average, because the average buries what you most need: which questions you own, which you are on the edge of, and which you are absent from. The edge cases are where next month's work pays off fastest.

Share of voice, across the panel. Your mentions divided by all business mentions in the same runs. If 90 runs produced 210 business names and 12 were yours, share of voice is 5.7%. That is the number that answers "are we winning, or is everyone winning", and the one that moves when a competitor gets aggressive.

We track a third, blunter number because it is easy to read at a glance: how many of the 15 prompts have a presence rate above zero. A count out of 15 flatters nobody. Do not confuse it with our guarantee bar, which is stricter and written into our terms: 3 of the 15 prompts at a presence rate of 40% or better by day 90.

What not to compute: an "AI ranking" or a single visibility score built from a formula nobody publishes. If you cannot rebuild a number from the raw log, it is not a measurement, it is a logo for a dashboard.

Step 6: Read the source diet, the part that tells you what to do next

The source diet is every domain the engines cited across your panel, counted and ranked. Ten minutes to build from the "domains cited" column, and the most actionable output of the whole exercise.

Presence rate tells you where you stand. The source diet tells you which pages on the internet are deciding the answer. Two findings from our own tracking, July 2026 (90 runs across 15 buying prompts):

  • Reddit was cited in 27 of the 45 Google AI Overview runs. Community threads are not a fringe source in this category, they are the majority source.
  • Perplexity's citations were dominated by agency listicles, the "best X agency" roundups from publishers like Thrive and First Page Sage. Whoever is on those lists is in the answer.

Read that as a work order. A roundup cited in half your runs that you are not on is next month's outreach. A thread cited constantly where your name never comes up is a reputation gap, not a website problem. A cited page carrying your old hours does more damage than any amount of writing on your own site undoes.

Watch for one pattern in particular: your own domain gets cited, but you are not named in the answer. The engine reads you and does not consider you the answer. That is a very different problem from being invisible, and it usually points at corroboration rather than crawling. Our guide on getting cited by ChatGPT and AI search covers what to change once the measurement tells you which of the two you have, and the gym version of the same diagnosis is in why your gym doesn't show up in Google AI Overviews.

Step 7: Re-run monthly, same method, or the series is worthless

Same prompts, same engines, same number of runs, same clean-session rules, same week of the month, written down as a procedure so a different person can run it and produce comparable numbers.

Change anything in the method and you have started a new series. Say so in the log, keep the old numbers visible, and do not draw a line between the two. A chart that goes up because the method changed is worse than no chart, because now you trust it.

Expect the first honest read at 90 days. Foundations take 30 to 90 days to land, engines refresh on their own schedule, and month-to-month wobble in a 90-run panel is normal. Report the misses next to the wins. A monthly report that only ever goes up is a report somebody is editing.

What can you measure free, and what needs a tool?

Honest version, from someone who runs these panels by hand and does not sell software:

What you want to knowFree, by handWhat a paid platform adds
Are we named for our 15 prompts? Yes. 90 runs, a spreadsheet, about two hours a month. Nothing you cannot do. It does the clicking.
Share of voice and competitors Yes, from the same log. Automatic name matching across hundreds of prompts and spelling variants.
Source diet Yes. Copy the citation links, count domains. Aggregation over months, deduplication, trend lines per source.
Daily or weekly tracking No. Nobody sustains this by hand. The first real reason to pay.
Hundreds of prompts, many locations No. The second real reason. Multi-location brands need it.
Evidence nobody can quietly edit Partly. Dated screenshots in a folder you control. Tamper-evident history and exports built for reporting.
Traffic actually caused by AI answers Partly. Referrals from chatgpt.com or perplexity.ai in your analytics. Some platforms stitch sessions to sources. Treat any "AI traffic" figure as an estimate.

We do not sell a tracking tool and there are no affiliate links on this page. For a single-location business running 15 prompts a month, a spreadsheet is genuinely enough.

One thing no tool and no analytics package gives you: a clean "AI search" number in Google's own reporting. Google states that sites appearing in AI features "are included in the overall search traffic in Search Console" inside the Web search type, and Search Console counts an AI Overview as a single position where clicking an external link "counts as a click". There is no AI filter to switch on. Add that most AI answers end with no click, and referral traffic is a floor on your visibility, never the measurement.

What mistakes invalidate the measurement?

Every one of these produces a number that looks fine and means nothing:

  • Asking once. The whole reason this guide exists. One run is an anecdote with a screenshot attached.
  • Rewording prompts between months. Even small edits change the answer. Your series becomes two unrelated series wearing the same chart.
  • Testing logged in with memory on. You are measuring your own history. The most common self-inflicted false positive, and why owners are so often sure they are recommended when they are not.
  • Counting a citation as a recommendation. A link in the sources panel is not your name in the answer. Track both, never merge them.
  • Comparing months measured differently. Different engine list, run count, person or rules. Not comparable, however tidy the chart looks.
  • Tracking yourself and nobody else. Competitor names in the same runs are your only benchmark. Without them you cannot tell a bad month from a hard category. And if the folder only contains wins, the folder is marketing.
  • Letting your agency own the only copy. The prompt list, the log and the screenshots live in your drive. If the relationship ends, the history is yours. Any vendor who resists that is telling you something.

Here is our own measurement, in public

We ran this exact method on ourselves before running it on anyone else, and published the result on our work page including the parts that make us look bad. Baseline frozen : GEO score 47 out of 100, share of voice 0.0 with zero mentions across 90 runs, 0 of 15 prompts at our own guarantee bar, and 0 clicks from 39 impressions in 28 days of Search Console data.

That is a four-week-old site measured honestly. We publish it because a method you only show when the numbers flatter you is not a method, and because the next measurement, late August 2026, needs a baseline nobody can quietly revise. Same 15 prompts, same engines, same three runs each, and the deltas go on that page whichever way they move.

So do you need an agency for any of this?

For the measurement, no. There is no secret method here, which is why we published ours in full. An owner or office manager with a spreadsheet and two focused hours a month can run this panel as well as we can, and should, whether or not anyone works for you. Knowing your own number makes every marketing conversation you have this year cheaper.

What we sell is the other half: doing the work that changes those numbers, and carrying the risk when it does not. The audit is $1,500 and credits 100% toward your first month. Retainers are Foundation at $1,500 a month, Growth at $2,500 and Dominate at $5,000, all listed on the pricing page. On Growth and up, the guarantee is measured against exactly the panel described on this page, frozen on day one: named in the AI answers for your core services within 90 days, or we keep working free. And we take one business per city, because these answers only name two or three names.

If you are weighing this against the SEO you already pay for, GEO agency vs SEO agency lays out what actually differs and when the honest answer is to hire nobody. If you would rather just see your starting point, the free GEO score gives you the first read in a few minutes.

Guide FAQ

What owners ask about tracking AI answers.

How many times should I run each prompt?
Three runs per prompt, per engine, per month is the floor. That is what makes the difference between an anecdote and a rate you can compare next month. Fifteen prompts across two engines at three runs each is 90 runs, roughly two hours of work. If two of the three runs disagree on whether you are named, run five and record all five.
Can Google Analytics or Search Console tell me if ChatGPT recommends my business?
No. Google folds AI Overviews and AI Mode into ordinary Search Console web traffic, so there is no AI filter to check, and a click out of an AI Overview looks like any other search click. Analytics can show referrals from chatgpt.com or perplexity.ai, but most AI answers end with no click at all. Referrals are a floor, not the measurement.
Do I need a paid AI visibility tracking tool?
Not for 15 prompts once a month. A spreadsheet, a browser and two hours cover it, and the dated screenshots are better evidence than most dashboards. Tools earn their price on frequency and scale: daily checks, hundreds of prompts, many locations, and a history nobody on your team can quietly edit.
Should I be logged in to ChatGPT when I test?
No. Logged in with memory on, you are measuring your own history, not the answer a stranger gets. Turn memory off in Settings, Personalization, Memory, or use a Temporary Chat, which does not use or create memories. Temporary chats still follow your custom instructions, so switch those off too, and put the city in the prompt instead of relying on your location.
What counts as a good share of voice?
There is no published benchmark worth quoting, and anyone who gives you one is guessing. The only honest comparison is your own baseline plus the competitors named in the same runs. Our own site started at 0.0, zero mentions in 90 runs. The bar we hold our client work to is being named for at least 3 of the 15 frozen prompts by day 90. Once you know where you stand, the next question is who should do something about it, and we compared the four kinds of provider in who can get your business recommended by ChatGPT.

Last updated: July 24, 2026

Want the first read without the two hours?

The free GEO score checks how visible your business is to ChatGPT, Perplexity and Google AI right now, and hands you the three fastest fixes. Run your own panel afterwards either way.

Get your free GEO score

Free. No sales call attached.

Get your free GEO score