AI Search
How to Track Whether ChatGPT, Perplexity, and Google AI Mode Recommend Your Business
A repeatable process for logging AI assistant answers, calculating share of answer, diagnosing why you weren't named, and measuring downstream traffic from AI search.

Most local service businesses have no idea whether ChatGPT, Perplexity, or Google AI Mode currently recommend them for the questions their customers are actually asking. There is no dashboard for this the way there is for organic rankings — no Search Console equivalent that tells you "you appeared in 14% of relevant AI answers this month." If you want that number, you have to build it yourself.
The good news is that building it does not require expensive tooling. It requires a fixed set of real customer questions, a spreadsheet, and the discipline to run the same test the same way every month. This guide walks through constructing that prompt set, logging results consistently, calculating a usable share-of-answer metric, and — the part most guides skip — what to actually do when the answer is "they don't mention you."
This is the practical companion to answer engine optimization and AI search optimization for local services. Those explain what to fix; this explains how to know whether the fix worked.
Why AI citation share matters even without traffic numbers to prove it yet
AI assistants are increasingly the first stop for research-stage questions that used to start with a Google search: "what does a new roof cost in [city]," "how long does a mortgage pre-approval take," "do I need a permit to replace my water heater." When an assistant answers that question by name-checking a business, it is functioning as a recommendation engine, not just an information source — and the businesses it names get a trust transfer that a plain search result does not carry.
The volume of AI-originated traffic to any single local business site is still small relative to organic and paid channels for most categories as of 2026. That is exactly why tracking now matters: the businesses that establish a citation baseline and improve it steadily will hold a structural advantage as AI-assisted research keeps growing, while businesses that wait until the traffic is undeniable will be starting from zero against competitors with a two-year head start.
Citation tracking also functions as a diagnostic for your broader content and entity strength, independent of AI traffic. An assistant that cannot name you when your competitor's customer would type the exact same question into it is telling you something true about how legible your site and your reviews footprint are — a signal worth having even if you never look at an AI referral number again.
Step 1: Build a fixed prompt set of 20 to 40 real questions
The single biggest mistake in AI tracking is testing vague prompts like "best roofer in Minneapolis" and treating the result as meaningful. Real customers do not ask that way, and assistants respond very differently to specific, intent-bearing questions. Build your prompt set from actual language: pull from your CRM call notes, your Google Business Profile Q&A, your team's most common phone questions, and your FAQ pages.
Spread the prompts across three intent stages so you are not just measuring bottom-of-funnel visibility. Early-stage prompts are informational ("how do I know if my furnace needs to be replaced or repaired"). Mid-stage prompts are comparative or cost-driven ("average cost to replace a furnace in [city]"). Late-stage prompts are near-transactional ("HVAC company that does same-day furnace replacement in [city]").
Example set for a roofing company
"How do I know if my roof needs to be replaced instead of repaired," "average cost of a roof replacement in [city] 2026," "how long does a roof replacement take," "do roofers in [city] work with insurance claims for storm damage," "roofing companies in [city] with same-week estimates."
Example set for an HVAC company
"Why is my AC blowing warm air," "cost to install a new HVAC system in [city]," "how often should I replace my furnace filter," "emergency HVAC repair near [city] on weekends," "is it worth repairing a 15 year old AC unit."
Example set for mortgage lending
"How much do I need for a down payment on a house in [state]," "how long does mortgage pre-approval take," "difference between pre-qualified and pre-approved," "best mortgage lenders in [city] for first-time buyers," "can I get a mortgage with a 620 credit score."
Step 2: Log results the same way every month
Consistency is what makes this data usable over time. Set up a spreadsheet with one row per prompt-per-assistant-per-month, and run the full set on the same day each month across at least three assistants: ChatGPT, Perplexity, and Google AI Mode (add Copilot or Claude if your customers skew toward them). Use fresh sessions with no memory or personalization carried over where the tool allows it, since personalization can quietly bias results toward businesses you have already interacted with.
The columns below cover what you need to interpret results later, not just record them.
Recommended spreadsheet columns
Prompt text | Assistant | Date tested | Named? (yes/no) | Position (1st mentioned, 2nd, buried) | Facts attributed to you (services, pricing, area, hours) | Accuracy of those facts (correct/incorrect/outdated) | Competitors named | Sources cited (your site, directory, review platform, none visible) | Notes.
Step 3: Calculate share of answer
Share of answer is the percentage of your prompt set, per assistant, where your business is named at all — regardless of position or accuracy. If you are named in 9 of your 30 prompts on ChatGPT this month, your ChatGPT share of answer is 30%. Calculate it separately per assistant rather than averaging, because the three behave differently: Perplexity leans heavily on live web sources and cites more visibly, ChatGPT blends training data with browsing depending on the query, and Google AI Mode draws heavily on the same signals that drive local pack results.
Track a second, stricter number alongside it: accurate share of answer, meaning named with facts that are current and correct. A business that is named often but with a disconnected phone number or a service area that is wrong is not actually benefiting from the mention, and in the worst case it is sending a frustrated caller to a dead end.
Illustrative share of answer by assistant, month one baseline
Directional example only, not a benchmark drawn from a specific study. Typical starting points vary widely by category and market competitiveness.
- Perplexity22% of prompts where business was named
- Google AI Mode18% of prompts where business was named
- ChatGPT11% of prompts where business was named
- Copilot9% of prompts where business was named
Interpreting the results beyond a raw percentage
A low share-of-answer number is useful but incomplete on its own. Read each unnamed or poorly-named response for what it reveals. If a directory listing (Angi, HomeAdvisor, a chamber of commerce page) gets cited instead of any individual business, the assistant is defaulting to an aggregator because no single business has established enough distinct authority for that query — a fixable content and entity problem, not a dead end.
If a competitor is consistently named instead of you for the same prompts, compare what is different about their footprint: do they have a dedicated page answering that exact question, more third-party corroboration (review volume, local press, association listings), or clearer schema. If you are named but with wrong facts — an outdated price range, a discontinued service, the wrong service area — that is often worse than not being named, because it damages trust with the fraction of users who do call. Prioritize fixing wrong facts before chasing new mentions.
Diagnosing why you weren't named
Four causes account for most cases, and they call for different fixes.
Not crawlable or not indexed
If your relevant pages are blocked, noindexed, orphaned, or simply not in Google's index, no assistant relying on retrieval can surface you. Check indexation status before assuming a content problem.
No answer-shaped content
If your site never states, in plain sentence form, the answer to the exact question being asked, there is nothing for an assistant to lift. A page that talks around furnace replacement cost without ever giving a number or range is invisible to a cost-comparison prompt even if it ranks fine in traditional search.
Weak entity corroboration
If your name, address, service area, and category are inconsistent across your Google Business Profile, directories, and site — or if third parties rarely mention you at all — the assistant has less confidence you are the right entity to name for that location and query.
Thin or old reviews
Review volume and recency are a proxy for legitimacy that both search engines and AI assistants weight. A business with 8 reviews from three years ago reads as a weaker signal than one with steady recent volume, independent of actual service quality.
The fix loop
Run the diagnosis above prompt by prompt, group the failures by cause, and fix in order of leverage: indexation and crawlability first (it blocks everything else), then answer-shaped content for your highest-volume unanswered prompts, then entity corroboration (schema, consistent NAP, directory cleanup), then review velocity. Re-test the same prompt set the following month rather than a new one, so the comparison is clean.
Treat this as a loop, not a project with an end date. Assistants change models and retrieval behavior on their own schedule, competitors improve their own content, and your prompt set should be refreshed every couple of quarters as customer language shifts — but the core discipline of testing the same questions the same way stays constant.
Measuring the downstream signals
Share of answer is a leading indicator; it does not by itself prove revenue impact. Pair it with downstream signals in your analytics. Most AI assistants that link out pass referral traffic tagged with their domain (perplexity.ai, chatgpt.com) in your analytics referral reports — segment this traffic and watch it over time relative to your citation improvements.
Because a meaningful share of AI-influenced visits arrive as branded or direct traffic rather than a clean referral (a user reads an AI answer, then opens a new tab and searches your name or types your URL), also track branded search volume and direct traffic as secondary indicators. Finally, watch organic conversion rate on the pages you have optimized for answer-shaped content — an increase here alongside a flat or improving keyword ranking suggests visitors are arriving better pre-qualified, which is consistent with AI-assisted research.
Tooling vs. manual tracking
A handful of paid platforms now automate prompt testing across assistants and produce citation dashboards. They save time at volume and are worth considering once you are tracking more than one location or a large prompt set across many assistants monthly. For a single-location business testing 20 to 40 prompts across three assistants, manual tracking in a spreadsheet is entirely sufficient and has the advantage of forcing you to actually read every answer rather than skim a score.
Whichever approach you use, resist the temptation to test dozens of assistants at low volume each. Depth on the two or three assistants your customers are most likely to use beats shallow coverage of everything on the market.
What a realistic 90-day improvement looks like
Expect modest, uneven movement in the first 90 days rather than a dramatic jump. Indexation and schema fixes can show up in citation checks within two to four weeks once crawled. Answer-shaped content additions typically take four to eight weeks to be reflected, since they depend on the page being re-crawled and, for some assistants, incorporated into a retrieval index rather than fetched live. Entity corroboration and review velocity move more slowly and compound over months rather than weeks.
A reasonable target for a business starting near zero is a share-of-answer improvement into the high single digits or low teens on at least one assistant within 90 days, concentrated on the specific prompts you targeted with new content — not a blanket lift across every query in your set. Treat that as validation that the loop works, then widen the prompt set and keep going.
Frequently Asked
Questions & answers
How many prompts do I need to get a meaningful reading?
Twenty to forty prompts per assistant is enough for a single-location business to see meaningful patterns without the testing burden becoming unsustainable monthly. Fewer than fifteen tends to produce noisy month-to-month swings that are hard to act on.
Should I test with a logged-in or logged-out account?
Test logged out or in a fresh session whenever the assistant allows it, since personalization and memory can bias results toward businesses you have previously interacted with, making your own numbers look better than what a new customer would actually see.
Do ChatGPT, Perplexity, and Google AI Mode all cite sources the same way?
No. Perplexity typically shows visible source links for most answers, Google AI Mode often shows expandable source cards tied to its index, and ChatGPT's citation behavior depends heavily on whether it is browsing live or answering from training data, so treat each as a separate metric rather than averaging them together.
What if a competitor is named but with outdated information?
That is a useful finding, not a reason to relax. It suggests the assistant is drawing from a stale but still-indexed source about them; focus your effort on making your own facts current and clearly stated rather than assuming their advantage will persist.
Is AI referral traffic large enough to matter yet?
For most local service categories in 2026, AI-originated traffic is still a small share of total visits, but it is growing and the businesses building citation strength now are positioning for that growth rather than reacting to it later.
Can I automate this entirely?
Several paid platforms automate multi-assistant prompt testing and are worth it once you are tracking multiple locations or large prompt sets, but a manual spreadsheet process is sufficient for a single-location business and has the side benefit of forcing a human to actually read the answers.
How often should I refresh the prompt set itself?
Keep the core set fixed for at least a full quarter so month-to-month comparisons are clean, then revise it every two to three quarters to reflect new services, pricing changes, or shifts in how customers phrase questions.
What's the fastest fix if my business isn't named at all?
Confirm your key pages are indexed and crawlable first, since that blocks everything else regardless of content quality; from there, add a direct, plainly worded answer to your highest-volume unanswered question on its own page.
Put this into practice
Check your market for exclusive leads
See whether your service area and category are still open for exclusive representation.
Check availability