Playbooks
The AI visibility audit: a step-by-step first assessment
Updated July 15, 2026
An AI visibility audit answers one question with evidence: when buyers ask AI assistants the questions that decide purchases in your category, does your brand appear in the answers? The method is to write roughly twenty real buyer questions, run each one several times through the engines your buyers use, store every answer, and count a mention only when your brand name actually appears, excluding questions that contain your own name. Run by hand it takes most of a day and produces your honest baseline. The stakes are documented: G2 research found 51 percent of B2B tech brands have zero citations across ChatGPT, Perplexity and Gemini. Reachroller runs this exact audit automatically from $29 per month, with a free three-day trial that covers a complete first report.
Why the audit is worth a day of your time
Most brands have never once checked what AI assistants say about them, while their buyers check constantly. G2's 2026 research found that 51 percent of B2B software buyers now start research with an AI chatbot more often than with Google, up from 36 percent just seven months earlier, and McKinsey reported in October 2025 that half of consumers already use AI-powered search intentionally as a primary way to find information and make buying decisions. Forrester's survey of 18,000 buyers puts AI somewhere in 94 percent of recent purchases. The conversations that decide your deals are happening inside answer boxes you have never read.
The distribution of outcomes in those answer boxes is brutal and specific. G2 found that 51 percent of B2B tech brands have zero citations across ChatGPT, Perplexity and Gemini: half the market simply does not exist where a growing share of buying research happens. Meanwhile 69 percent of buyers chose a different vendor than they originally expected because of AI chatbot output, and 33 percent bought from a brand they had never heard of before the AI named it. Which half of that story you are in is an empirical question, and the audit is how you answer it with evidence instead of vibes.
What follows is the full manual method: question list, engine selection, run protocol, scoring rules and interpretation. It costs a focused day for the first pass. It is also, transparently, the exact procedure Reachroller automates, so at each step we will note what the tool does differently, and you can decide at the end whether your second audit should be hand-run or scheduled. If you want the conceptual grounding first, the AI visibility primer covers why this discipline exists at all.
Step 1: write the twenty questions that decide deals
The audit is only as good as its question list, because you are sampling an infinite space of possible prompts and the sample has to represent what buyers actually type. Aim for about twenty questions. Fewer under-samples your category; many more becomes unmanageable once each question is run several times per engine. Source them from reality: the questions prospects ask on sales calls, the searches that bring signups, the questions in your category's communities, and the comparisons G2's data says dominate, since comparing vendor strengths and weaknesses is the top AI use case in software research at 41 percent.
Cover the intents a buyer moves through. Problem questions: how do I stop losing deals I never see. Category questions: what tools track brand mentions in AI answers. Evaluation questions: what should I look for in an AI visibility tool, is one worth paying for. Comparison questions: best X for a small team, alternatives to the category leader. Write each in plain buyer language, the slightly messy way a person types, because polished marketing phrasing is not what the engines are being asked.
Then label every question branded or unbranded, and be strict about it. A branded question contains your company or product name; an unbranded one does not. The distinction runs through everything downstream, because an answer to a branded question mentions you by construction, and counting those mentions inflates the score into meaninglessness. Keep a few branded questions in the audit, they surface wrong claims about you, but they live outside the headline number. The full argument is in branded vs unbranded prompts, and it is the single most common way first audits lie to their authors.
Step 2: pick engines by where your buyers actually ask
You do not need every engine on day one; you need the ones your buyers use. ChatGPT is non-negotiable: G2 found 72 percent of B2B software buyers use it during vendor evaluation, it is the dominant chatbot for software research at 63 percent, and Semrush's clickstream analysis shows it commanding roughly 92 percent of trackable LLM referral traffic. Whatever else is true of your category, ChatGPT's answers are part of your funnel.
Add Perplexity for research-heavy categories, since G2 found 44 percent of buyers use it during shortlisting, and add Gemini where your audience lives in Google's ecosystem, given the Gemini app's reported 750 million monthly users. The reason to audit engines separately rather than assuming one stands in for the rest is measured: cross-platform citation analyses find only about 11 percent of domains are cited by both ChatGPT and Perplexity. Winning one engine says almost nothing about the others, a point developed in which AI engines actually matter for your brand.
For a hand-run first audit, two engines is a sane scope and three is the ceiling; every added engine multiplies the run count. Decide up front whether web search is enabled and keep it consistent, because a retrieval-grounded answer and a memory answer are different measurements. When Reachroller runs this step it queries engines through official APIs only, ChatGPT live today with Claude, Gemini, Perplexity and Grok built and rolling out, which keeps conditions identical across runs in a way browser sessions never quite do.
Step 3: run the protocol, because one pass is noise
Here is where most self-audits quietly fail. You ask ChatGPT your twenty questions once each, feel relieved or horrified, and write up the result. But AI answers are probabilistic: SparkToro measured under a 1 percent chance that two identical ChatGPT runs return the same brand list. A single pass gives you twenty coin flips dressed up as a score. The honest protocol runs each question at least three times per engine, in a fresh session each time so chat history cannot contaminate the answer, ideally spread across two or three days.
Record everything as you go, and record it raw. The full answer text, pasted whole into a spreadsheet or document, along with the engine, the date, and the citations if the engine shows them. Summaries and checkmarks feel efficient and destroy the audit's value, because next month's comparison needs the actual words, and because wrong claims about your brand hide in sentences a checkmark never captures. The evidence file is the audit; the score is just its summary. This is also the step that hurts by hand: twenty questions, three runs, two engines is 120 answers to collect and file, which is most of the focused day.
Resist the urge to rephrase questions mid-audit when an answer disappoints. Changed wording is a changed measurement, and the baseline only works if the identical question can be re-run next month. If a phrasing feels wrong, note the improved version for the next audit cycle and finish this one consistently. The deeper measurement theory behind all of this, trend lines, stored evidence, volatility, is laid out in how to measure AI visibility without lying to yourself.
The audit scorecard, field by field
| Field | What to record | Why it matters |
|---|---|---|
| Question | The exact prompt text, in buyer phrasing | Small wording changes shift answers; the audit must be repeatable verbatim |
| Type | Branded or unbranded | Branded questions mention you by construction and must be excluded from the headline score |
| Engine and date | Which assistant, which day, web search on or off | Coverage differs per engine and answers drift over time |
| Full answer text | The complete response, pasted, not a summary | Stored evidence is what makes the score auditable and the recheck comparable |
| Your mention | Present or absent, literal name match only | Paraphrase and wishful reading inflate scores; literal matching keeps it honest |
| Competitors named | Every rival that appears, per run | Share of answer against rivals is the competitive read |
| Claims about you | Any factual statement, right or wrong | Wrong claims become their own fix queue |
| Cited sources | Every URL the engine shows | The citation list tells you which pages to fix or earn a place on |
One row per run, not per question. The per-question score is computed afterward from the runs, never eyeballed during collection.
Step 4: score it with rules you would accept from a vendor
Scoring is where self-deception creeps in, so borrow the strictest rules available. A mention counts only when your brand name literally appears in the answer text. A description that sounds like your product but never names you is a zero, however flattering; buyers cannot follow a mention that is not there. Compute your headline number from unbranded questions only, as the share of unbranded runs in which you were named. Then compute the same share for each competitor that appeared, which turns the audit into a share-of-answer table and usually delivers the audit's rudest surprise: the same two or three rivals recurring across questions you consider your home turf.
These rules are Reachroller's scoring rules, stated publicly on its methodology page: literal name match in stored answer text, branded questions excluded from the headline score, every number linked to the raw answer behind it. We publish them because a visibility score you cannot audit answer by answer is a number you have to take on faith, and the manual audit deserves the same standard. Apply them by hand and your baseline will be comparable to anything a tool later reports.
While scoring, harvest the two by-products that are often worth more than the headline number. First, every factual claim the engines made about your brand, marked right or wrong; the wrong ones are a correction queue with its own playbook in when AI gets your brand wrong. Second, the citation lists: which domains the engines leaned on, question by question. Those domains are your influence map, the specific pages where the next months of content and outreach work should land.
Step 5: read the baseline like a diagnosis
A finished audit sorts every unbranded question into one of three states, and each state has a different prescription. Questions you win, where you appear in most runs: protect these by keeping their supporting pages current, and note which of your pages the engines cite so you do not accidentally break them in a redesign. Questions a rival dominates: study what the engines cite for them, because the citations tell you whether the rival wins through their own content, through reviews and comparison sites, or through community threads, and each of those is a different counter-move. Questions nobody wins cleanly, where answers are generic and citations scattered: these are the cheapest ground to take, since you are competing with a vacuum.
Expect the baseline to be low, and do not let a low number read as failure. With half of B2B tech brands at zero citations per G2's research, a first audit that finds you named in a handful of unbranded runs already puts you ahead of the median, and the audit's purpose was never to flatter. It was to convert a vague anxiety into a ranked list of specific, fixable gaps, which is what a losing question with stored evidence and a citation trail is.
Read the citation lists in aggregate as well as per question, because they usually reproduce a known pattern at your scale. 5W Research measured Wikipedia at 13.15 percent and Reddit at 11.97 percent of ChatGPT citations in the U.S., over a quarter of the total between them, with a long fragmented tail behind the head. If your audit's citations show the same shape, your influence work splits accordingly: the head surfaces reward careful, rules-respecting presence work, while the tail rewards owned pages that answer specific losing questions better than the scattered sources currently being stitched together. If instead one niche review site or one community keeps appearing across your losing questions, you have found the single door most worth knocking on, and the audit has paid for itself in that one observation.
Prioritize the gaps by commercial weight rather than by wounded pride. A lost comparison question in your core segment outranks a lost broad-category question, because comparison is where G2's data says buying decisions concentrate. Pick the top three to five losing questions, fix those, and leave the rest for the next cycle. Focus beats coverage at this stage, and the fixes themselves, answer-first pages built to be citable, indexed properly, are the subject of the getting-mentioned playbook.
Step 6: turn the audit into a loop, or it expires
A single audit is a photograph of weather. Answers drift as engines re-crawl, models refresh and competitors publish, so the baseline starts expiring the day you finish it, and every fix you ship needs a recheck to prove it worked. The audit only becomes management information when it repeats: same questions, same protocol, same scoring, monthly at minimum and weekly once fixes are in flight. That cadence is where the honest cost of the manual method lands, because the focused day was tolerable once and becomes a tax at every repetition.
Set expectations for the second cycle before running it. Fixes grounded in live retrieval need their pages indexed first, which takes one to two weeks after submission through Google Search Console and Bing Webmaster Tools, so a recheck run five days after publishing measures nothing but your impatience. And even a working fix shows up as a shifted rate rather than a clean flip: a question you lost in all runs last month and win in two of three runs this month is a success, probabilistically speaking. Write those two rules into the audit doc itself, because the person running cycle two, possibly future you, will otherwise misread both delays and partial wins as failure.
This is the point where a tool earns its keep, and the math is not subtle. The manual audit consumes a day of someone senior every cycle. Reachroller's Starter plan is $29 per month for 400 credits and 25 tracked questions, where one credit is one AI answer: the whole audit, re-run on schedule, with answers stored, scores computed under the rules above, and a generated publish-ready fix page for any question you lose at ten credits each. If the budget is zero, stay manual or start with the no-cost options in free ways to check your AI visibility, including HubSpot's free one-time grader for a quick outside snapshot.
The fair caveat about our own tool: Reachroller is young, and its engine coverage beyond ChatGPT is still rolling out, so a maximally thorough audit today pairs its scheduled ChatGPT tracking with occasional manual passes on Perplexity and Gemini until those adapters ship. The three-day trial includes 50 credits and every feature with no card, which is exactly enough to run your first automated audit against the manual baseline you just built and see whether the numbers agree. If they do, you never hand-run the day-long version again.
Frequently asked questions
What is an AI visibility audit?+
A structured assessment of whether AI assistants like ChatGPT, Perplexity and Gemini mention your brand when asked the questions buyers use to make purchase decisions in your category. It produces a baseline score, a list of questions you lose, the competitors winning them, any wrong claims about your brand, and the sources the engines cite.
How many questions should an AI visibility audit include?+
Around twenty is the practical floor for a first audit: enough to cover the main buying intents in your category without making repeated runs unmanageable by hand. Each question should be run several times per engine, since answers vary, so twenty questions already means well over a hundred recorded answers across two or three engines.
Why do I have to run each question more than once?+
Because AI answers are probabilistic. SparkToro measured under a 1 percent chance that two identical ChatGPT runs return the same brand list, so a single run per question produces a score built on noise. Three or more runs per question per engine, spread across days, turns the noise into a readable rate.
Which AI engines should the audit cover?+
Start with ChatGPT: G2 found 72 percent of B2B software buyers use it during vendor evaluation, and it commands roughly 92 percent of trackable LLM referral traffic. Add Perplexity if your buyers research deeply, since 44 percent of buyers use it during shortlisting. Cross-engine analyses find only about 11 percent of cited domains overlap between ChatGPT and Perplexity, so results from one engine do not transfer to another.
What does a bad audit result actually look like?+
The common patterns: your brand absent from unbranded buying questions while two or three rivals recur, wrong or stale claims appearing when you are mentioned, and citation lists dominated by pages you have never touched. Absence is the most common. G2 research found 51 percent of B2B tech brands have zero citations across ChatGPT, Perplexity and Gemini.
Can I do an AI visibility audit for free?+
Yes. The manual method in this guide costs only time, and HubSpot's AEO Grader offers a free one-time snapshot score. The manual audit gives you depth the free checks lack, and a tool earns its keep when you want the audit repeated on schedule: Reachroller's three-day trial includes 50 credits, enough for a complete first report, with no card required.
How often should the audit be repeated?+
Monthly is a reasonable manual cadence; weekly is better once fixes are shipping, because you want the recheck close to the change. The baseline only becomes a trend line through repetition, and the trend is the honest measurement. This is the main reason to automate: the second and every later audit is where the hand-run method starts consuming days.
Sources referenced
- G2, B2B buyer AI research, 2026
- Forrester, 2026 Buyers' Journey Survey (18,000 global business buyers)
- McKinsey, consumer AI search adoption, October 2025
- SparkToro, consistency of repeated ChatGPT brand recommendations, 2025
- Semrush, ChatGPT traffic analysis, 17 months of clickstream data
- Profound and cross-platform citation analyses of domain overlap between AI engines, 2026
- 5W Research, ChatGPT citation share analysis, 2026
Run your first audit in minutes instead of a day.
Three days, 50 credits, every feature, no card. Enough for a full first report and a generated fix on your own domain.
Check my brand free