MultiplierAI
ArticlesLog in
Back to Articles
Analytics

How to Run an AI Visibility Audit

A step-by-step AI visibility audit you can run this week without buying anything: prompt set, sampling, four metrics and the off-site target list.

M
MultiplierAI Research Team·September 3, 2026
In Brief
  • Core Answer: An AI visibility audit is a structured, repeatable measurement of whether AI engines mention and cite your brand across a defined prompt set — producing a baseline, a competitor picture, and a list of the third-party sources you are missing from.
  • Why It Matters: Without a dated baseline, every claim about improvement afterwards is unfalsifiable, including your agency's and your own.
  • Best For: Teams about to start GEO work, or about to buy a tool, who need to know their starting position first.

An AI visibility audit measures four things across a fixed prompt set and several engines: whether your brand appears, whether it is cited with a link, whether what is said is accurate, and which other brands and sources appear instead. It takes a few hours to run manually and produces the one artefact every later decision depends on — a dated baseline.

What follows is a process you can execute this week without buying anything.

Step 1: Build the Prompt Set

The prompt set is the audit. Everything downstream inherits its quality.

Sources to draw from, in order of usefulness: recorded sales discovery calls, support tickets, the query report in Search Console filtered to question-shaped queries, and the questions your sales engineers answer repeatedly. A keyword tool is a distant fifth, because it returns phrasings rather than questions.

Structure it in three tiers:

  • Brand (5–10): "What is [company]?", "Is [company] good?", "How much does [company] cost?", "What are alternatives to [company]?"
  • Category (15–30): the problem-shaped questions asked before vendor names are known.
  • Comparison (10–20): "best [category] tools", "[company] vs [competitor]", "[category] for [segment]".

Thirty to fifty prompts total is the practical range. Fewer is too noisy to trend; more is unsustainable to sample properly.

Step 2: Choose Engines and Sampling

Cover at minimum ChatGPT, Perplexity and Google AI Mode or AI Overviews. Add Claude and Copilot if your buyers are enterprise. If you sell in more than one country, each locale is a separate measurement — answers differ materially.

Sampling rules that make the numbers mean something:

  1. Three to five runs per prompt per engine. Generated answers vary; one run is a sample of one.
  2. Fresh sessions each time. Conversation history and personalisation contaminate results.
  3. Record the date, engine, locale and run number against every observation.
  4. Include one control prompt unrelated to your brand, run the same number of times, to establish the noise floor.

Step 3: Record the Right Fields

For every run, capture six things:

Field

Values

What it feeds

Mentioned

Yes / no

Mention rate

Cited

Yes / no, with URL

Citation rate; which pages work

Position in answer

First named / later / passing

Prominence

Description verdict

Accurate / drifted / wrong

Correction backlog

Brands named

Ordered list

Share of answer

Sources cited

All domains

Off-site target list

Also paste the full answer text. Flags tell you what changed; only the text tells you why.

Step 4: Compute Four Numbers

  • Mention rate = runs where you were named / total runs. Segment by tier — brand, category and comparison prompts behave differently and averaging them hides the problem.
  • Citation rate = runs where a link to your domain appeared / total runs. Usually much lower than mention rate; the gap is a retrievability signal.
  • Share of answer = your mentions / all brand mentions across the set. This is the competitive number and the one that matters for shortlist inclusion.
  • Accuracy rate = runs where the description was correct / runs where you were described. Below about eighty percent this is the most urgent finding in the audit.

Report each with its sample count. A mention rate of forty percent from ten runs and from two hundred runs are different claims.

Step 5: Diagnose

Four diagnostic patterns cover most results.

  • Low mention, low citation. You are not a category member in the model's view. This is a consensus problem: work the cited-source list.
  • High mention, low citation. The models know you but are not reaching your pages. Check crawler access per agent, rendering, and whether your pages answer questions in extractable form.
  • Decent mention, low accuracy. An entity and canonical-fact problem. Publish an unambiguous fact page, correct third-party records, consolidate naming.
  • Good on brand prompts, absent on category prompts. The most common B2B pattern. You are findable when asked for by name and invisible during the research that precedes it.

Each implies a different first action, which is why a single composite visibility score is the wrong output for an audit.

Step 6: Produce the Off-Site Target List

Count every third-party domain cited across the whole set and rank by frequency — how models choose citations explains why the list concentrates. In most categories four to eight domains account for the majority of citations — typically a review platform, a comparison site, one or two industry publications, and a community forum.

For each, record the mechanism of inclusion: a claimable profile, a listicle you could be added to, a directory record you can correct, a publication that accepts contributed pieces. That list, prioritised, is the highest-value output of the audit and the thing most audits fail to produce.

Step 7: Check the Technical Layer

Fast and binary. Confirm that Googlebot, GPTBot, OAI-SearchBot, ChatGPT-User, PerplexityBot, ClaudeBot and Bingbot are each permitted in robots.txt unless you have a deliberate reason otherwise. Confirm that your key pages render server-side. Confirm they are indexed. Confirm Organization schema is present and correct.

A blocked agent is the single most common cause of a mention-without-citation pattern, and it is often the result of a bot-mitigation rule nobody reviewed. The current crawler list changes; re-check quarterly.

Step 8: Set the Re-Measurement Date

An audit that is not repeated is a screenshot. Schedule the next full run at ninety days, with brand-tier prompts checked weekly in between because accuracy problems compound. Freeze the prompt set — changing it between runs makes the comparison meaningless — and version it if you must add prompts, keeping the original subset intact for continuity.

What an AI Visibility Audit Will Not Tell You

It measures visibility, not value. It will not tell you whether the visibility produced pipeline; that requires branded search baselines, self-reported attribution and referrer classification on your side. It will not tell you why a model chose a source. And it will not, on its own, produce improvement — the audit's job is to make the next quarter's work falsifiable. Connecting visibility to revenue is the separate half of the problem.

Common Mistakes in a First Audit

Six errors account for most audits that produce numbers nobody can use.

  1. Sampling once. A single run per prompt makes every subsequent comparison noise. This is the error that invalidates the most work, and it is entirely avoidable.
  2. Using a signed-in session with history. Assistants personalise. If you have spent six months asking an assistant about your own company, it will name you more readily than it names you for a stranger. Fresh sessions, or a private window, are non-negotiable.
  3. Writing marketing prompts. "What is the best AI-native revenue attribution platform?" contains your own positioning language. A buyer asks "how do I find out if ChatGPT is sending us business?" Prompts that presuppose your category framing measure your framing, not your visibility.
  4. Averaging across tiers. Brand prompts will always score higher than category prompts. A blended number hides the only interesting finding — that you are known when named and invisible before that.
  5. Ignoring the sources column. Teams record mentions diligently and skip the cited domains, discarding the most actionable output of the exercise.
  6. Auditing after making changes. A baseline measured after a content refresh is not a baseline. Measure first, even if the site is in a state you are not proud of.

Turning the Audit Into a Plan

The audit produces four artefacts. Each maps to a workstream, and the mapping should be explicit before anyone starts writing.

  • The accuracy backlog → a canonical facts page on your own site, plus corrections to the specific third-party records that carry the wrong information. Fastest to fix, and usually the highest-severity finding.
  • The retrievability gaps → crawler permissions, server-side rendering, indexation. A two-day technical ticket, not a quarter of work.
  • The question-ownership map → for each category prompt where you were absent, which page should have answered it? Where no page exists, that is the content brief. Where two pages half-exist, that is a consolidation.
  • The off-site target list → the ranked cited domains, with the mechanism of inclusion noted against each. This becomes a quarterly programme, and it is the one that keeps producing after the on-site work is finished.

Sequence them in that order. Accuracy first because it is cheap and the damage is ongoing; retrievability second because nothing downstream works without it; content third; off-site continuously, starting immediately, because it has the longest lead time.

Reporting the Audit Upward

Executives do not want a spreadsheet of prompts. Three slides work:

One: share of answer against the three competitors the models actually name, with the sample size. This is the competitive position and it is usually the slide that creates urgency.

Two: the accuracy finding, quoted verbatim. A screenshot of a model describing the product incorrectly does more to fund the programme than any rate.

Three: the off-site source list, framed as "here is where buyers are being told about our category, and here is where we are absent." It converts an abstract problem into a procurement-shaped one with named targets and a plausible cost.

Deliberately omit any composite score. Scores invite the question "what should it be?", which has no answer, and they obscure the four diagnoses that actually drive the plan. Ongoing brand monitoring is the natural follow-on once the baseline exists.

Frequently Asked Questions

What is an AI visibility audit?

An AI visibility audit is a structured measurement of whether AI engines mention and cite your brand across a fixed prompt set, covering presence, citation, accuracy, competitive context and the third-party sources being drawn on.

How do I check my AI visibility?

Define thirty to fifty buying questions, run each several times in fresh sessions across ChatGPT, Perplexity and Google's AI surfaces, and record whether you were mentioned, whether you were linked, what was said, which competitors appeared and which sources were cited.

How long does an AI visibility audit take?

A manual first audit across thirty prompts and three engines takes roughly four to six hours including analysis. Repeat runs are faster because the prompt set and spreadsheet already exist.

Do I need a tool to audit AI visibility?

No. A spreadsheet is sufficient for a first baseline and often better, because it forces you to read the answers. Tooling pays off when you need multiple locales, weekly cadence or year-long trends.

What is a good AI visibility score?

There is no absolute benchmark, and cross-vendor scores are not comparable. Judge against your own baseline and against the competitors named in the same answers — share of answer is the meaningful comparative figure.

How often should you repeat an AI visibility audit?

A full audit quarterly, with brand-tier prompts monitored weekly. Keep the prompt set frozen between runs so the comparison remains valid.

References

  1. https://platform.openai.com/docs/bots
  2. https://docs.perplexity.ai/guides/bots
  3. https://developers.google.com/search/docs/crawling-indexing/google-common-crawlers
  4. https://developers.google.com/search/docs/appearance/structured-data/organization

Related Articles

AI Marketing

AI Brand Monitoring: Watching What Models Say

AI Marketing

What Is AI Visibility? How to Define and Measure It

Analytics

AI Visibility Tracking: Metrics, Cadence and Prompt Design

Your Free AI Referral Report

Is AI referring you or your competitor?

AI is becoming your market's biggest referral source. Your report shows where those referrals are going, and what winning them is worth.

What you'll get

  • Where AI sends buyers in your market
  • Who's capturing them today
  • Your AI Search Revenue Gap
Book an AI Revenue ForecastLog in

Built for your market, walked through with you on a 10-minute call.

MultiplierAI

We engineer the system that produces your revenue. Measurable, attributable, and compounding.

Book an AI Revenue Forecast
Product
  • The Revenue Brain
  • The Revenue Engine
  • The Intelligence Layer
  • The Revenue Chain
Company
  • Free AI Revenue Forecast
  • Contact
  • Privacy
  • Terms
© 2026 MultiplierAI·Revenue Growth Engine
All systems operational