In Brief
- Core Answer: AI citations are the source links attached to a generated answer. Systems attach them to the passages that supported specific claims, which means citation is won at the passage level by content that makes a checkable, attributable statement.
- Why It Matters: A citation is the only part of a generated answer that sends a reader anywhere. Being mentioned without one is exposure; being cited is exposure plus a route.
- Best For: Content and SEO teams trying to work out why competitors get linked in AI answers and they do not.
AI citations are the attributions a generative system attaches to claims in its answer — links back to the sources that supported them. Citations are granted at the passage level, not the page level, and they follow a consistent logic: the system links the source it actually used for that specific sentence.
Understanding that logic explains most of the otherwise puzzling patterns, including why a page ranking eighth gets cited while the page ranking first does not.
How AI Citations Are Selected
The pipeline is broadly consistent across engines:
- The question is decomposed into sub-queries.
- Documents are retrieved for each.
- Passages are selected that support the claims the answer will make.
- The answer is composed, often merging several sources into one sentence.
- Links are attached to the sentences their sources supported.
The consequence: a citation is a record of which passage did the work. It is not a reward for domain authority, and it is not proportional to ranking. It goes to whichever text most cleanly supported the claim being made.
The Four Properties of Citable Content
1. It Makes a Checkable Claim
A citation attaches to an assertion. Content that asserts nothing specific supports nothing specific. "Many organisations struggle with attribution" is not a claim; "assistant referrals frequently arrive without a referrer header, so analytics records them as direct traffic" is. The second can be cited because it is the kind of statement an answer needs backing for.
The practical test: could a sceptical reader disagree with this sentence? If not, it is filler and it will not be cited.
2. It Stands Alone
The selected passage is extracted from its context. A paragraph that depends on the preceding two to make sense cannot be lifted intact, so the system either reconstructs the point from a competitor's cleaner passage or paraphrases yours without attribution.
Self-contained means: the subject is named, not implied by a pronoun; the claim is complete within the paragraph; and the paragraph makes sense with the heading removed.
3. It Is Attributable
Content that identifies who is making the claim, when, and on what basis is easier to cite than anonymous prose. A named author, a visible date, a stated methodology and links to underlying sources all raise the confidence a system can place in the passage. This is not an E-E-A-T incantation; it is the practical difference between a statement a system can stand behind and one it cannot.
4. It Is Not Already Better Covered Elsewhere
Citation is competitive. If your page restates what a widely-cited source already says, there is no reason to cite yours. The content most likely to earn citations is the content that supplies something absent from the corpus: original data, a specific method, a documented failure mode, a comparison nobody else has made, a number nobody else has published.
This is the uncomfortable one, because it means content strategy in an answer-engine world is closer to publishing than to marketing.
Mentioned But Not Cited
The most common diagnostic pattern, and it has three distinct causes worth separating.
Cause | Signal | Fix |
|---|---|---|
Not retrievable | Named, but the cited sources are all third-party | Check crawler access and rendering |
Not extractable | Your page ranks but is never linked | |
Known from training, not retrieval | Mentioned with no source at all | Publish canonical facts; improve retrievability |
The third case is worth dwelling on. When a model names you without citing anything, the information came from its training rather than a live fetch. That is flattering — you are known — and fragile, because it is frozen at the training cutoff and cannot be corrected until the model is replaced.
Where Citations Actually Come From in Your Category
Run twenty category questions across three engines and tabulate every domain cited. In most B2B categories the result is concentrated: four to eight domains account for the majority of citations, and they are rarely the ones a communications plan would have chosen.
Typical composition:
- One or two review or comparison platforms.
- One or two industry publications with strong topical depth.
- A practitioner community — a forum, a subreddit, a Q&A site.
- Documentation or reference sites, when the question is technical.
- Occasionally a single vendor blog that has become the de facto explainer for the category.
That last one is the opportunity. Categories in formation frequently have no canonical explainer, and the vendor that publishes the clearest, most complete, most attributable version of "what this category is and how to evaluate it" often becomes the cited source for years. It is available to anyone willing to write it properly.
What Does Not Earn Citations
- Domain authority alone. Helps with retrieval, does not determine which passage supports a claim.
- Volume of content. More pages means more weak candidates unless each makes a distinct, checkable claim.
- Keyword optimisation. Semantic retrieval does not need density and does not reward it.
- Asking to be cited. There is no markup for it and no consumer that would honour one.
- Restating consensus. If the claim is already well-sourced elsewhere, your version adds nothing to cite.
Measuring Citation Performance
Track four things across a fixed prompt set, sampled repeatedly:
- Citation rate — runs where a link to your domain appeared, as a proportion of total runs.
- Mention-to-citation ratio — the gap between being named and being linked. A wide gap is a retrievability or extractability problem.
- Which pages get cited — usually a small subset, and usually not the pages you expected. Study what those pages do differently.
- Citation position — first-cited sources carry disproportionate attention.
The third of these is the most useful and the most neglected. Whichever three pages earn most of your citations are a working template for the rest of the site, derived from evidence rather than theory. A structured audit produces this list as a byproduct.
The Original-Contribution Test
If citation goes to whatever supplied something the corpus did not already have, the planning question becomes concrete: what can you publish that nobody else can?
Five categories of content pass this test reliably, and almost nothing else does.
- First-party data. Aggregate figures from your own product or customer base, published with a stated methodology and sample size. Nobody else has your data, and a specific number with a source is the most citable object on the web.
- Documented methods. A step-by-step process with real parameters — thresholds, sequences, decision rules. Generic advice is everywhere; a specified method is not.
- Failure modes. What goes wrong, under what conditions, and how to recognise it. Vendor content systematically under-covers this, which leaves the space open.
- Comparisons nobody has made. Structured, fair, table-based comparisons of approaches or categories. These are heavily retrieved because comparison questions are common and good comparison content is rare.
- Definitions in forming categories. Where a term is new and no canonical explainer exists, the clearest version becomes the cited one, sometimes for years.
What fails the test: rewritten versions of what three better-established sources already say, list posts assembled from other list posts, and anything whose distinguishing feature is length.
Reverse-Engineering a Competitor's Citations
When a competitor is consistently cited and you are not, the diagnosis is usually available in twenty minutes.
- Open the cited page. Not their homepage — the specific URL the answer linked to.
- Find the paragraph that supports the claim. It is normally the opening paragraph under a heading, forty to sixty words, declarative.
- Check what it contains. A number, a named source, a specific mechanism, or a clean definition — usually one of the four.
- Find your equivalent page. Compare the corresponding paragraph. In most cases yours either buries the claim, hedges it, or omits the specific detail that made theirs quotable.
- Rewrite that paragraph. Not the page. The paragraph.
Done across ten priority questions this is a few hours of work and produces a more accurate picture of what earns citations in your category than any general guidance, including this article. The evidence is sitting in the answers you are already losing.
Citation behavior differs by engine; how to get cited by Claude covers the engine that sends the most B2B referrals per visit.
Frequently Asked Questions
What are AI citations?
AI citations are the source links a generative system attaches to claims in its answer, indicating which page supported that specific statement. They are granted at the passage level rather than the page level.
How do you get cited by AI?
By being retrievable, and by writing self-contained passages that make specific, checkable, attributable claims which an answer needs backing for. Content restating well-covered consensus rarely earns a citation because a better-established source already exists.
Why is my brand mentioned but not cited?
Three common causes: your pages are not reachable by that engine's crawler, your passages are not extractable because the answer is buried, or the model knows you from training rather than live retrieval and therefore has no source to link.
Does domain authority affect AI citations?
It influences whether you are retrieved, not which passage is selected to support a claim. A lower-authority page with a cleaner, more specific answer is frequently cited over a higher-authority page without one.
Can you ask an AI to cite your site?
No. There is no markup or directive that requests attribution, and no major system honours one. Citation follows from being the source that supported the claim.
How many sources does a typical AI answer cite?
Usually a small handful, varying by engine and question complexity. Because the number is small and there is no second page, absence from the cited set is a complete exclusion rather than a low ranking.