A close look at the preprint research on whether traditional domain authority predicts who gets cited in AI-generated answers—and what SEO professionals should take from it.
- Princeton NLP researchers found that LLMs disproportionately cite .gov and .edu domains, but not strictly because of PageRank-style authority signals.
- The study suggests source selection in retrieval-augmented generation (RAG) systems follows different patterns than traditional search ranking.
- High-authority domains correlate with citation frequency, but causation remains unproven—content type and retrieval index composition matter.
- Traditional link-based metrics (Domain Authority, PageRank) are weak predictors of LLM citation compared to content clarity and factual density.
- For GEO practitioners, the implication is clear: optimising for AI visibility requires new measurement, not just DA building.
The study in context
In late 2023, researchers at Princeton NLP released a preprint examining how large language models select and cite sources when generating answers. The paper, "Attributed Question Answering: Evaluation and Modeling for Attributed Large Language Models" (Bohnet et al., 2023), analysed citation behaviour across multiple LLMs integrated with retrieval systems.
The findings challenged a common assumption: that domains with high traditional authority (measured by metrics like Moz Domain Authority or Ahrefs Domain Rating) would automatically dominate AI-generated citations.
This matters for anyone tracking brand visibility in AI search. If the old playbook—build links, raise DA, rank higher—doesn't transfer cleanly to LLM outputs, then GEO requires different tactics and different measurement.
What the researchers actually found
The Princeton team evaluated citation patterns in retrieval-augmented LLMs by analysing which sources appeared in model outputs when answering factual questions.
Key observations from the preprint:
| Finding | Implication |
|---|---|
| .gov and .edu domains were cited 2–3× more often than their index share would predict | Institutional sources receive disproportionate weight |
| Wikipedia appeared in roughly 15–20% of attributed answers | Encyclopaedic, neutral-tone content performs well |
| Commercial domains (.com) were cited less frequently than their prevalence in training data | Brand content faces a citation disadvantage |
| Citation frequency did not correlate strongly with Moz DA or Ahrefs DR | Traditional link metrics are poor predictors |
(Source: Bohnet et al., "Attributed Question Answering," Princeton NLP, 2023 preprint)
The researchers noted that retrieval systems favour "authoritative-seeming" content—but the signals they use are not identical to PageRank.
Why .gov and .edu domains over-index
The preprint offered several hypotheses for the institutional domain preference:
-
Training data composition. LLMs trained on Common Crawl and curated datasets over-sample government and academic sources relative to the open web.
-
Retrieval index bias. RAG systems often use filtered indexes (like Wikipedia or curated knowledge bases) that skew toward institutional content.
-
Content structure. Government and academic pages tend to be factually dense, clearly structured, and free of commercial intent—qualities that retrieval models reward.
-
Explicit domain heuristics. Some retrieval pipelines apply domain-level filters or boosts for .gov/.edu as a proxy for trustworthiness.
None of these factors are captured by traditional link-based authority scores.
What this means for GEO practitioners
The Princeton findings suggest a shift in how we should think about AI visibility:
Domain Authority is not a citation predictor. A DR 80 commercial site may still be invisible in ChatGPT answers if its content reads as promotional or lacks factual specificity.
Content quality signals matter more. Clarity, structure, and informational density appear to influence retrieval ranking. This aligns with earlier observations from Perplexity and Bing Chat behaviour.
Institutional sources have structural advantages. If your competitors include government agencies or universities, you may be fighting an uphill battle for share of voice in certain query categories.
Measurement must be AI-native. Tracking your Ahrefs DR tells you nothing about whether Claude or ChatGPT mentions your brand. You need direct observation of AI outputs.
The limits of the research
The preprint has constraints worth noting:
- It examined a snapshot in time. Retrieval systems and model behaviour evolve.
- Sample size was limited to factual question-answering. Brand queries and commercial intent were underexplored.
- The study predates GPT-4 Turbo and Claude 3. Citation behaviour may differ in newer models.
Still, the directional insight holds: traditional authority metrics are insufficient for predicting AI visibility.
Tracking what actually matters
If DA doesn't predict citations, what should you measure?
The answer is direct observation. Query the AI engines with the prompts your customers use. Record whether your brand appears, in what context, and which competitors are cited alongside you.
This is the problem Mentio is built to solve. It probes ChatGPT, Claude, and Perplexity with your target queries, tracks brand mentions and citations over time, and surfaces share-of-voice data across engines. No guesswork. No proxy metrics.
For teams serious about GEO, the Princeton study is a useful reminder: the old scorecards don't apply. Visibility intelligence requires new tools.
Frequently asked questions
Does high Domain Authority help with AI citations at all?
Indirectly. High-DA sites often have content that is well-structured and factually dense—qualities that retrieval systems reward. But the DA score itself is not a direct input to LLM citation ranking. The Princeton preprint found weak correlation between link-based authority metrics and actual citation frequency.
Should I focus on getting .gov or .edu backlinks for GEO?
The research does not suggest that backlinks from institutional domains improve AI citation rates. The .gov/.edu advantage appears to stem from content characteristics and retrieval index composition, not link equity transfer. Earning mentions on authoritative sources may help indirectly, but link building alone is not a GEO strategy.
How do I know if my brand is being cited by AI assistants?
You need to query the AI engines directly and systematically. Manual spot-checks are unreliable because model outputs vary by session, region, and prompt phrasing. Tools like Mentio automate this process—probing multiple engines with your target queries and tracking citation presence over time.
See how AI engines answer for your brand.
Mentio tracks whether ChatGPT, Claude and Perplexity mention and cite you — own your data, self-host anytime.
Start tracking