A clear-eyed breakdown of the CMU/Georgia Tech GEO paper's nine optimization strategies—what the data supports, what remains unproven, and what practitioners should actually take away.
- The foundational GEO paper tested nine content optimization methods across 10,000 queries and multiple AI engines (Aggarwal et al., 2024)
- Three strategies showed consistent gains: citing sources, adding quotations from authorities, and including relevant statistics
- Results varied significantly by query type—informational queries responded differently than transactional ones
- The study used a controlled benchmark, not live production environments—real-world replicability remains an open question
- No strategy produced universal improvement across all engines and query categories
The Paper That Started the GEO Conversation
In early 2024, researchers from Carnegie Mellon University and Georgia Tech published what became the field-defining paper on Generative Engine Optimization. Titled "GEO: Generative Engine Optimization," the study by Aggarwal et al. introduced both the term and a systematic framework for testing how content modifications affect visibility in AI-generated answers.
The paper matters because it moved GEO from speculation to measurement. But it also has limits that practitioners should understand before treating its findings as universal laws.
What the Researchers Actually Tested
The study evaluated nine distinct optimization strategies across a benchmark of 10,000 queries. These queries were drawn from diverse domains and tested against multiple generative engines including BingChat (now Copilot) and perplexity.ai.
The researchers measured impact using a "position-adjusted word count" metric—essentially tracking how much of the AI's response drew from each source and weighting earlier mentions more heavily (Aggarwal et al., 2024).
The Strategies That Worked
Three methods produced the most consistent visibility improvements across the study:
| Strategy | Relative Improvement | Best-Performing Domain |
|---|---|---|
| Cite Sources | Up to 40% | Factual/informational queries |
| Quotations | Up to 30% | Opinion and debate topics |
| Statistics | Up to 30% | Data-driven questions |
(Aggarwal et al., 2024)
The researchers found that adding citations—explicit references to sources within the content—showed the strongest gains for factual queries. This aligns with how retrieval-augmented generation systems work: they prefer content that itself demonstrates verification.
Quotations from recognized authorities performed well for subjective topics where credibility signals matter more than raw data.
Statistics helped most when queries involved comparative or quantitative elements.
What the Paper Did Not Prove
Here's where intellectual honesty matters. The study has real limitations:
Benchmark vs. production environment. The GEO-bench dataset is synthetic. Real-world queries, index freshness, and model updates all introduce variables the study couldn't control.
Point-in-time snapshot. Models change constantly. The paper tested systems as they existed in late 2023. Today's GPT-4o, Claude 3.5, and Perplexity use different retrieval architectures.
Relative improvement, not absolute visibility. A 40% improvement on a baseline of near-zero is still near-zero. The paper measured relative gains, not whether content reached the top of results in competitive categories.
No traffic correlation. The study measured citation and word-count inclusion. It did not measure click-through, referral traffic, or conversion. Visibility in an AI answer does not automatically equal business impact.
Single-pass optimization. Each strategy was tested in isolation. The paper did not fully explore how combining strategies affects results—or whether combinations create diminishing returns.
What's Replicable for Practitioners
Despite the caveats, several principles hold up:
Credibility signals matter. If you want AI engines to trust and cite your content, demonstrate trustworthiness through sourcing.
Domain context shapes strategy. Technical documentation may benefit more from authoritative tone; health content from citations; comparison articles from statistics.
Keyword stuffing doesn't work. The study found this traditional SEO tactic produced minimal improvement and sometimes negative effects in generative engines (Aggarwal et al., 2024).
The Honest Takeaway
The GEO paper established that content optimization for AI engines is measurable and improvable. That's valuable. But it didn't create a universal playbook.
What practitioners should do now: treat the paper's strategies as hypotheses for your own testing. Monitor whether your content appears in AI responses. Track which competitors show up for the same queries. Iterate based on observed data, not assumed lifts.
This is where systematic measurement becomes essential. Understanding whether your optimization efforts actually affect AI visibility requires consistent tracking across engines, queries, and time.
Mentio provides that measurement layer—tracking whether and how often your brand appears in ChatGPT, Claude, and Perplexity responses. No invented metrics. Just visibility data you can act on.
Frequently asked questions
Does the GEO paper guarantee that following its strategies will improve my AI visibility?
No. The paper demonstrated measurable effects under controlled conditions, but results varied by query type, engine, and domain. Treat the strategies as informed starting points, not guarantees. Real improvement requires testing against your specific content and competitive landscape.
Is the GEO paper's methodology still relevant given how quickly AI models change?
The principles—credibility signals, source attribution, domain-appropriate content—remain sound. However, the specific percentage improvements were measured on 2023-era systems. Ongoing monitoring is necessary to understand current impact, since retrieval and ranking mechanisms evolve with each model update.
Should I combine multiple GEO strategies or focus on one at a time?
The original study primarily tested strategies in isolation, so there's limited evidence on optimal combinations. A reasonable approach: implement the highest-performing strategies first (citations, statistics, quotations where appropriate), then measure the actual effect on your AI visibility before adding additional optimizations.
See how AI engines answer for your brand.
Mentio tracks whether ChatGPT, Claude and Perplexity mention and cite you — own your data, self-host anytime.
Start tracking