What This Paper Is

In November 2023, a research team from Princeton University, Georgia Tech, and other institutions published GEO: Generative Engine Optimization (arXiv:2311.09735) on arXiv. This was the first systematic academic definition of generative engine optimization as a concept, accompanied by quantitative research.

The paper's core contribution: establishing an experimental framework for observing how content strategies affect AI citation rates, and identifying multiple content strategies that significantly improve AI visibility under experimental conditions.

Research Methodology

The research team constructed a test set of queries spanning multiple verticals (law, finance, health, etc.) and tested using multiple generative search engines. They systematically modified article content by adding different optimization strategies, then measured changes in brand or article exposure in AI answers.

An important constraint: experiments were conducted under fixed-context conditions, meaning articles were already retrieved—the test only measured content strategy effects on citation rates within retrieved results, not simulating the full competitive scenario of earning visibility from scratch.

Key Findings: Which Content Strategies Work

The strategies with the most significant effects: verifiable statistics—including specific data, percentages, and checkable research findings significantly increases AI citation willingness, as AI models prefer citing content with precise numbers because such content is easier to verify. Primary sources—directly citing or linking to authoritative primary sources raises content credibility scores. Attributed quotations—including direct quotations attributed to specific individuals (name plus title or institution), as AI prefers citing perspectives with clear attribution. Content fluency—naturally flowing language with clear paragraphs outperforms dense keyword-stuffed content in AI answer inclusion rates.

Applying These Findings in Practice

When translating this paper's findings into operational methods: first, experimental findings hold under specific domain conditions and already-retrieved scenarios, and actual effects vary by target market, competitive landscape, and model version, so do not treat the specific uplift percentages from the paper as guarantees. Second, the essence of these strategies is improving content quality—making content more data-backed, better sourced, and more genuinely opinionated, which also benefits SEO.

Practical suggestions: add industry data (with sources) to product pages and blog posts; cite internal research or authoritative reports in FAQ content; include named user quotations in About and case study pages; ensure natural language and clear paragraph hierarchy throughout the site.

Limitations: What This Paper Cannot Tell You

Honestly understanding this research boundaries is important: the experimental environment differs from real-world scenarios, since in actual AI search, content first needs to be retrieved before content strategies can have any effect. AI models are continuously updated, so the 2023 experimental results were valid at that time, but specific effects will vary as model versions, training data, and retrieval systems evolve—which is why continuous observation is more valuable than a single analysis. Effect sizes differ significantly across domains, meaning there is no fixed uplift percentage that applies universally.

How Broccoli AI GEO Uses This Research

The Broccoli AI GEO platform uses this paper to inform several audit dimensions: the GEO Audit module checks whether your content includes verifiable statistics, whether there are primary source citations, and whether page entities are clearly defined. These audit items directly correspond to the effective strategies identified in the paper.

At the same time, the platform preserves the research usage boundaries: we display analysis results and optimization recommendations based on the current sample, not promises of specific uplift. Every report includes sample size, sampling timestamp, and confidence notes so you can return to the evidence itself for judgment.