Generative Engine Optimization: A Critical Look at the GEO Study

A digital representation of the intersection of AI and search

AI optimization is currently in its earliest stages, with limited empirical research available. However, a recent study titled GEO: Generative Engine Optimization is one of the first attempts to move the discussion beyond speculation and test how changes to website content may affect its visibility within AI-generated answers.

The researchers found that adding quotations, statistics, and citations could substantially increase how prominently a source appeared within a generated response. They introduced the term Generative Engine Optimization, or GEO, to describe a framework for improving visibility within systems that retrieve information from multiple sources and synthesize it into a direct answer.

In my previous article, Generative Engine Optimization: What the GEO Study Found, I examined the study’s methods, strongest-performing treatments, and potential applications. This follow-up takes a more critical look at the research. The central question I examine here is what those results actually tell us about visibility in generative search.

The GEO study does not evaluate the complete path from publishing a webpage to appearing in an AI-generated answer. It begins with a fixed set of sources that have already been retrieved and measures how changes to one source affect its representation during answer generation. That makes the paper highly relevant to synthesis, but less conclusive about discovery, retrieval, citation eligibility, user behavior, and business outcomes.


Why the GEO Study Deserves Attention

The GEO paper addresses a problem that traditional search metrics were not designed to measure. In a typical search engine, visibility can usually be approximated through ranking position, impressions, clicks, and traffic. A generative response is structured differently. It may combine information from several sources, attribute multiple sentences to one citation, mention another source briefly, and position those references at different points in the answer.

The researchers recognized that a source appearing in a generated response cannot be evaluated through ranking position alone. They therefore created metrics intended to measure how much of the response relied on a source, where the source’s information appeared, and how influential the citation seemed within the answer.

That is an important contribution. Generative search introduces a form of visibility that sits between retrieval and user engagement. A source may not receive a traditional ranking, yet its information may shape a substantial portion of the final response. The GEO paper provides one of the first frameworks for studying that form of influence.

However, the paper’s headline finding, that some methods improved visibility by as much as 40%, can easily be interpreted more broadly than the methodology supports. The study does not show that adding citations or statistics will increase traffic by 40%, make a page 40% more likely to be retrieved, or produce a 40% increase in citations across a commercial AI platform. It measures a narrower outcome: the optimized source’s relative prominence within answers generated from a predefined group of documents.

The Study Measures Influence After Retrieval

The most important limitation is built into the experimental design. For each query, the researchers collected the top five Google search results and supplied the cleaned content from those pages to the generative system. One source was then selected for modification, and the researchers measured whether the modified version received more representation in the generated answer.

Generative visibility pipeline showing discovery, retrieval, source selection, extraction, synthesis, attribution, and user response, with the GEO study focused on the middle stages.
The GEO study primarily evaluates what happens after a source has already been retrieved. A complete model of generative visibility also includes discovery, selection, attribution, and user response.

The optimized source did not need to compete against the wider web after it was changed. It had already been discovered, indexed, ranked, and included in the source set.

This means the experiment begins after several important stages have already occurred. Before a source can influence an AI-generated answer, the system must generally be able to access the page, interpret its content, determine that it is relevant, select it from a larger body of possible sources, and make its information available to the generative model. Only then can the model extract information, combine it with other sources, and generate a response.

The GEO study focuses primarily on those later stages. It measures how an already retrieved source competes for representation during synthesis.

This does not make the findings less useful. It makes their scope clearer. Adding a statistic may make a page more useful once it has entered the retrieval set. The experiment does not show that the statistic caused the page to enter that set.

The experiment modified a source that was already included in a fixed five-document retrieval set. Real-world systems select sources from a much larger and changing corpus.

A complete model of generative visibility would need to account for the entire sequence:

Discovery and access → relevance and retrieval → source selection → extraction → synthesis → attribution → user response

The GEO paper examines the middle of that process. Its findings should be interpreted accordingly.

GEO Is Not a Complete Visibility Framework

The researchers position GEO as a new optimization paradigm for generative engines. That framing is understandable because generated answers create problems that traditional rank-based metrics cannot address. But, the study does not provide a complete framework for earning visibility across generative search. It provides a framework for improving source representation after retrieval.

That difference becomes especially important when GEO is discussed as an alternative to SEO. A source cannot influence a generated answer if the system never discovers or retrieves it. Technical accessibility, indexability, content relevance, internal linking, structured information, authority, reputation, and external references may all affect whether the source becomes available for synthesis.

The paper’s own methodology demonstrates this relationship. Google’s top five results formed the retrieval set for the primary experiment. Conventional search visibility was therefore a prerequisite for inclusion.

A more useful interpretation is that SEO and GEO operate at different but connected layers. SEO may help a source become discoverable, understandable, relevant, and retrievable. GEO may affect how that source is used once retrieved.

The boundaries are unlikely to remain perfectly separate. Changes that improve clarity, structure, and evidence may affect both retrieval and synthesis. Even so, treating the two disciplines as complementary provides a more accurate model than presenting GEO as a replacement for SEO.

The Study Takes a Narrow View of SEO

The paper’s treatment of keyword stuffing is one example of how its representation of traditional SEO becomes too narrow. The researchers define keyword stuffing as modifying content to include more keywords from the query, describing this as something expected in classical SEO optimization. The treatment performed poorly on the primary visibility metric.

The result is useful, but the treatment is not a meaningful test of SEO as a discipline. It is better understood as query-term insertion. Modern SEO includes far more than increasing keyword frequency. It involves understanding search intent, organizing information, strengthening internal relationships, improving crawlability, clarifying entities, addressing related questions, and building content that satisfies the underlying need represented by the query.

The study did not test those practices. It tested whether inserting additional query terms into an already retrieved document increased that document’s influence during answer generation. The study results showed that it did not.

That finding supports a limited conclusion: mechanical repetition did not make the source more useful to the generative model. It does not demonstrate that keyword research, relevance, topical coverage, or on-page SEO are ineffective in generative search.

This is important because the paper occasionally uses a weak proxy for traditional SEO and then draws a broader contrast between SEO and GEO. The experiment provides evidence against query-term insertion, not against the wider set of practices used to improve search visibility.

The 40% Gain Is a Share-of-Answer Measurement

The widely cited 40% result also requires more context. The researchers normalized their impression metrics so the total visibility assigned across citations in a response equaled one. This allowed them to compare the relative contribution of each source which also made the measurement competitive.

When one source gained a larger share of the answer, another source could lose share. The optimized source did not necessarily create more total exposure for all publishers. It captured a greater portion of the visibility available within the generated response. This is closer to share of answer than traditional search impressions.

That changes the interpretation of the result. A 40% improvement does not necessarily mean the answer became longer, the source received 40% more citations across the platform, or users became 40% more likely to visit the website. It means the optimized source received a larger proportion of the response under the study’s metrics.

Share of answer may become an important measurement in generative search. Publishers may eventually need to track how much of an answer relies on their information, where they are cited, how their contribution compares with competitors, and whether their brand or entity is represented accurately.

However, share of answer is only one part of visibility. A complete framework would also need to measure inclusion frequency, citation frequency, citation prominence, brand portrayal, source sentiment, click-through behavior, user recall, and downstream business impact.

Lower-Ranked Sources Gained More, but the Meaning Is Limited

One of the study’s most interesting findings was that lower-ranked sources often received larger relative gains than sources occupying the first position. For example, the Cite Sources method increased the fifth-ranked source’s visibility by 115.1%, while the first-ranked source declined by 30.3%. Quotation Addition and Statistics Addition produced similar patterns, with much larger gains for fifth-position sources than for those already ranking first or second.

The researchers suggest that GEO could therefore benefit smaller creators and independent websites. That is possible, but the experiment did not measure business or website size.

A source ranking fifth is not necessarily a small publisher. A source ranking first is not necessarily a large corporation. The study did not classify sources by revenue, domain authority, backlinks, brand recognition, publishing resources, or organizational scale.

The evidence supports a narrower conclusion. Lower-position sources that had already entered the retrieval set sometimes gained more from the tested modifications than higher-position sources.

This may indicate that a source can compensate for a weaker initial position by contributing information that is particularly useful during synthesis. It does not establish that small websites have an inherent advantage in generative search.

The normalized visibility metric may also contribute to this pattern. Because the sources compete for a finite share of the answer, increasing the representation of a lower-position source can reduce the relative share assigned to a higher-position source.

The finding remains valuable, but it should be framed as evidence about competition within a fixed source set rather than as proof that generative engines level the playing field across the web.

The Study Uses Modeled Attention, Not Human Attention

The study’s Subjective Impression metric attempts to capture qualities that word count and citation position cannot fully represent. These include relevance, influence, uniqueness, diversity, citation prominence, and the likelihood that a user will follow the citation.

The limitation is that these judgments were made by GPT-3.5 rather than by users.

The model evaluated the generated responses and estimated how influential or clickable each citation appeared. The resulting scores were then normalized to make them comparable with the position-adjusted metric. This is modeled attention, not observed attention.

The difference is substantial. A language model’s estimate that a citation is likely to receive a click does not tell us whether users noticed it, trusted the source, remembered the publisher, or selected the link. No user behavior was observed in the study.

There is also a potential shared-model bias. GPT-3.5 was used to generate the experimental answers and to evaluate the Subjective Impression metric. The evaluator may prefer styles or structures that resemble those favored by the generator.

That does not make the metric useless. LLM-based evaluation can provide a scalable way to compare large numbers of responses. However, the metric should be treated as an approximation that requires validation against human behavior.

Future research should test whether source prominence affects citation recognition, trust, click-through rates, follow-up questions, brand recall, and decision-making. Without that validation, we cannot assume that a source receiving more modeled visibility will receive more meaningful attention.

The Evidence Added by GEO Requires Verification

The strongest-performing treatments included the addition of statistics, quotations, and citations. These methods are often described as simple content improvements, but they involve more than style.

Adding a statistic introduces a factual claim. Adding a quotation introduces language attributed to a person or organization. Adding a citation establishes a relationship between a claim and a supporting source. Each of those requires verification.

The study used language models to modify source content. The paper measured whether those modifications increased representation, but it paid less attention to whether every added claim, source, or quotation was accurate, current, and appropriately contextualized.

That creates a practical risk. A generated statistic may increase specificity while introducing unsupported information. A plausible citation may create the appearance of credibility even when the cited source does not support the claim. A quotation may be altered, misattributed, or removed from its original context.

If those additions increase the likelihood that a source influences a generated answer, they may also increase the likelihood that inaccurate information is amplified.

This is where optimization and editorial responsibility become inseparable. Publishers should not interpret the study as permission to add evidence-like elements without validation. Every statistic, quotation, and citation should be checked against the original source before publication.

An optimization that increases visibility while reducing accuracy is not a successful optimization. It is a more efficient way to spread misinformation.

The Findings May Reflect Information Utility

The paper describes citations, quotations, and statistics as improving the credibility and richness of content. That is a reasonable interpretation, but credibility was not isolated as the mechanism responsible for the gains. Another possibility is that these methods increased information utility.

A statistic gives the model a concise and measurable claim. A quotation provides distinctive wording connected to an identifiable speaker or source. A citation creates an explicit relationship between a statement and evidence. Fluency may make those relationships easier to interpret.

Each treatment may improve the source’s usefulness during answer construction. The content becomes easier to extract, summarize, differentiate, connect to the query, and attribute.

This interpretation explains why the strongest methods may work together. A statistic becomes more useful when it is clearly written. A quotation becomes more useful when its relevance is explained. A citation becomes more useful when the claim it supports is specific and understandable.

The study’s combination experiment supports this possibility. Fluency Optimization combined with Statistics Addition produced the strongest result in the smaller subset analysis, suggesting that presentation and evidence may reinforce each other.

The practical implication is more substantial than “add more quotes and statistics.” Content should contribute information that is specific, distinctive, verifiable, and easy to incorporate into an answer.

The Benchmark Is Broad, but Primarily Informational

GEO-bench contains 10,000 queries drawn from multiple sources, including anonymized search queries, question-answering datasets, debate prompts, Perplexity Discover, ELI5, and GPT-4-generated queries. Approximately 80% of the benchmark consists of informational queries. Transactional and navigational queries each account for about 10%.

This matters for commercial applications. Businesses are often most interested in queries involving comparison, recommendation, suitability, location, availability, pricing, reputation, and choice. These may depend on different source characteristics from general informational questions.

A query asking how a process works may respond to clear explanations, quotations, and statistics. A query asking which contractor, product, software platform, or healthcare provider to choose may rely more heavily on reviews, service areas, availability, pricing, expertise, reputation, and third-party corroboration.

The study includes some transactional and navigational queries, but its strongest overall conclusions are shaped primarily by informational tasks. More research is needed before the findings can be generalized to product recommendations, local search, provider selection, ecommerce, and other decision-oriented environments.

The Category Findings Are Difficult to Apply

The researchers found that different treatments performed better for different query tags. Quotations performed well for people and society, explanations, and history. Statistics performed well for law and government, debate, and opinion. Citations performed well for statements, facts, and law and government. These findings suggest that GEO should be context-dependent rather than applied as a universal checklist.

The category structure is difficult to operationalize because it combines several dimensions. “Business” and “Health” describe subject areas, while “Facts,” “Explanation,” and “Opinion” describe the form of the query or expected response.

A business query may also be factual, comparative, local, transactional, or explanatory. These categories are not mutually exclusive.

A more useful taxonomy would separate subject, intent, task type, answer format, complexity, geographic dependence, freshness, and commercial relevance. That would make it easier to determine which treatment is appropriate for a specific page or query.

The current findings support the idea that different questions reward different forms of evidence. They do not yet provide a reliable decision framework for choosing a GEO method by industry or search intent.

Generalizability Remains Limited

The primary experiment used a custom architecture combining Google’s top five results with GPT-3.5-turbo. The researchers also tested selected treatments in Perplexity. The Perplexity test strengthens the paper because it shows that some effects were not limited entirely to the custom system, but it does not establish that the methods will perform consistently across all generative products.

Commercial systems may differ in their search indexes, query reformulation, retrieval models, reranking, context windows, summarization, prompts, citation policies, freshness controls, personalization, and interface design.

A content treatment that improves representation in one architecture may have little effect in another. The same source may be retrieved for one product, excluded by another, or cited differently depending on the query and system.

The paper references Bing Chat and Google’s Search Generative Experience, but neither was included in the reported evaluation. Testing them would have introduced practical challenges because commercial systems change frequently and offer limited experimental control. Even so, their omission limits the conclusions that can be drawn.

The study provides evidence that certain methods can work across more than one system. It does not provide evidence that they will work everywhere.

Text Is Only One Part of Generative Search

The GEO experiment focuses on cleaned webpage text and text-based responses with inline citations. That makes the framework primarily a text-content optimization study.

Generative search results may also include images, video, maps, product cards, local listings, reviews, tables, charts, feeds, and other structured elements. Google’s Search Generative Experience demonstrates that an AI-generated result may function as a multimodal interface rather than a block of synthesized prose. The factors influencing those elements may differ substantially from those affecting text synthesis.

A local business may gain visibility through its Google Business Profile, reviews, service information, location relationships, and map data. An ecommerce product may be selected through structured attributes, product feeds, pricing, availability, reviews, and category relationships. A video may be surfaced because of its transcript, chapters, title, visual content, or engagement patterns.

The GEO paper does not address these systems. It does not explain how images are selected, how product attributes are compared, or how map-based recommendations are formed.

A broader model of generative optimization will need to incorporate structured data, multimedia, entity relationships, reputation, reviews, feeds, and third-party information in addition to webpage text.

Long-Term Effectiveness Remains Unknown

Generative systems are developing quickly. Platforms can change their retrieval methods, language models, prompts, source-selection criteria, and citation interfaces without public notice. A method that improves visibility in one version of a system may lose effectiveness after an update. Widespread adoption may also reduce the value of a tactic.

If every publisher begins adding more quotations, statistics, and citations, those features may no longer distinguish one source from another. Systems may place more weight on originality, reputation, consensus, first-hand experience, or source quality.

The paper acknowledges that GEO methods will need to adapt as generative engines and query behavior evolve. That is likely to be a defining characteristic of the field.

Long-term advantages may come less from applying a fixed collection of content treatments and more from producing evidence competitors cannot easily replicate. Original research, proprietary data, documented experience, strong reputation, expert analysis, and credible external validation may provide more durable value than stylistic modifications alone.

What Future GEO Research Needs to Measure

The GEO study provides a foundation, but future research needs to examine the stages it leaves outside the experiment. Retrieval should be tested across a wider corpus rather than assumed through a fixed group of five sources. Researchers should evaluate whether specific content changes affect eligibility, source selection, citation frequency, and performance across different architectures.

Human behavior also needs to become part of the evaluation. Users should be observed to determine whether they notice citations, trust the source, click through, remember the brand, ask follow-up questions, or change their decision.

Accuracy should be treated as an outcome alongside visibility. A method should not be considered successful if it increases source representation by introducing unsupported claims or misleading evidence.

Commercial and local queries deserve more attention. Recommendation, comparison, provider selection, product evaluation, and location-dependent searches may depend on signals that differ from those associated with informational questions.

Multimedia and structured information should also be included. Images, maps, product feeds, reviews, videos, and business profiles are increasingly part of generative search experiences.

Finally, studies should track how results change over time. A method that works in one model version or product interface may not remain effective as systems evolve.

What the Study Ultimately Contributes

The limitations of the paper should not overshadow its significance. Aggarwal et al. recognized that visibility within a generated answer differs from visibility in a ranked list. They proposed metrics for measuring source representation, created a large benchmark, and demonstrated that changes to content can affect how a generative system distributes attention among retrieved sources.

The study also introduced questions that will become increasingly important for publishers and SEOs. It may no longer be enough to know whether a page ranks. We may also need to know whether it was retrieved, how much of the answer relied on it, where its citation appeared, how the entity was described, and whether users acted on the attribution. The GEO paper begins to measure the middle of that process.

Its findings are best understood as evidence that an already retrieved source can increase its influence during synthesis by providing information that is clear, specific, attributable, and useful. That is a meaningful contribution. It is not a complete theory of AI search visibility.

Final Thoughts

GEO: Generative Engine Optimization is an important first step in the empirical study of source visibility within AI-generated answers. It demonstrates that changes to content can alter how prominently a retrieved source appears in the final response.

The most important qualification is that the experiment begins with a fixed set of Google-ranked sources. It does not establish how a page becomes eligible for retrieval across the wider web, whether users notice or click citations, or whether the same methods will produce consistent results across different platforms.

In the end, the study provides a framework for understanding source influence during synthesis rather than a complete framework for earning visibility in generative search. That makes the research more useful because it places GEO within a larger process. Technical access, relevance, retrieval, entity understanding, evidence, reputation, synthesis, attribution, and user response all contribute to whether a source is ultimately seen, trusted, and chosen.

The paper gives us one of the first measurable pieces of that process. The next stage of research will need to connect those pieces into a broader understanding of how visibility actually works in generative search.