Ever since the introduction of AI tools such as ChatGPT and Bing Chat, SEOs and businesses have been wondering how to effectively attain visibility within them. As Google develops its Search Generative Experience and platforms such as Perplexity provide answers synthesized from multiple sources, the familiar list of search results is beginning to share space with AI-generated responses.
A new study by Aggarwal et al., titled GEO: Generative Engine Optimization, provides one of the first formal attempts to examine how content creators might improve their visibility within these systems. The researchers found that adding relevant citations, quotations, and statistics to source content could increase its visibility within generated answers by as much as 40% across the queries they tested. They also found that improving fluency and making content easier to understand produced meaningful gains.
The study introduces the term Generative Engine Optimization, or GEO, to describe the process of optimizing content for visibility within generative engines. Its findings offer an early framework for understanding how content may be selected, represented, and attributed when an AI system constructs an answer from several sources.
There is also encouraging news for websites that do not occupy the first traditional search position. The researchers found that lower-ranked sources sometimes received significantly larger visibility gains than sources already ranking first. Although the study did not directly compare small and large websites, the results suggest that generative search may create additional opportunities for sources that are retrieved but do not hold the highest organic position.
These are still early findings, and generative search systems will continue to evolve. Even so, the research provides a useful starting point for understanding how content presentation may influence visibility in AI-generated answers.
Explore In-Depth Insights on GEO
SEO and GEO: How Search Visibility Works Across Retrieval and Synthesis
Learn how SEO and GEO support different stages of visibility, from crawling and retrieval through extraction, synthesis, attribution, and the way users ultimately encounter information.
Generative Engine Optimization: A Critical Look at the GEO Study
A closer examination of the GEO study’s limitations, including its controlled retrieval setup, visibility metrics, category findings, and why the results should not be treated as universal AI ranking factors.
Free GEO GPT: Improve Content for Search and AI-Generated Answers
Use the free GEO GPT to review content for clarity, evidence, distinctiveness, structure, entity clarity, and information utility across search and AI-generated answers.
What is Generative Engine Optimization (GEO)?
Generative Engine Optimization refers to modifying website content to improve how prominently it appears within responses generated by AI-powered search systems. Aggarwal et al. use the term generative engine to describe systems that combine search and retrieval with generative models.
Instead of returning only a ranked list of webpages, a generative engine may reformulate a user’s question, retrieve relevant sources, summarize those sources, and synthesize the information into a single response. The resulting answer may contain inline citations or other forms of attribution that allow users to identify and verify the sources behind the information.

The paper references systems such as Bing Chat, Google’s Search Generative Experience, You.com, and Perplexity as examples of this emerging model. Although each platform may use a different technical process, they share the general objective of retrieving information and using generative models to construct a direct answer.
GEO is intended to help content creators understand how their information appears within those answers and how changes to their content might improve that representation. This introduces a different visibility problem from the one addressed by traditional SEO.
In a conventional search engine, visibility can often be approximated through rankings, impressions, clicks, and traffic. Generative engines combine information from multiple sources within a single response, which makes visibility more difficult to define. A source may support several sentences near the beginning of an answer, appear only once near the end, or contribute information without receiving substantial prominence.
Why Seek Visibility in Large Language Models?
Generative search provides businesses and publishers with an opportunity to gain exposure beyond traditional organic listings. At the same time, it may reduce the amount of attention users give to the familiar search results positioned beneath the generated answer.
Google’s Search Generative Experience provides an early example of how this may change the search results page. For a query such as “Acadia National Park,” the generated result can include a synthesized overview, links, references, images, maps, and suggested follow-up questions. These elements occupy a substantial portion of the page and can push traditional results below the fold.

For businesses, this creates both an opportunity and a risk. A source included prominently in the generated response may receive valuable exposure, even when it does not hold the first organic position. Conversely, a page may rank well in traditional search while receiving little attention if the generated answer satisfies the user before they reach the organic listings.
Visibility in this environment is no longer limited to where a page ranks. It also involves whether the page is retrieved, whether its information is incorporated into the answer, how prominently that information appears, and whether the source receives attribution.
How the GEO Study Was Conducted
The researchers created a benchmark called GEO-bench containing 10,000 queries from several datasets and subject areas. These included anonymized search queries from Bing and Google, question-answering datasets, debate prompts, questions from Perplexity’s Discover section, questions drawn from the ELI5 community, and queries generated with GPT-4.
The benchmark covered 25 domains and included queries with different levels of difficulty and intent. Approximately 80% of the queries were informational, while transactional and navigational queries each represented about 10% of the dataset.
For each query, the researchers retrieved the top five Google search results and supplied their cleaned content to a generative system. One of the five sources was then selected for modification, and the researchers applied each GEO method separately to determine how the change affected the source’s representation within the generated answer.

The study tested nine methods: authoritative language, statistics addition, keyword stuffing, source citations, quotation addition, easier-to-understand language, fluency optimization, unique words, and technical terms. The source selected for optimization remained the same across the different methods for each query, allowing the researchers to compare the effects of each modification.
The primary experimental system used GPT-3.5-turbo to generate answers from the five retrieved sources. The researchers generated five responses at a temperature of 0.7 and ran the experiments across several random seeds to reduce the influence of normal variation.
They also conducted a separate experiment using Perplexity. This allowed them to evaluate whether some of the strongest-performing methods could produce similar improvements within a publicly available generative engine.
How the Researchers Defined Visibility
One of the study’s most important contributions is its attempt to define visibility within a generated response. Traditional ranking measurements do not adequately account for the way generative engines combine and present information from multiple sources.
A source cited near the beginning of an answer may receive more attention than a source cited near the end. Similarly, a source supporting an entire paragraph may have more influence than one attached to a brief statement. The researchers therefore developed metrics designed specifically for generative responses.
Position-Adjusted Word Count
The primary metric was Position-Adjusted Word Count. This measured the proportion of the generated response associated with a particular source while assigning greater weight to information presented earlier in the answer.
A source received a higher visibility score when more of the response relied on its information and when that information appeared in a more prominent position. This allowed the researchers to account for both the amount of content attributed to the source and its placement within the response.
Subjective Impression
The study also introduced a Subjective Impression metric that considered several qualitative aspects of source visibility. These included the relevance of the cited material, its influence on the generated answer, the uniqueness of the information, the prominence of its position, the amount of material associated with the citation, the likelihood that a user would follow the citation, and the diversity of information supplied by the source.
These factors were evaluated using GPT-3.5 through a methodology based on G-Eval. The researchers then normalized the results so they could be compared with the Position-Adjusted Word Count scores.
Visibility in the GEO study therefore refers to the amount, position, and apparent influence of a source’s information within the generated answer. The study did not measure website traffic, referral clicks, leads, sales, or conversions.
Which GEO Strategies Performed Best?
The strongest overall methods were quotation addition, statistics addition, and source citations. Across the benchmark, these methods produced relative gains of approximately 30% to 40% on the Position-Adjusted Word Count metric and between 15% and 30% on the Subjective Impression metric.

Quotation addition produced the strongest result on the primary metric. It increased the Position-Adjusted Word Count score from 19.3 for the unmodified baseline to 27.2, representing an improvement of approximately 41%.
Statistics addition and source citations also performed well, although their results varied by query type, subject area, metric, and the original search position of the source. The study’s widely reported “up to 40%” result is therefore accurate, but it should not be interpreted as a guaranteed improvement for every page or query.
The broader finding is that content containing specific, attributable, and evidence-supported information generally performed better than content modified primarily through stylistic changes or repeated keywords.
How Citations May Improve AI Search Visibility
The Cite Sources method added relevant citations from credible sources to the website content. The researchers found that this method could significantly improve visibility, particularly for queries involving factual statements.
Citations connect claims to identifiable evidence. They give readers a way to verify the information and may also make it easier for a generative system to identify which claims are supported by an outside source.
Consider the difference between a general statement and one supported by attribution. A passage might state that a regular slice of pizza contains a meaningful amount of protein. A more specific version could explain that, according to Pizza Bien, a regular slice contains between 13 and 22 grams of protein, or approximately 25% of the recommended daily value in a 2,000-calorie diet.
The second version provides a measurable claim and identifies its source. This gives the reader additional context and gives the generative system a more specific unit of information to extract and potentially incorporate into an answer.
The study did not separately evaluate hyperlinks, anchor text, citation formats, or reference sections. It tested the addition of relevant source attribution. From a publishing perspective, linking to the original source remains a useful practice because it allows readers to verify the information, but the study does not establish that the HTML link itself caused the visibility improvement.
How to Cite Sources Effectively
Citations should directly support the statement in which they appear. A source should not be included simply because it is well known or generally authoritative. It should contain evidence that is relevant to the specific claim being made.
Whenever possible, content creators should link to the original research, public record, interview, report, or dataset rather than relying on a secondary article that summarizes it. This gives readers access to the original context and reduces the chance that a finding will be misrepresented as it passes from one publication to another.
For non-digital sources, the article should include enough identifying information for the reader to locate the original work. A reference section using an established citation format may also be appropriate for research-heavy content.
These practices are valuable regardless of their effect on generative search. Citations demonstrate that the writer has conducted appropriate research, allow claims to be verified, and give credit to the people and organizations responsible for the original work.
The Effective Use of Quotations
Quotation addition produced the largest improvement in the study’s primary visibility metric. A relevant quotation provides information in a distinct and attributable form, often capturing an expert explanation, first-hand observation, or original perspective that cannot be communicated as effectively through a generic summary.
For example, in an interview with Studio Potter, Ben Cohen discussed the challenge of integrating social responsibility with profitability at Ben & Jerry’s. He explained, “The challenging part of Ben & Jerry’s business has been figuring out how to integrate a concern for community with making a profit.”
The quotation adds a direct perspective from someone involved in the organization. It contributes more than a general statement that Ben & Jerry’s values social responsibility because it shows how one of the company’s founders described the relationship between its commercial and community objectives.
Quotations should be selected because they add meaning, not simply because the study found that quotations performed well. A useful quote should clarify a point, support a conclusion, introduce first-hand experience, or provide wording that deserves to be preserved in its original form.
The original interview, speech, report, or publication should be linked whenever possible. This allows readers to review the quotation within its complete context and confirm that it has been represented accurately.
Statistics and Content Credibility
Statistics can make content more specific and information-dense. A broad statement may communicate a general trend, but a well-supported statistic allows readers to understand the scale, direction, or significance of that trend.
For example, an article could state that gummy candy is becoming more popular. A more useful version might explain that, according to The New York Times, sales of chewy candy in the United States, including gummies, reached $4.6 billion in 2021 and increased by nearly 15% from the previous year.
The statistic provides a measurable basis for the conclusion. It shows both the size of the market and the rate at which it was growing at that time.
Pizza, ice cream, and gummy bears appear to be a recurring theme in my examples. Apparently, writing about research while hungry has consequences.
Statistics should still be used selectively. A number does not improve content simply because it creates the appearance of precision. The statistic should come from a credible source, directly support the surrounding claim, and include enough context for the reader to interpret it correctly.
The date and methodology also matter. A statistic that accurately described a market several years ago may no longer reflect current conditions. Whenever possible, writers should use the original dataset or report and explain any qualifications that affect how the number should be understood.
Why Citations, Quotations, and Statistics May Work
The researchers demonstrated that citations, quotations, and statistics improved visibility, but they did not establish the exact mechanism responsible for those gains. Several explanations remain possible.
These additions may make a passage more specific, information-dense, distinctive, or easier to attribute. They may also help the generative system identify clear factual claims that can be incorporated into a response without requiring extensive interpretation.
A vague statement gives the system relatively little material to work with. A statement containing an identified source, measurable figure, or direct quotation provides a clearer piece of evidence that can be extracted and connected to the user’s question.
This suggests that generative systems may respond not only to topical relevance but also to the usefulness of the information during answer construction. Content that is well supported and clearly expressed may be easier to summarize, combine with other sources, and attribute within the final response.
This does not mean that publishers should add statistics or quotations to every section. The objective is not to decorate the article with signals that resemble authority. The objective is to provide evidence that improves the quality and usefulness of the explanation.
Optimizing Content for Fluency
The study also found that improving fluency could increase source visibility. Fluency Optimization performed well on the Position-Adjusted Word Count metric and produced a stronger result than the unmodified baseline.
The researchers used a language model to rewrite the source content according to a prompt requesting improved fluency. They did not separately test sentence length, paragraph length, transition words, active voice, reading level, or heading structure.
The study therefore supports the broader conclusion that clearer presentation may help a generative system use the content. It does not establish a specific writing formula.
For example, the statement “Impressions in the SERPs improved by 14%” assumes the reader understands both “impressions” and “SERPs.” A clearer version would state that search impressions increased by 14%, meaning the pages appeared more often in Google’s search results.
The revised version explains the terminology without adding an unsupported interpretation. Clarity does not require making every sentence longer. It requires making the relationship between the information and its meaning easier to understand.
Making Content Easier to Understand
The Easy-to-Understand method also improved visibility compared with the baseline. As with fluency optimization, the researchers used a broad rewriting prompt rather than testing individual formatting or writing rules.
The result suggests that generative systems may benefit from source content in which the main ideas and relationships are clearly expressed. Technical terms should be explained when the intended audience may not understand them, and acronyms should generally be defined when they first appear.
Headings should accurately describe the material that follows, and paragraphs should develop one primary idea without introducing unnecessary ambiguity. Sentence and paragraph length should be determined by the complexity of the idea rather than an arbitrary rule.
Writing for clarity does not require removing technical depth. Some subjects cannot be explained accurately without specialized terminology and detailed analysis. The objective is to provide enough context for the reader to follow the argument and understand how each concept relates to the larger topic.
A reader and a generative system should both be able to identify what claim is being made, which entity the claim refers to, what evidence supports it, and how it contributes to the answer.
Did an Authoritative Tone Improve Visibility?
The researchers also tested whether rewriting content to sound more persuasive and authoritative would improve visibility. The results were mixed.
The Authoritative method performed better than the baseline on several measurements, particularly the Subjective Impression metric. It also produced stronger results for certain categories, including debate, history, and science.
However, the method did not create the consistent improvements observed with quotations, statistics, citations, and fluency optimization. The researchers concluded that generative engines appeared relatively resistant to changes based primarily on persuasive tone.
This does not mean that authority is unimportant. It suggests that sounding authoritative is not the same as demonstrating authority.
A writer can adopt a confident tone without supplying evidence, expertise, or original insight. Authority is more convincingly established through first-hand experience, original research, transparent methodology, relevant credentials, documented results, and support from credible outside sources.
The findings therefore support substance over posture. Content creators should focus less on making claims sound authoritative and more on demonstrating why those claims deserve to be trusted.
What the Study Found About Keyword Stuffing
Keyword stuffing performed poorly on the study’s primary visibility metric. The researchers modified the source content to include additional keywords from the user’s query, following what they characterized as a traditional SEO tactic.
This approach did not improve the source’s representation within the generated response. In fact, its Position-Adjusted Word Count score was lower than the unmodified baseline.
The finding should not be interpreted as evidence that keywords, search intent, or topical relevance no longer matter. The study did not test keyword research, semantic coverage, titles, headings, internal linking, content depth, or on-page SEO as a complete discipline.
It tested the mechanical addition of more query terms. That method did not make the source more useful to the generative system.
Relevance remained essential to the experiment. The five documents supplied to the generative model were retrieved from Google because they were considered relevant to the query. The result suggests that once a relevant document has been retrieved, repeating additional query terms does not necessarily improve how the system uses it.
Results Varied by Query Type
The effectiveness of each GEO method varied across subject areas and query categories. This is one of the study’s most important findings because it suggests that a single optimization formula may not work equally well for every type of content.
Source citations performed particularly well for factual statements and queries related to law and government. Quotation addition produced strong results for queries involving people and society, explanations, and history. Statistics performed well for law and government, debate, and opinion-related queries.
The Authoritative method was more effective for debate and history queries, while Fluency Optimization performed well for business, science, and health topics. These differences suggest that the best way to improve content depends partly on the information the user is seeking.
An article explaining a scientific process may benefit from clear definitions, measured claims, and supporting statistics. A historical article may depend more heavily on primary sources and relevant quotations. A product comparison may require structured attributes, documented differences, and evidence tied to specific use cases.
GEO should therefore be approached as a context-dependent process rather than a checklist applied uniformly to every page.
Combining GEO Strategies
The researchers also examined the effects of combining the four strongest methods: fluency optimization, statistics addition, source citations, and quotation addition.
The best-performing combination was Fluency Optimization with Statistics Addition. This pairing outperformed the strongest individual strategy by more than 5%, although the combination analysis was conducted on a smaller subset of 200 examples because of the cost of the experiment.
Citations also performed more effectively when combined with other methods than when used alone. This suggests that attribution becomes more useful when the surrounding information is also specific, clear, and relevant.
The result reflects how high-quality content is normally created. A statistic becomes more useful when its meaning is explained clearly. A quotation becomes more valuable when its relevance is established. A citation works best when the claim it supports is precise enough for the reader to understand what is being verified.
The findings suggest that GEO should not be reduced to isolated tactics. The strongest content may combine clear writing, specific evidence, credible attribution, and information selected for the needs of the query.
Can GEO Help Lower-Ranked Websites?
The researchers also examined how GEO methods affected sources occupying different positions among the five Google results used in the experiment.
Lower-ranked sources frequently experienced substantially larger relative gains than sources already ranking first. For example, adding citations increased the visibility of the fifth-ranked source by 115.1%, while the visibility of the first-ranked source declined by 30.3%.
Quotation addition and statistics addition followed a similar pattern. Sources in the fourth and fifth positions generally received larger improvements than sources in the first and second positions.
The researchers argued that this could create opportunities for smaller publishers and independent creators that often struggle to compete with larger organizations in traditional search. Their interpretation is reasonable, but the study did not directly compare businesses according to size.
It did not measure revenue, website size, brand recognition, backlink strength, or domain authority. The comparison was based on the source’s original position among the five Google results provided to the generative engine.
The study therefore shows that lower-ranked sources may gain more from certain content improvements after they have entered the retrieval set. This may benefit smaller publishers, but it does not demonstrate that smaller websites automatically hold an advantage in generative search.
Retrieval Comes Before Generation
The experimental design highlights an important limitation of the findings. The researchers began with the top five Google search results, meaning the source selected for optimization had already been found, indexed, ranked, and retrieved.
The study then measured whether modifying that source changed how prominently its information appeared in the generated answer. The experiment therefore focused primarily on the synthesis stage rather than the complete process through which a source earns visibility.
Before a generative engine can use information from a webpage, the system generally needs to access the content, understand its subject, determine that it is relevant, select it for retrieval, extract useful information, and incorporate that information into the final answer.
The GEO methods tested in the paper appear most directly related to the later stages of this process. Adding a statistic may make a source more useful after retrieval, but it does not guarantee that the source will be retrieved.
This is one reason GEO should not be treated as a replacement for SEO. Technical accessibility, relevance, internal linking, content architecture, authority, backlinks, and other established search considerations may still influence whether a page becomes part of the source set.
Traditional SEO may help a page become eligible for retrieval. GEO may help its information become more useful once it is retrieved.
Testing GEO in Perplexity
The researchers also tested selected GEO methods within Perplexity. This provided an opportunity to determine whether some of the findings would carry over from the experimental system to a publicly available generative engine.
Quotation addition improved the Position-Adjusted Word Count metric by approximately 22% compared with the baseline. Statistics addition produced an improvement of up to 37% on the Subjective Impression metric.
Keyword stuffing again performed poorly on the primary visibility metric. Its Position-Adjusted Word Count score was approximately 10% lower than the baseline.
These results provide some real-world support for the primary experiment, but the Perplexity evaluation was more limited. It tested fewer methods and does not demonstrate that the same modifications will perform equally across every platform.
Different generative engines may use different search indexes, retrieval methods, language models, prompts, citation processes, and ranking systems. The researchers acknowledge that GEO methods may need to change as generative engines and user behavior evolve.
How to Apply the Findings to Your Content
The GEO study gives content creators several practical ideas to test, but the methods should be applied according to the needs of the article rather than as mechanical requirements.
Begin by reviewing the important claims in the content. Determine whether they can be supported with original research, public data, expert commentary, case studies, or other credible evidence. General statements should be made more specific when a measurable result, documented example, or direct quotation would improve the explanation.
The clarity of the writing should also be evaluated. Important information should not be hidden beneath unnecessary introductions or unexplained terminology. The meaning of a statistic should be explained, and the relevance of a quotation should be clear from the surrounding context.
Most importantly, every modification should improve the content for the reader. An outdated statistic does not become useful because a generative engine may respond to numbers. An unrelated quotation does not strengthen the article simply because quotation addition performed well in the study.
The evidence must be credible, relevant, and connected to the question being answered.
GEO and E-E-A-T
The GEO study did not test Google’s E-E-A-T framework, quality raters, or traditional ranking systems. It would therefore be inaccurate to claim that E-E-A-T caused the visibility improvements observed in the experiment.
However, several of the strongest-performing methods align with established practices for producing credible content. Citations allow readers to verify claims, quotations provide identifiable first-hand or expert perspectives, statistics offer measurable evidence, and clear writing makes the information easier to interpret.
These practices may contribute to trust, but their relationship to the study should be described carefully. The research demonstrates that supported and clearly presented information gave the generative system more useful material to incorporate into its answers.
That finding is compatible with the broader objective of producing trustworthy content, even though the experiment did not evaluate E-E-A-T directly.
Limitations of the GEO Study
The researchers identify several limitations that should be considered when applying the findings.
The study evaluated two generative environments: an experimental system built with Google results and GPT-3.5-turbo, and a more limited test using Perplexity. Other generative engines may respond differently to the same modifications.
The benchmark includes a broad range of queries, but user behavior is likely to change as people become more familiar with conversational search. Queries may become longer, more specific, and more dependent on personal context.
The researchers also did not evaluate how the GEO modifications affected traditional Google rankings. A change that improves representation in a generated response could have a neutral, positive, or negative effect on organic performance.
Most importantly, the study does not show that adding citations, statistics, or quotations will automatically cause a page to be retrieved, cited, clicked, or visited. It also does not measure leads, sales, or other business outcomes.
These limitations do not diminish the value of the research. They define what the study demonstrates: once a source is available to a generative engine, the way its information is written and supported can affect how prominently it appears within the resulting answer.
Improving Visibility in Generative Search
The GEO study provides an important starting point for understanding how visibility may work when search systems generate answers from multiple sources.
The strongest-performing content was not the content that repeated the most query terms or adopted the most persuasive tone. It was content that offered specific, supported, and attributable information.
Relevant quotations, credible source citations, and useful statistics produced the largest gains. Clear and fluent writing also improved performance, particularly when combined with evidence-rich content.
These are not entirely new publishing practices. Good writers have always supported their claims, identified their sources, explained complex ideas clearly, and provided readers with enough information to evaluate the argument.
What is changing is the role those practices may play within search. A webpage is no longer competing only for a position within a list of links. Its content may also be retrieved, extracted, combined with information from other sources, and presented within an AI-generated answer.
Generative Engine Optimization begins with understanding that process. The objective is not simply to mention the correct keywords or make the article sound more authoritative. It is to create information that a generative system can understand, use accurately, connect to the question, and attribute to the source.
The research remains preliminary, and generative search will continue to develop. Even so, the study offers a valuable framework for preparing content for a search environment in which visibility may depend not only on being found, but also on being useful enough to become part of the answer.
