Back

GEO: The New Power of LLMs, or When AI Decides Who Exists

Summary

Large language models (LLMs) have become the new arbiters of informational visibility. Brands, researchers, media outlets, individuals: any entity absent from their responses is, in effect, invisible to a growing share of users. This power of selection rests on documented structural biases (notoriety bias, gender bias, recency bias, AI-to-AI self-reference bias), an extreme concentration of cited sources, and a large-scale hallucination phenomenon that pollutes collective knowledge by attributing an “existence” to fictitious entities. Understanding these mechanisms is a first-order strategic imperative for any organization engaged in content production, reputation management, or scientific research.

The question “who appears in AI responses” is on the verge of supplanting “who appears on Google’s first page.” This shift is not incidental. ChatGPT now handles 2.5 billion daily queries from 883 million monthly users, and AI-driven referral traffic grew by 527% between early 2024 and early 2025. In this context, the act of citing, or not citing, is no longer the neutral gesture of a search engine listing URLs: it is a decision, algorithmic in nature, but whose consequences for the reputation, credibility, and perceived existence of an entity are now measurable. LLMs have become de facto publishers, without bearing the declared responsibilities that come with that role.

GEO-LLM-Power

LLMs as the New Gatekeepers of the Informational Space

An Unprecedented Concentration of Sources in Web History

The open web theoretically allowed any indexed site to appear in search results. The architecture of LLMs follows a fundamentally different and far more restrictive logic. A generative engine cites an average of two to seven domains per response, compared to the ten traditional blue links served by Google. This radical compression of the source spectrum creates an unprecedented winner-takes-all effect at this scale. An analysis of 150,000 LLM citations conducted by Semrush in June 2025 reveals that Reddit accounts for 40.1% of cited sources, Wikipedia for 26.3%, and YouTube for 23.5%, with no other platform reaching 5%. Three platforms thus absorb nearly 90% of global visibility. For any company, researcher, or media outlet absent from these dominant ecosystems, the probability of appearing in an AI response is marginal, regardless of the intrinsic quality of their content.

The RAG Logic and the Filter of Perceived Authority

Modern generative engines operate primarily through Retrieval-Augmented Generation (RAG): they query external sources in real time to supplement their parametric knowledge. This retrieval logic prioritizes three criteria: relevance, recency, and trust. E-E-A-T (Experience, Expertise, Authority, Trustworthiness), a cornerstone of SEO, remains central to GEO. What changes is the nature of the authority signal. Brand search volume is the strongest predictor of LLM citations, with a correlation coefficient of 0.334, surpassing the impact of traditional backlinks. In other words, being known now precedes being well-positioned. Entities that benefit from pre-existing recognition in training data hold a considerable structural advantage over emerging entities, regardless of their actual expertise.

An Already Highly Concentrated AI Visibility Market

Brands with the highest share of voice in LLM responses are typically those that invested earliest in SEO. Technical health, structured data, and authority signals remain the foundation of AI visibility. This means that the competitive advantage accumulated on the traditional web transfers, and is often amplified, within the LLM environment. Gartner projects a 25% decline in traditional search volume by 2026, and IDC anticipates that companies will spend up to five times more on LLM optimization than on classic SEO by 2029. The economic stakes of this concentration are therefore no longer prospective: they are immediate.

The Structural Biases That Manufacture Existence and Invisibility

Notoriety Bias: The Powerful Cite the Powerful

Algorithmic neutrality is a fiction that empirical data consistently refutes. In the scientific domain, a particularly well-documented phenomenon illustrates the mechanics of reinforcing existing inequalities. A study published in May 2026, analyzing 111 million references extracted from 2.5 million publications on arXiv, bioRxiv, SSRN, and PubMed Central, estimates that LLMs produced 146,932 hallucinated citations in 2025. These errors disproportionately assign credibility to already prominent researchers and to men. The LLM does not create inequality from scratch: it reads it in its training data and reproduces it at scale. Entities already visible become even more visible; already marginal entities disappear further still.

Gender and Representation Bias: A Documented Discrimination

A study published in PNAS in July 2025 reveals that LLMs favor communications produced by other LLMs, and that they are likely to introduce anti-human discrimination into decision-making processes. Furthermore, 91% of LLMs are trained on data extracted from the open web, where women are underrepresented in 41% of professional contexts and minority voices appear 35% less frequently. This representational bias in training data directly translates into outputs. LLMs used in academic writing assistance tools filter or rank candidate references, and if their internal scoring is biased, those discriminations propagate into literature syntheses, automated evidence reviews, and other derivative applications. The scientific existence of a female researcher, an African laboratory, or an institution from the Global South thus remains structurally threatened by systems that present themselves as objective.

Recency Bias: Only the Recent Present Exists

The temporal dimension of LLM selection constitutes another massive filter of existence. An analysis of citation habits across ChatGPT, Perplexity, and AI Overviews reveals that 65% of cited content was published within the past year, 79% within the past two years, and 89% within the past three years. For any content predating 2021, the probability of citation collapses to 6%. This recency bias produces institutional amnesia: foundational works, in-depth analyses, and reference journalistic archives disappear from the active informational space in favor of a continuous news stream that is sometimes superficial. An organization whose content production is old or irregular finds itself structurally excluded from the operational memory of LLMs.

Hallucinations and Fictitious Existence: When AI Creates and Destroys

The Fabrication of Non-Existent Entities as a New Systemic Risk

The power of LLMs over existence is not limited to ignoring real entities: it extends to the creation of fictitious entities endowed with an appearance of legitimacy. An analysis of papers accepted at the NeurIPS 2025 conference, one of the most prestigious in global AI, identified 100 hallucinated citations spread across 53 publications, despite review by three to five experts per article. These citations are not typographical errors; they are references to non-existent sources, with invented authors, fabricated titles, and false publication information. Hallucination therefore confers documentary existence upon works that never existed, while potentially erasing those that do exist but do not match the statistical patterns expected by the model.

The Paradox of Fictitious Authority

This phenomenon produces a troubling paradox for collective integrity. A hallucinated reference from a non-existent journal or an invented author enters bibliographic catalogs alongside legitimate citations, with no indication of its fictitious nature. Bibliographic databases such as Web of Science, Scopus, or Google Scholar provide verification features that require manual querying of each citation, a process that does not scale. The systemic risk is real: entire scientific corpora may incorporate ghost references that will in turn be used to train the next generations of LLMs, creating a self-sustaining contamination loop. Fictitious existence begets fictitious existence.

Practical Consequences for Content and Reputation Actors

For organizations that depend on their informational reputation, the LLM existence/invisibility dynamic imposes a profound strategic reassessment. Adding statistics to content increases AI visibility by 22%, and including citations increases that visibility by 37%. These actionable levers confirm that content structure is now designed as much for human comprehension as for algorithmic ingestion. Online presence is no longer a passive state: it is an active infrastructure to be maintained according to the specific logic of each inference system.

What Brands Can Do Today to Exist in LLM Responses

Build a Multi-Signal Presence Rather Than Optimizing a Single Channel

The temptation to treat GEO as an enhanced version of SEO is a framing error that leads to insufficient strategies. LLMs construct their representation of a brand from the totality of available public signals, including conversational, editorial, and community signals, not from a single well-structured website. The first priority is therefore to act simultaneously on three levels: what third parties say about the brand (customer reviews, Reddit mentions, sector forums), what the brand itself produces (structured content, semantically dense and regularly updated), and how the brand responds publicly (social care, community management, crisis management). A dissatisfied customer whose complaint goes unanswered publicly is a negative training data point for future models, just as a poorly sourced article is. An effective GEO strategy treats every public interaction as micro-content potentially absorbed by an LLM, which requires coordinating teams and tools that, until now, operated in silos.

Structure Content to Be Extractable, Not Just Readable

Content structure is the primary actionable lever for improving the probability of citation by an LLM. A generative engine does not read an article the way a human does: it looks for extractable passages, direct assertions, quantified data, and clearly identifiable named entities. Concretely, this means opening each section with a direct answer to the reader’s implicit question, integrating at least one sourced figure per thematic block, and using H2/H3 headings formulated as questions or complete assertions rather than generic labels. The FAQ format is also a direct lever: each verbatim from social listening (“I’m torn between X and Y for…”) is a real question that thousands of users will ask an LLM, and answering it explicitly in brand content amounts to positioning on a conversational query before it is even formulated.

Actively Monitor Your Algorithmic Reputation as You Would Your E-Reputation

Visibility in LLM responses can be measured and managed, yet this monitoring dimension is often the first to be neglected. Dedicated tools such as Profound, Brandlight, or Semrush’s AI modules now make it possible to track how frequently a brand is cited by ChatGPT, Gemini, or Perplexity, on which queries, with what associated sentiment, and against which competitors. This data is strategically different from classic SEO rankings: a brand may dominate Google page 1 and be nearly absent from AI responses, or vice versa. Tracking share of voice in LLM responses is an emerging KPI to integrate into any digital performance dashboard. Regularly verifying the consistency of entity data (name, description, positioning) across all external properties, including Reddit, Trustpilot, LinkedIn, and Wikipedia, constitutes a baseline hygiene upon which the consistency of the brand signal that models absorb depends.

Netino: The Only Player Acting Across the Entire LLM Signal Ecosystem

Netino’s GEO approach is fundamentally distinct from that of SEO agencies or traditional content marketing providers, precisely because its core competencies correspond point by point to the signals that LLMs consult to construct their representation of a brand. While most players intervene on the production of optimized content, a single lever among the seven identified, Netino acts simultaneously on mass reputation via its 800 multilingual community managers, on preventive moderation that maintains a favorable signal-to-noise ratio (500 million pieces of content moderated at a 96% quality score), on the public social proof generated by a 95% Social Care response rate in under 30 minutes and a CSAT score above 85%, and on the intelligence of real conversational queries through its 50 insights specialists covering more than 30 markets and languages. This last point is particularly structural: verbatims collected via social listening constitute the exact raw material for the questions audiences will ask an LLM tomorrow, and treating them as a GEO FAQ corpus allows brands to anticipate queries rather than merely react to them. Netino is also the only player operating at the level of the training layer itself, through its Data Annotation capabilities compatible with LLM standards (question/answer pairs, intent classification, sentiment detection), opening the door to direct influence on future sector-specific models, and not merely on their real-time retrieval behavior.

Conclusion

The challenge for content, research, and communication actors is no longer simply to produce quality: it is to make that quality legible, structured, and traceable to systems that do not read like humans, do not judge like editors, and do not remember like archivists. Understanding the selection mechanisms, biases, and hallucinations of LLMs has become a critical competency, not in order to “optimize for AI” in the superficial sense of the term, but to avoid letting the algorithm alone decide what deserves to exist.

Sources

  • Semrush / Visual Capitalist, Analysis of 150,000 LLM Citations, June 2025
  • Zhenyue Zhao et al., LLM hallucinations in the wild: Large-scale evidence from non-existent citations, arXiv, May 2026
  • Samar Ansari, Compound Deception in Elite Peer Review: A Failure Mode Taxonomy of 100 Fabricated Citations at NeurIPS 2025, arXiv, February 2026
  • PNAS, AI–AI bias: Large language models favor communications generated by large language models, July 2025
  • arXiv, Who Gets Cited? Gender- and Majority-Bias in LLM-Driven Reference Selection, AAAI 2026
  • Princeton / KDD 2024, GEO: Generative Engine Optimization
  • Seer Interactive, Study: AI Brand Visibility and Content Recency, June 2025
  • Gartner, Traditional Search Volume Projections 2026, 2025
  • Profound / Mersel AI, Generative Engine Optimization Guide 2026
  • Contently, Top 10 Sources LLMs Cite Most in 2026, April 2026

Netino
Netino

Leave a Reply

Your email address will not be published. Required fields are marked *