Category: Credibility

  • The GEO Advice You Followed Was Written for a World That No Longer Exists

    The GEO Advice You Followed Was Written for a World That No Longer Exists

    In January 2025, ChatGPT held 86.7% of AI chatbot web session share. By January 2026, that figure had fallen to 64.5%. US mobile share had dropped below 40% for the first time. In the same twelve months, Google Gemini grew from 5.7% to 21.5% — nearly a fourfold increase. Perplexity grew 370% year-on-year.

    Most published GEO advice was written when ChatGPT had 87% of the market. At that concentration, treating “optimise for AI” and “optimise for ChatGPT” as synonymous was reasonable. It is no longer reasonable. The market has fragmented faster than the advice has updated.


    What was built for 87%

    The practical playbooks for AI visibility — which outlets to target, which content formats to prioritise, which technical signals matter — were largely derived from studying ChatGPT’s citation behaviour. ChatGPT was the obvious choice: it was available for testing, it had public documentation, and at 87% it was functionally the entire market. When practitioners said “AI search,” they meant ChatGPT.

    That produces a body of advice that is not wrong so much as platform-specific. Targeting Reuters, the Financial Times, and Axios makes excellent sense for ChatGPT — Muck Rack’s analysis of 1 million+ citations (July 2025) confirms these are among ChatGPT’s most-cited journalism outlets. Prioritising fresh coverage from the last twelve months makes sense for ChatGPT, which draws 56% of its journalism citations from that window.

    This advice is well-evidenced and worth following. It is just not advice for “AI.” It is advice for ChatGPT — written at a moment when that distinction did not seem to matter.


    Why Gemini’s rise is not the threat it looks like

    The obvious interpretation of the market share shift is that Gemini has eaten into ChatGPT’s dominance. That is true. The less obvious interpretation is that Gemini’s growth is, paradoxically, an argument for traditional search quality — not against it.

    The reason is architecture. University of Toronto researchers tested domain overlap between each major AI engine and Google’s top-10 search results across 1,000+ queries (arXiv:2601.16858, January 2026). The results by engine:

    • GPT-4o: 4.0% overlap with Google
    • Gemini: 11.1% overlap with Google
    • Claude: 12.6% overlap with Google
    • Perplexity: 15.2% overlap with Google

    Gemini uses Google Search as its retrieval grounding. Its domain overlap with Google is nearly three times ChatGPT’s. The platform taking market share from ChatGPT is the one most tightly coupled to the search index that SEO builds. As Gemini’s share grows, the fraction of AI interactions that run through Google’s retrieval infrastructure grows with it.

    The claim that AI search requires a strategy separate from traditional search is structurally weaker today than it was a year ago — not stronger.


    One partial reprieve

    There is a second piece of data that reduces the complexity somewhat. Muck Rack’s outlet analysis finds that ChatGPT and Gemini share an almost identical journalism citation profile: Reuters, Financial Times, Time, Forbes, Axios. The two biggest non-Google AI engines — one holding 64% of the market, the other growing fastest — are drawing from the same outlet pool.

    For brands that built their earned media strategy around ChatGPT’s citation preferences, Gemini’s rise may not require wholesale reprioritisation. The outlet list that serves ChatGPT is likely to serve Gemini at similar rates. The outlet divergence problem is currently concentrated in Claude, not Gemini.

    This does not mean the advice translates frictionlessly. Gemini’s higher Google-grounding means that its citation behaviour is also more dependent on search ranking than ChatGPT’s — the 4% vs 11.1% domain overlap gap is not just trivia, it is a statement about what prerequisite work you need to have done. But the outlet-level strategy is more transferable than the platform shift headline implies.


    Where the fragmentation actually bites

    The sharper problem is not ChatGPT-to-Gemini substitution. It is the growth of everything else.

    Claude is structurally different. It cites Reuters approximately 50 times less than ChatGPT (Muck Rack, July 2025). Its top journalism sources — Good Housekeeping, TechRadar, Harvard Business Review — have almost nothing in common with the wire-service profile that serves ChatGPT and Gemini. Claude also operates on a longer temporal window: only 36% of its journalism citations come from the last twelve months, compared to 56% for ChatGPT.

    Perplexity grew 370% year-on-year from a smaller base but 1.2 billion monthly AI chatbot sessions (Similarweb / Vertu, 2026) means even minority platforms carry volume. Perplexity’s source mix draws heavily on domain overlap with organic search (15.2%, the highest of the four) but blends in social and video content in ways the other engines do not.

    Semrush’s longitudinal tracking (October 2025) documents the platform-specific volatility: Reddit dropped 82% in ChatGPT’s citation share while rising 74% in Google AI Mode in a single quarter. The same source, moving in opposite directions simultaneously, on the two biggest platforms. That is not noise — it is a structural signal that platform-specific dynamics are already operating at a level that single-channel strategy cannot capture.


    What a Gemini-first strategy looks like

    For most brands, the market share data points to the same practical conclusion from two directions.

    First: Gemini’s growth strengthens the case for search quality. Its retrieval architecture means that ranking in Google’s index is not just a prerequisite for ChatGPT citation — it is a more direct input for Gemini. The fraction of AI sessions where search ranking materially influences citation outcomes is growing, not shrinking.

    Second: the outlet strategy for the ChatGPT/Gemini bloc (now roughly 85% of the market combined) still centres on wire services, financial press, and authority publications — and it requires consistent recent coverage, not historical presence.

    What is not covered by that strategy is Claude — and to a lesser extent, Perplexity. For brands whose customers index toward research-oriented, knowledge-worker, or specialist professional contexts, Claude’s growing share matters in ways a wire-service PR strategy will not address.

    The advice most brands received about AI visibility is not wrong. It is increasingly incomplete — and the incompleteness has a specific shape. The market that the advice was written for is gone. The question now is which slice of the new market your customers actually live in.

  • Why Mid-Sized Brands Are Locked Out of AI Knowledge

    The Structural Exclusion Problem: Why Mid-Sized Brands Are Locked Out of AI Knowledge

    Most GEO and EC advice frames the SME visibility problem as a competitive disadvantage. You are behind the large players. Here is how to close the gap.

    That framing is wrong — or at least, it is not wrong enough. The problem is not that mid-sized brands are losing a race. It is that the race was designed without them.


    The Training Data Problem Is Not Random

    Before AI systems answer questions about your brand, your industry, or your category, they have already formed a view. That view was assembled during pre-training — the process by which a model learns the world’s knowledge from an enormous corpus of text, before it ever answers a query.

    That corpus was not a neutral sample of what exists. It was a weighted sample of what had been digitally published, cited, covered in large-circulation media, and documented in the reference sources that trained models weight most heavily: Wikipedia, academic databases, major editorial outlets, trade press with decades of archive depth.

    Large brands — the multinationals, the household names, the category incumbents — accumulated exactly these kinds of presence over decades. They had Wikipedia pages. They had Reuters coverage. They had academic case studies, analyst reports, Financial Times profiles. That presence existed before any training data was assembled. When the models trained on the web, they trained on a web that had already organised itself around the visible and the established.

    Mid-sized SMEs, by design, had none of this. A regional services firm with thirty years of operational excellence but no analyst coverage and no Wikipedia entry had produced no signal that pre-training data collection would recognise as authoritative. The AI system did not decide that firm was unimportant. The training data never recorded its importance in the first place.

    This is not market inefficiency. It is structural exclusion embedded in how knowledge was assembled.


    What the Evidence Shows

    The academic basis for this is not theoretical. Chen et al. (arXiv:2601.16858, January 2026) ran perturbation experiments on GPT-4o: they manipulated the evidence the model received about brands — shuffling retrieved snippets, restricting retrieval to only provided content, injecting brand names into unrelated material — and measured how much those manipulations moved the model’s output rankings.

    For well-known, popular brands, the results were stark. Average rank deviation under snippet manipulation: 2.30–2.60. For niche entities: 4.15–4.63. Popular brand rankings barely shifted regardless of what the retrieved evidence said. The model already knew the answer. The citation miss rate for Cadillac was 58%. For Infiniti, 73%. Those brands appeared in AI answers without any supporting retrieved content more than half the time — drawn entirely from training priors. (Chen et al., arXiv:2601.16858)

    The mechanism is clear. Popular entity rankings are governed by pre-trained knowledge. Retrieved evidence is used to confirm what the model already believes, not to discover who deserves to be recommended. For niche brands — where the model holds no stable prior — retrieval actually drives the answer. The two populations are not playing the same game.

    The training-data corpus itself reflects this asymmetry. Analysis of earned media citation patterns shows that 82% of AI citations come from earned media sources — but those sources are heavily concentrated in a small cluster of high-authority outlets with long publication histories. (Muck Rack, What Is AI Reading?, 2025) The outlets that trained AI systems to recognise credibility are the same outlets that were historically accessible only to companies with significant PR infrastructure. Small and mid-sized businesses have always been systemically underrepresented in major national media. That underrepresentation was baked into training data.


    The Philosophical Reframe

    Mark Coeckelbergh (2025) draws on Dotson’s (2014) concept of epistemic oppression — “a persistent and unwarranted infringement on the ability to utilize persuasively shared epistemic resources that hinder one’s contribution to knowledge production” — and extends it to AI-mediated knowledge environments. The argument is that AI does not merely repeat existing power asymmetries. It embeds them structurally.

    The relevant extension for brands is not just about producing knowledge. It is about the knowledge environments where customers form beliefs. If an AI system has no training-data basis to surface a brand in the answers your potential customers receive, that brand is excluded from the epistemic environment where purchase decisions begin. A customer who asks an AI assistant “what are the best firms for X?” receives an answer shaped entirely by what training data recognised as authoritative before any query was submitted. If your brand was not visible to the training data, you are absent from the answer — not because you lack capability, but because the epistemic infrastructure never recorded it. (Coeckelbergh, Social Epistemology 39(1), 2025)

    This is a consumption-side exclusion, not just a production-side one. The SME is not only excluded from contributing knowledge; it is excluded from the environments where knowledge shapes customer belief.

    Framing the SME AI visibility problem as a competitive gap misses this point. The gap did not emerge because large competitors worked harder or invested more in the last two years. It emerged because AI training data weighted types of presence — Wikipedia coverage, academic citations, major-media editorial — that mid-sized businesses have never had the infrastructure to accumulate. That is not a level playing field with a laggard on one side. That is a structural condition. And structural conditions require structural responses.


    What EC Work Actually Is

    The standard commercial framing for AI visibility tools — GEO vendors, entity optimisation platforms, AI citation monitoring services — presents the work as a competitive instrument. Get cited before your competitors do.

    That framing is not false. But it understates what the work is.

    Earned media in credible outlets, consistently maintained across a publication cadence, does two things simultaneously. In the retrieval layer, it provides fresh evidence for AI systems to draw on. In the training layer — for future model updates — it begins to build the kind of cross-source corroboration that pre-training data collection recognises as authoritative. Structured entity data (schema markup, sameAs identifiers, Wikidata records) creates the machine-readable signals that allow AI systems to resolve which firm you are and connect disparate mentions into a coherent entity record. Original research and defined analytical frameworks give AI systems something causally grounded to cite — not just statistical pattern-matching from general web content.

    Each of these is not just a tactic for climbing citation rankings. Each is a mechanism for entering the epistemic infrastructure from which training data bias has structurally excluded mid-sized brands.

    The EC toolset, properly understood, is not a route to competitive advantage. It is a route to epistemic access — the ability to participate in the knowledge environments where your potential customers form beliefs. Large brands already have that access, because their historical presence built it for them. Mid-sized brands are building it now, with tools that did not exist when the training data was assembled.

    That distinction matters for how you explain the work, how you measure its value, and how you evaluate whether a GEO vendor is offering you a tactical campaign or a structural solution.


    The Limit of the Argument

    This framing should not be stretched beyond its evidence base. Coeckelbergh’s epistemic justice concept is a normative philosophical framework, not an empirical claim about AI systems. The pre-training bias data from Chen et al. is robust, but it covers a single model (GPT-4o) and a narrow category (automotive brands). The training data composition claims are grounded in observed citation patterns, not disclosed training corpus analyses.

    What can be said with confidence: AI training data demonstrably weighted large-media presence, reference database coverage, and academic documentation in ways that structurally disadvantaged brands without those resources. That weighting was not deliberate exclusion — it was an artefact of using the web as training material. But the effect is structural regardless of intent. And the practical response — earned media, entity data, original research — addresses it at the level it operates: the evidence base from which AI systems draw their knowledge.


    The Close

    Generic GEO vendors offer tactics. Citation audits. Source gap analysis. Content briefs for the current citation map.

    The problem with that framing is that the citation map moves — 120% average source volatility in three months in late 2025 (Semrush AI Visibility Index, Geaney, Oct 2025) — while the underlying epistemic exclusion does not. Chasing this quarter’s top-cited sources does not change the structural fact that a brand with no pre-training signal is operating from a deficit that tactical content alone cannot close.

    The bias is not in the algorithm. It is in the data that trained it. And data — accumulated evidence, earned recognition, consistent cross-source corroboration — can be changed. It just requires understanding what it is you are actually trying to change.


    Sources: Coeckelbergh, Social Epistemology 39(1), 2025; Chen et al., arXiv:2601.16858 (University of Toronto, January 2026); Semrush AI visibility trend update, October 2025; earned media AI citation analysis, multiple sources.

    Article: #0049
    Date: 2026-07-16

  • E-E-A-T Is Entity Confidence

    E-E-A-T Is Entity Confidence — So Why Don’t We Say That?

    Summary: A provocation piece arguing that Google’s E-E-A-T framework and the practitioner concept of “entity confidence” describe identical phenomena — and examining what is genuinely new about the AI measurement layer.

    Last updated: 2026-07-07

    Type: Original article / insight post


    The GEO industry has spent two years building new vocabulary. Entity confidence. Citation frequency rate. AI visibility scores. Semantic authority. Confidence language analysis. These terms fill product decks, vendor websites, and optimisation guides published from 2024 onwards.

    Google invented most of this framework in 2014 and called it something different.

    But there is a catch — and it matters more every month. E-E-A-T was designed for one search engine. Half the search behaviour it was built for is now happening somewhere else entirely.


    What E-E-A-T Measures

    E-E-A-T — Experience, Expertise, Authoritativeness, Trustworthiness — is Google’s formal quality framework. Google’s HJ Kim described it as “a template they use to rate every single site for every single query.” Marie Haynes distilled the core of it more precisely: “E-E-A-T is a measure of the legitimacy of your entity as a destination for the topics you cover.”

    The measurement mechanism was described by Gary Illyes at Pubcon 2018: “E-A-T is largely based on links and mentions on authoritative sites. If the Washington Post mentions you, that’s good.”

    Danny Sullivan made the ranking connection explicit: E-E-A-T is not a directly disclosed score but a cluster of proxy signals — third-party mentions, backlinks, review reputation, entity recognition — that approximate what human quality raters would assess.

    And the loop closes formally through Pandu Nayak’s antitrust testimony: quality rater assessments generate the Information Satisfaction (IS) Score, and IS-scored documents train the deep learning systems that power Google Search. E-E-A-T signals → rater feedback → IS Score → ranking model training. It is not theoretical; it is a documented production mechanism.


    What Entity Confidence Measures

    Entity confidence score — as described by the vendors building AI citation measurement tools — is a composite measure of how certain an AI model is about the accuracy and authority of information associated with a brand. Primary vendors measure it through: citation frequency (how often AI mentions a brand for relevant queries), confidence language (whether AI uses “according to Brand X” versus “some sources suggest”), and response position (whether the brand appears first or sixth in an AI answer).

    The signals that build a high entity confidence score: authoritative off-site mentions, structured entity data (schema, sameAs markup), cross-platform consistency, third-party recognition from credible sources.


    The Mapping

    Hold the two frameworks side by side:

    E-E-A-T ComponentEntity Confidence Equivalent
    Authoritativeness — cited by credible third partiesRelationship Mapping — industry ecosystem recognition
    Expertise — topic depth, credentials, definitional claritySemantic Authority — topic clustering, definitional precision
    Experience — first-hand, original knowledgeInformation gain, non-AI-replicable original content
    Trustworthiness — transparent, consistent, no manipulationContextual Consistency — coherent, verified, cross-platform aligned

    The signal overlap is not approximate — it is complete. E-E-A-T is built from off-site mentions, links, entity recognition, review reputation, schema, and community presence. Entity confidence is built from the same list.

    The data confirms this at the empirical level. Ahrefs analysed 75,000 brands against AI Overview citation outcomes. The strongest correlate with AI citation was branded web mentions (off-site) at a correlation coefficient of 0.664. Not schema. Not content length. Not author bios. Earned media mentions — the Gary Illyes signal, showing up again in 2025 data for an entirely different platform.

    The finding replicates across independent datasets. 82% of AI citations across 1M+ analysed prompts came from earned media (Muck Rack/MacroLingo). And 76% of AI Overview citations came from pages already in the top-10 search results — meaning AI answers are substantially inheriting Google’s E-E-A-T rankings rather than computing something new from scratch.


    Did AI Engines Inherit E-E-A-T — or Reinvent It?

    The most important question the data raises is whether AI citation systems independently arrived at the same signal set as Google, or simply inherited it.

    For Google’s own AI surface, the evidence points clearly to inheritance. In the Ahrefs 1.4M-prompt study, 88% of ChatGPT citations came from the search channel — pages indexed and ranked in traditional web search. For Google AI Overviews and AI Mode, this is even more direct: those systems are grounded against Google’s index by design. AI Overviews are, to a significant degree, Google ranking with a synthesised output layer.

    The chain is documented end to end for this surface:

    E-E-A-T signals → off-site mentions and links → IS Score trains ranking models → pages rank in search → AI retrieval pulls from ranked pages → AI citation emerges

    There is no layer in that chain where AI engines are separately deciding what constitutes a trustworthy source. They are, in large part, delegating that judgment to Google’s ranking systems — which were themselves trained on E-E-A-T rater feedback.

    But this chain only describes part of what is now happening.


    The Search That Google Doesn’t See

    In June 2025, OpenAI published its first major usage research. The finding that matters here: 51.6% of all ChatGPT interactions are now search-like information queries — the primary use case for the platform, having overtaken content generation in a single year. ChatGPT has 700 million weekly active users. At 11–12% monthly growth from early 2025, it passed the mass adoption threshold around April 2026.

    These are not Google searches with a different interface. They are queries going directly to an AI engine that is not Google — and they are not passing through Google’s ranking systems before generating a response.

    The University of Toronto’s peer-reviewed study of 1,000+ queries (Chen et al., arXiv:2601.16858, January 2026) makes the structural difference measurable. Domain-level overlap between GPT-4o’s citations and Google’s top-10 results: 4%. Claude’s overlap: 12.6%. Perplexity’s: 15.2%.

    GPT-4o is making its own authority decisions on 96% of the domains it cites. It is not routing those decisions through Google’s ranking systems. For that 96%, E-E-A-T → search ranking → AI citation is not the chain. The AI is drawing on training data, its own retrieval logic, and entity signals that exist independently of Google’s index.

    The pre-training bias finding from the same study sharpens this further. For well-known brands, AI rankings remain highly stable even when retrieved supporting evidence is removed or shuffled entirely — the model’s pre-trained understanding of brand authority dominates regardless of what is currently ranking on Google. For those queries, Google’s ranking is not the input. Training data is. And training data is shaped by years of accumulated web content, entity signals, and third-party coverage — not by last week’s ranking positions.

    Zero-click search has been driving this shift at the structural level: approaching 65–70% of queries now resolve in AI-generated answers without a click-through. Google AI Mode runs at approximately 93% no-click. The search that is happening is increasingly not search in the traditional sense — it is AI-mediated information retrieval across multiple platforms, only some of which route through Google.


    Then What Is Entity Confidence Actually Adding?

    If E-E-A-T is entity confidence, the concept is not new. But what it covers — and how it is measured — genuinely is.

    E-E-A-T was built to track brand authority in one system: Google Search. The signals are right. The gap is the scope.

    Entity confidence, properly defined, covers two surfaces that E-E-A-T tracks only partially:

    The Google-adjacent surface: AI Overviews, Google AI Mode, and AI systems that retrieve primarily through search indices. Here, E-E-A-T signals are the right inputs and search ranking is a reasonable proxy for AI citation eligibility. This is the surface where the inheritance chain holds.

    The direct AI surface: ChatGPT queries without web search enabled, Claude, Perplexity, and all AI interactions where the model answers from training knowledge or its own retrieval logic rather than Google’s index. Here, E-E-A-T signals still matter — the underlying signals (earned media, entity clarity, cross-platform consistency) are what train these models and what their retrieval systems learn to trust. But Google ranking is not a reliable proxy for AI citation on this surface. A brand can hold Google position 3 and have no meaningful AI presence if its entity signals are thin in the training corpus.

    This is the surface that is growing fastest. The 51.6% of ChatGPT interactions that are search-like queries is not a static number — it is a rising proportion of a user base that grew 75% in five months.

    The second genuine addition is the measurement layer. E-E-A-T has no disclosed score. It is tracked by proxy: rank improvement, traffic, domain authority. These are position-based, deterministic metrics.

    AI citation is stochastic. Only 30% of brands maintain consistent visibility across multiple regenerations of the same query. A brand is not in AI position 3. It appears in approximately 65 out of every 100 relevant AI responses. That is a probability distribution, not a rank — and it can differ significantly across AI platforms using different retrieval architectures and training histories.

    The third addition is the narrative framing layer. Google returns a ranked list. AI platforms synthesise a recommendation and present it as a direct answer. There is a meaningful difference between being described as “the industry standard” and “a budget-friendly alternative” — both are appearances, but at very different confidence levels. That qualifier is invisible in any rank tracking tool.


    What This Means for Strategy

    Brands being sold AI visibility programmes as something new and distinct from their existing SEO and PR investment should ask: what specifically are we being asked to do that we are not already doing?

    For the Google-adjacent AI surface, the honest answer is: at the foundational level, very little. The inputs are shared:

    • Build authoritative off-site mentions → E-E-A-T Authoritativeness = EC Relationship Mapping = earned media programme
    • Demonstrate expertise in original content → E-E-A-T Expertise/Experience = EC Semantic Authority = owned content programme
    • Maintain entity clarity via schema and sameAs markup → E-E-A-T entity signals = EC entity disambiguation = technical SEO
    • Build consistent brand presence across platforms → E-E-A-T Trustworthiness = EC Contextual Consistency = brand management

    For the direct AI surface — the faster-growing, non-Google half — the inputs are still the same signals. But the tracking is not. Google rankings tell you nothing about how Claude or ChatGPT without web search represents your brand. Measuring that requires different instruments: repeated sampling across regenerated prompts, platform-specific citation profiling, confidence language analysis. This is what EC measurement adds that rank tracking cannot.

    The framing shift matters too. Google Search is where you compete for rank position. Direct AI search is where AI decides whether to recommend you based on what it has learned about your brand — and that learning happened before the user typed a word.


    The Vocabulary Question

    Why doesn’t the entity confidence industry say “we’re measuring E-E-A-T for AI engines”?

    The uncharitable read: novelty sells. New vocabulary justifies new products and new budget lines.

    The more generous read — and the more accurate one — is that E-E-A-T is a Google framework, and the problem is now multi-platform in a way it was not when the GEO vocabulary was being built. Saying “build your E-E-A-T” implies Google Search is the destination. Saying “build your entity confidence” points to the full citation surface: Google AI Overviews, ChatGPT, Perplexity, Claude, Gemini — including the surfaces that have no relationship to Google’s ranking systems.

    That vocabulary change reflects a real scope change. The underlying signals — what you need to do to build brand authority — are E-E-A-T. The measurement layer, the platform coverage, and the probabilistic framing of visibility scores are new. And the urgency is new: a search behaviour that took Google a decade to accumulate is now being replicated across multiple AI platforms at mass-adoption pace.

    The question worth asking is not whether your E-E-A-T programme is doing its job. It is whether you know what your brand looks like on the surfaces that E-E-A-T tracking was never designed to see.


  • Why ChatGPT Still Needs Google

    Why ChatGPT Still Needs Google (And What That Means for Brand Strategy)

    Published: 2026-07-02
    Status: Active


    One of the more persistent myths in AI search is that ChatGPT and Google are in competition — two separate systems fighting for the same users, pulling from separate pools of content. The practical version of this belief shows up in marketing strategy discussions: if AI search is taking over, does SEO still matter? Should budgets shift away from search optimisation and toward something else?

    The answer is no, and the reason is structural. ChatGPT is not independent of Google’s infrastructure. It is, to a significant degree, built on top of it.


    The 88% Finding

    Ahrefs analysed 1.4 million ChatGPT prompts and traced where the cited pages actually came from. The search channel — content retrieved via the web search index — accounts for 88% of all ChatGPT citations. News accounts for most of the rest. Reddit, YouTube, and academic sources together contribute less than 3%.

    This is not a peripheral finding. It is the foundational architecture of how ChatGPT browses. When a user asks ChatGPT a question that triggers web retrieval, the system queries a search index, assembles candidate URLs, filters them for relevance, reads the most promising pages, and cites selected content. The index it queries is built substantially on the same crawl infrastructure that powers web search.

    A brand not present in that index is not retrievable. It does not matter how well the brand’s content is structured, how authoritative its coverage, or how precisely its messaging matches a query. If the page is not indexed, it is invisible to the retrieval pipeline.

    The operational implication: making a page indexable and crawlable by search engines is not SEO housekeeping. It is the minimum requirement for AI citation eligibility.


    Indexed, Not Necessarily Ranked

    The 88% figure is sometimes misread. It does not mean that ChatGPT cites pages because they rank well in Google. The relationship is looser than that.

    Research across multiple sources finds that 80% of LLM citations do not rank in Google’s top 100 for the specific query being answered. A separate study found that 28.3% of the most-cited ChatGPT pages rank nowhere visible in Google at all for the relevant query. These pages are indexed — they are in the database — but they are not necessarily winning on traditional search ranking signals for the particular question being asked.

    The distinction matters. What ChatGPT needs from Google’s infrastructure is discovery and access: a mechanism to find pages and retrieve their content. It does not need those pages to have won Google’s ranking competition. A page that sits at position 47 for a given query, or that ranks well for related queries but not this exact one, can still be retrieved and cited by ChatGPT if it passes the semantic relevance filters.

    The threshold is being indexed, not being ranked. But the two are related in practice. Pages that are not indexed are invisible. Pages with severe technical issues — slow load times, blocked crawl paths, thin content signals that discourage indexing — are poorly represented even when nominally indexed. And pages that rank well for a topic tend to be well-indexed, freshly crawled, and more likely to appear in AI retrieval sets.

    The correct frame is not “rank #1 or be invisible to AI” — it is “be a legitimate, accessible, well-structured presence in the web index, and AI can reach you.”


    Google AI Overviews: An Even Tighter Relationship

    The dependency between AI citation and search performance is strongest, not weakest, in Google’s own AI product.

    76% of AI Overview citations come from pages already ranking in Google’s top 10. For Google’s own AI-generated answers, traditional search ranking is not a rough correlation — it is the dominant selection mechanism. The E-E-A-T signals that determine page ranking are directly inherited by AI Overview citation selection. Branded web mentions correlate with AI Overview citation at 0.664 — the strongest single signal measured in Ahrefs’ original research.

    This is architecturally unsurprising. Google’s AI Overviews are powered by the same quality signals that power its search rankings because they are drawing from the same evaluated content pool. AI Overviews are not a separate system that happens to overlap with Search — they are an interface layer built on top of Search’s infrastructure.

    For brands targeting visibility in Google’s AI products specifically, the most direct path is also the most obvious one: rank well in Google Search.


    Where the Dependency Breaks Down

    The search-channel dependency is not uniform. Two scenarios produce exceptions to the 88% pattern.

    The training-mode exception. When a user asks ChatGPT a question it can answer from pre-trained knowledge — without triggering any retrieval — the search channel plays no role. The model draws from parametric knowledge built during training, shaped by the web’s cumulative discussion of a topic. For well-known brands with strong training-data presence, a significant share of AI mentions may be training-mode responses that retrieve nothing and cite nothing.

    This is not an opportunity to avoid search investment. It is a separate channel (documented in detail elsewhere), and it is largely inaccessible to short-cycle optimisation — it is determined by years of web presence, not recent content. For the niche and mid-market brands that make up most commercial AI search competition, the pre-training channel is weak by definition. The AI has limited knowledge of them at training time, so it relies on retrieval — and retrieval runs through the search index.

    The Perplexity exception. Perplexity’s citation pool overlaps with ChatGPT’s by less than 1%. Its architecture weights Reddit more heavily (46.7% of top Perplexity citations), favours freshness aggressively (content deprioritised after 2–3 days), and draws from a different mix of authoritative sources. Optimising for ChatGPT citation does not automatically produce Perplexity visibility.

    This matters for brands choosing where to invest. ChatGPT and Perplexity are not interchangeable targets, and a strategy built entirely around the search-channel dependency will be less effective for Perplexity than for ChatGPT or Google AI Overviews.


    What This Means for SEO Investment

    The 88% finding does not mean SEO and AI search are identical — it means SEO is the infrastructure layer on which most AI citation depends. Three practical conclusions follow from this.

    Search indexability is non-negotiable. Any brand whose pages have crawl issues, thin-content penalties, or poor technical foundations is not just underperforming in search — it is actively reducing its AI citation eligibility. Technical SEO is not a legacy practice in an AI-first world. It is the precondition for entering AI retrieval pipelines.

    Search ranking improves AI citation probability, but the relationship is not linear. A page that moves from position 12 to position 3 in Google Search will be more likely to appear in AI retrieval sets — not because AI citation tracks ranking directly, but because high-ranking pages are fresher, better crawled, and more likely to pass AI semantic relevance filters. The investment that improves ranking also tends to improve citation eligibility, through overlapping mechanisms.

    AI citation requires an additional layer that search ranking alone does not provide. Getting indexed and ranked makes a page reachable. What determines whether it is actually cited is a second filter: semantic relevance to the AI’s internal sub-questions, content structure (direct answers, clear headings, attributed data), and source credibility (authoritative third-party coverage far outperforms brand-owned content). These are not ranking factors in the traditional sense — they are citation selection factors that apply after the retrieval pipeline has already found the page.

    The sequence is: be indexed → be retrievable → be selected for citation. Search investment covers the first two stages. AI-specific content and authority work covers the third.


    The Strategic Error to Avoid

    The error is treating AI search and traditional search as substitutes — allocating budget to “AI SEO” while reducing investment in the infrastructure that makes AI citation possible in the first place.

    ChatGPT’s 88% search-channel dependency is not a transitional feature that will disappear as AI search matures. It reflects a structural reality: AI retrieval systems need curated, authority-evaluated, freshness-maintained content databases to draw from, and the web’s search index is the most comprehensive one that exists. Google built that index over two decades. AI systems are using it because there is no better alternative.

    The brands that will be cited reliably in AI answers are the ones that are indexed, crawlable, authoritative, and structured for extraction — then additionally covered by the earned media and community presence that AI systems use to evaluate authority. The first set of requirements is not new. They are what good SEO has always required.

    The second set of requirements is the genuine addition. But it sits on top of the first set, not beside it.


    Sources: Ahrefs 1.4M ChatGPT prompt study (Why ChatGPT Cites One Page Over Another); Ahrefs AI Overview brand correlation research; ConvertMate GEO Benchmark Study 2026 (12,500+ queries, 8,000 domains); Chen et al. arXiv:2601.16858 (University of Toronto, January 2026); BrandFeatured AI ranking factors cross-platform divergence data.

  • The Role of Third-Party Validation in AI Recommendations

    In the world of AI recommendations, what others say about you carries far more weight than what you say about yourself. This principle shapes how AI systems evaluate businesses and explains why some companies with modest marketing efforts outperform heavily promoted competitors.

    The Credibility Asymmetry

    Consider how you personally evaluate claims. If a company’s website says ‘We provide exceptional service,’ you might note it but remain appropriately sceptical. If an independent review says ‘Their service was exceptional—they went above and beyond,’ you weight it differently. If a trade publication names them ‘Service Provider of the Year,’ you pay serious attention.

    AI systems have learned this same asymmetry. They’ve been trained on vast amounts of text that includes both marketing language and genuine third-party assessments. They’ve learned to distinguish between claims a business makes about itself and validation that comes from independent sources—and to weight the latter more heavily.

    The Spectrum of Validation

    Third-party validation exists on a spectrum of credibility. At one end are casual mentions—a social media post praising your business, a forum comment recommending you. These help, but carry limited weight. Further along are customer reviews on established platforms, where verification processes lend credibility. Further still are media mentions, industry awards, professional accreditations, and academic or expert citations.

    AI systems appear to calibrate the weight they give different validation types. A mention in a respected industry publication signals more than a positive review, which signals more than a casual social mention. The cumulative effect of multiple validation types across the spectrum creates the strongest foundation for confident recommendations.

    Reviews: Quantity, Quality, and Diversity

    Customer reviews deserve particular attention because they’re the most common form of validation and the most accessible for businesses to influence. But not all review presence is equal. AI systems appear to consider quantity, quality, recency, and diversity.

    Quantity matters because a single glowing review could be an anomaly or a planted endorsement, while consistent positive reviews over time suggest genuine customer satisfaction. Quality matters because detailed, specific reviews demonstrate authentic experience while generic praise may be discounted. Recency matters because recent reviews confirm current service quality. And diversity—reviews across multiple platforms—matters because it’s harder to manipulate and suggests broader customer engagement.

    The Earned Media Advantage

    Beyond reviews, earned media coverage represents particularly valuable validation. When a publication chooses to write about your business, interview your leadership, or feature your work, it implies editorial judgement about your relevance and credibility. You can’t buy genuine editorial coverage; it must be earned through newsworthiness, expertise, or excellence.

    This is why businesses with media presence often outperform in AI recommendations despite potentially having less optimised websites. The AI recognises that independent journalists and editors have already done verification work.

    Building Validation Systematically

    The good news is that third-party validation, while earned rather than purchased, can be cultivated strategically. Businesses can encourage satisfied customers to leave reviews. They can pursue relevant professional accreditations. They can develop genuine thought leadership that attracts media attention. They can participate in industry activities that generate mentions and recognition.

    The key is understanding that this validation ecosystem matters for AI visibility and deserves the same strategic attention that businesses have historically given to their website or advertising.

    Third-party validation can’t be manufactured artificially, but it can be developed deliberately. Understanding which forms of validation your business most lacks—and which would have the greatest impact—enables focused effort with meaningful returns.


    Next read – ‘How it works’ right here to see how you can benefit.

  • How AI Systems Verify Whether Your Business Is Real

    Before an AI will recommend your business, it must first be confident that your business actually exists as a legitimate, operational entity. This verification process happens automatically, drawing on patterns the AI has learned from analysing millions of businesses—both genuine and fraudulent.

    The Problem AI Systems Are Solving

    The internet is full of fake businesses. Shell companies created for fraud, abandoned websites for defunct operations, placeholder pages that never became real businesses, and deliberately misleading listings designed to capture traffic or payments. AI systems must distinguish genuine businesses from this noise to provide useful recommendations.

    This isn’t a hypothetical concern—it’s a practical necessity. If an AI recommended fake or defunct businesses, users would quickly lose trust in its suggestions. The systems have therefore developed sophisticated verification heuristics.

    Cross-Platform Consistency

    One key verification signal is consistency across platforms. A legitimate business typically has a presence across multiple authoritative platforms—a website, Google Business Profile, Companies House registration (for UK companies), professional directory listings, social media profiles, and so forth. When these sources align in their basic details—company name, address, contact information, nature of business—the AI gains confidence that it’s dealing with a real entity.

    Fake businesses struggle to maintain this consistency. They might have a website but no verifiable registration. Their address might not match any actual location. Their phone number might be disconnected or lead somewhere unexpected. Each inconsistency raises doubt.

    Evidence of Activity Over Time

    Legitimate businesses leave traces of activity over time. They accumulate reviews. They appear in dated news articles or blog posts. Their websites show evidence of updates. Their social media has a history. This temporal dimension helps distinguish established businesses from recently created facades.

    AI systems are particularly attentive to this for newer businesses. A company that appears to have sprung into existence fully formed, with no discernible history, triggers caution. A company with clear evidence of operating over months or years, accumulating the normal digital artefacts of a real business, passes verification more readily.

    Human Verification Signals

    Real businesses are run by real people, and AI systems look for evidence of those connections. Do the business’s principals have credible professional profiles? Are they associated with other legitimate entities? Do their claimed qualifications appear verifiable? Does anyone mention them in professional contexts?

    This personal dimension of verification explains why businesses with visible, credentialed leadership often perform better in AI recommendations. The humans behind the business provide another layer of authentication that purely anonymous businesses cannot offer.

    The Implications for Legitimate Businesses

    Understanding this verification process matters because legitimate businesses sometimes inadvertently fail it. They might have inconsistent information across platforms because nobody has audited it. They might lack the temporal footprint because they’ve been operating primarily offline. Their principals might have minimal digital presence despite substantial real-world credentials.

    These verification gaps don’t necessarily trigger outright rejection—the AI doesn’t think you’re fake—but they do reduce confidence. And reduced confidence translates to less frequent or less emphatic recommendations. Ensuring your business clearly passes verification isn’t about proving something contentious; it’s about making the obvious easily discoverable.

    Verification is the foundation of AI visibility—without it, nothing else matters. Yet many businesses have never assessed how clearly they demonstrate their legitimacy to systems that cannot assume good faith.


    Next read – ‘How it works’ right here to see how you can benefit.

  • The Difference Between Being Found and Being Recommended

    There’s a crucial distinction that many businesses overlook: appearing in search results is not the same as being recommended. This difference is becoming increasingly important as AI transforms how customers discover and choose businesses.

    The Old Model: Lists and Rankings

    Traditional search engines present users with lists. Type in a query, receive ten blue links, scroll through, click on a few, compare, decide. The search engine’s job is to return relevant results and rank them by some combination of relevance and authority. The user does the work of evaluating and choosing.

    In this model, being found means appearing on that list—ideally near the top. Success is measured by rankings and click-through rates. The search engine is essentially a librarian pointing you toward the right section of the library; you still have to read the books and make up your own mind.

    The New Model: Curated Answers

    AI search works fundamentally differently. When someone asks an AI assistant for a recommendation, they’re not expecting a list to investigate. They’re expecting the AI to have already done that investigation and to provide a direct answer. ‘Which law firm should I use for commercial property?’ expects a name, perhaps with reasoning, not a reading list.

    This means the AI isn’t just finding businesses that match the query—it’s making qualitative judgements about which businesses deserve to be specifically mentioned. It’s acting less like a librarian and more like a knowledgeable friend who happens to know the answer.

    The Implications of Being Recommended

    Being recommended carries a form of endorsement that being listed never did. When an AI suggests your business specifically, users naturally assign weight to that recommendation. They may not investigate alternatives at all. The competitive dynamics shift dramatically when you move from ‘one option among many’ to ‘the suggested answer.’

    But this also raises the bar considerably. AI systems are cautious about making explicit recommendations because their credibility depends on those recommendations being sound. They want substantial evidence before they’ll confidently suggest a specific business. Being merely findable isn’t enough; you need to be convincingly recommendable.

    The Confidence Threshold

    Think of AI systems as having a confidence threshold for recommendations. Below that threshold, they might mention your business as one possibility among several, or list you in a category, or say they’re not sure who to recommend. Above that threshold, they’ll specifically suggest you as the answer to the user’s question.

    What drives confidence above that threshold? The cumulative weight of positive signals: strong independent validation, consistent information, depth of expertise demonstrated, relevance to the specific query, and absence of concerning negatives. Businesses hovering just below the threshold may appear occasionally or in certain contexts, but those above it capture a disproportionate share of AI-driven recommendations.

    A Different Kind of Competition

    This creates a different competitive landscape. In traditional search, you competed for rankings against everyone in your category. In AI recommendations, you’re competing to be the business the AI feels most confident suggesting. This is often a smaller, more intense competition—but one with higher rewards for those who succeed.

    Understanding where you stand in this confidence hierarchy—and what factors are holding you back from more frequent, more confident recommendations—becomes essential for businesses that want to capture this emerging channel of customer acquisition.

    The shift from being found to being recommended represents a fundamental change in digital visibility. Businesses that understand this distinction and position themselves accordingly will capture opportunities that their competitors don’t even realise exist.


    Next read – ‘How it works’ right here to see how you can benefit.

  • Why Your Website Alone Won’t Get You Recommended by AI

    If you’ve invested time and money into your website, you might reasonably expect it to be your primary asset for attracting customers through AI search. After all, it’s where you control your message completely. Unfortunately, when it comes to AI recommendations, your website is necessary but far from sufficient.

    The Fundamental Difference

    Traditional search engines work primarily by crawling websites and ranking them based on various signals including content relevance, site structure, loading speed, and inbound links. Your website is the primary object being evaluated. AI systems take a fundamentally different approach: they try to understand the business itself, using the website as just one source of evidence among many.

    Think of it this way: if you were personally recommending a business to a friend, you wouldn’t base your recommendation solely on how impressive their website looked. You’d consider what you’d heard from other people, whether they had a good reputation, how long they’d been operating, whether they’d been mentioned in credible publications, and whether their claims seemed substantiated. AI systems attempt something similar at scale.

    The Corroboration Problem

    Any business can claim anything on its own website. ‘Award-winning service,’ ‘industry-leading expertise,’ ‘trusted by hundreds of satisfied clients’—these phrases appear on countless sites. AI systems have learned to be appropriately sceptical of self-proclaimed excellence. What they look for is corroboration: do independent sources confirm what the business claims about itself?

    This corroboration can take many forms. Customer reviews on third-party platforms provide evidence of service quality. Mentions in trade publications suggest industry recognition. Listings in professional directories confirm legitimate operation. Social media engagement demonstrates an active, responsive business. Each of these external signals helps AI systems calibrate how much weight to give your own claims.

    A business with a beautiful website but no external validation faces a credibility gap. The AI has only one source of information—the source with the most obvious incentive to present things favourably—and must discount accordingly.

    The Consistency Imperative

    Beyond corroboration, AI systems look for consistency. Does the information on your website match what appears elsewhere? Are your contact details, service descriptions, and business information uniform across all platforms where you appear? Inconsistencies create uncertainty, and uncertainty leads AI systems to hedge their recommendations or favour competitors with cleaner, more consistent information.

    Many businesses inadvertently create these inconsistencies over time. An old directory listing shows a previous address. A review platform has an outdated phone number. LinkedIn describes services differently from the website. Each discrepancy, however minor, chips away at the AI’s confidence in its understanding of your business.

    The Holistic View

    What AI systems are really attempting is to build a holistic, verified understanding of your business as an entity that exists in the world. Your website contributes to this understanding, but it cannot single-handedly establish it. The businesses that get recommended most readily are those with rich, consistent, externally validated digital footprints that extend far beyond their own domains.

    This doesn’t mean your website doesn’t matter—it absolutely does. But it means that website optimisation alone is an incomplete strategy for AI visibility. The broader ecosystem of information about your business requires equal attention.

    Most businesses have never audited their complete digital footprint or assessed how they appear across the diverse sources that AI systems consult. Understanding this full picture is essential for any meaningful improvement strategy.


    Next read – ‘How it works’ right here to see how you can benefit.