Why Mid-Sized Brands Are Locked Out of AI Knowledge

The Structural Exclusion Problem: Why Mid-Sized Brands Are Locked Out of AI Knowledge

Most GEO and EC advice frames the SME visibility problem as a competitive disadvantage. You are behind the large players. Here is how to close the gap.

That framing is wrong — or at least, it is not wrong enough. The problem is not that mid-sized brands are losing a race. It is that the race was designed without them.


The Training Data Problem Is Not Random

Before AI systems answer questions about your brand, your industry, or your category, they have already formed a view. That view was assembled during pre-training — the process by which a model learns the world’s knowledge from an enormous corpus of text, before it ever answers a query.

That corpus was not a neutral sample of what exists. It was a weighted sample of what had been digitally published, cited, covered in large-circulation media, and documented in the reference sources that trained models weight most heavily: Wikipedia, academic databases, major editorial outlets, trade press with decades of archive depth.

Large brands — the multinationals, the household names, the category incumbents — accumulated exactly these kinds of presence over decades. They had Wikipedia pages. They had Reuters coverage. They had academic case studies, analyst reports, Financial Times profiles. That presence existed before any training data was assembled. When the models trained on the web, they trained on a web that had already organised itself around the visible and the established.

Mid-sized SMEs, by design, had none of this. A regional services firm with thirty years of operational excellence but no analyst coverage and no Wikipedia entry had produced no signal that pre-training data collection would recognise as authoritative. The AI system did not decide that firm was unimportant. The training data never recorded its importance in the first place.

This is not market inefficiency. It is structural exclusion embedded in how knowledge was assembled.


What the Evidence Shows

The academic basis for this is not theoretical. Chen et al. (arXiv:2601.16858, January 2026) ran perturbation experiments on GPT-4o: they manipulated the evidence the model received about brands — shuffling retrieved snippets, restricting retrieval to only provided content, injecting brand names into unrelated material — and measured how much those manipulations moved the model’s output rankings.

For well-known, popular brands, the results were stark. Average rank deviation under snippet manipulation: 2.30–2.60. For niche entities: 4.15–4.63. Popular brand rankings barely shifted regardless of what the retrieved evidence said. The model already knew the answer. The citation miss rate for Cadillac was 58%. For Infiniti, 73%. Those brands appeared in AI answers without any supporting retrieved content more than half the time — drawn entirely from training priors. (Chen et al., arXiv:2601.16858)

The mechanism is clear. Popular entity rankings are governed by pre-trained knowledge. Retrieved evidence is used to confirm what the model already believes, not to discover who deserves to be recommended. For niche brands — where the model holds no stable prior — retrieval actually drives the answer. The two populations are not playing the same game.

The training-data corpus itself reflects this asymmetry. Analysis of earned media citation patterns shows that 82% of AI citations come from earned media sources — but those sources are heavily concentrated in a small cluster of high-authority outlets with long publication histories. (Muck Rack, What Is AI Reading?, 2025) The outlets that trained AI systems to recognise credibility are the same outlets that were historically accessible only to companies with significant PR infrastructure. Small and mid-sized businesses have always been systemically underrepresented in major national media. That underrepresentation was baked into training data.


The Philosophical Reframe

Mark Coeckelbergh (2025) draws on Dotson’s (2014) concept of epistemic oppression — “a persistent and unwarranted infringement on the ability to utilize persuasively shared epistemic resources that hinder one’s contribution to knowledge production” — and extends it to AI-mediated knowledge environments. The argument is that AI does not merely repeat existing power asymmetries. It embeds them structurally.

The relevant extension for brands is not just about producing knowledge. It is about the knowledge environments where customers form beliefs. If an AI system has no training-data basis to surface a brand in the answers your potential customers receive, that brand is excluded from the epistemic environment where purchase decisions begin. A customer who asks an AI assistant “what are the best firms for X?” receives an answer shaped entirely by what training data recognised as authoritative before any query was submitted. If your brand was not visible to the training data, you are absent from the answer — not because you lack capability, but because the epistemic infrastructure never recorded it. (Coeckelbergh, Social Epistemology 39(1), 2025)

This is a consumption-side exclusion, not just a production-side one. The SME is not only excluded from contributing knowledge; it is excluded from the environments where knowledge shapes customer belief.

Framing the SME AI visibility problem as a competitive gap misses this point. The gap did not emerge because large competitors worked harder or invested more in the last two years. It emerged because AI training data weighted types of presence — Wikipedia coverage, academic citations, major-media editorial — that mid-sized businesses have never had the infrastructure to accumulate. That is not a level playing field with a laggard on one side. That is a structural condition. And structural conditions require structural responses.


What EC Work Actually Is

The standard commercial framing for AI visibility tools — GEO vendors, entity optimisation platforms, AI citation monitoring services — presents the work as a competitive instrument. Get cited before your competitors do.

That framing is not false. But it understates what the work is.

Earned media in credible outlets, consistently maintained across a publication cadence, does two things simultaneously. In the retrieval layer, it provides fresh evidence for AI systems to draw on. In the training layer — for future model updates — it begins to build the kind of cross-source corroboration that pre-training data collection recognises as authoritative. Structured entity data (schema markup, sameAs identifiers, Wikidata records) creates the machine-readable signals that allow AI systems to resolve which firm you are and connect disparate mentions into a coherent entity record. Original research and defined analytical frameworks give AI systems something causally grounded to cite — not just statistical pattern-matching from general web content.

Each of these is not just a tactic for climbing citation rankings. Each is a mechanism for entering the epistemic infrastructure from which training data bias has structurally excluded mid-sized brands.

The EC toolset, properly understood, is not a route to competitive advantage. It is a route to epistemic access — the ability to participate in the knowledge environments where your potential customers form beliefs. Large brands already have that access, because their historical presence built it for them. Mid-sized brands are building it now, with tools that did not exist when the training data was assembled.

That distinction matters for how you explain the work, how you measure its value, and how you evaluate whether a GEO vendor is offering you a tactical campaign or a structural solution.


The Limit of the Argument

This framing should not be stretched beyond its evidence base. Coeckelbergh’s epistemic justice concept is a normative philosophical framework, not an empirical claim about AI systems. The pre-training bias data from Chen et al. is robust, but it covers a single model (GPT-4o) and a narrow category (automotive brands). The training data composition claims are grounded in observed citation patterns, not disclosed training corpus analyses.

What can be said with confidence: AI training data demonstrably weighted large-media presence, reference database coverage, and academic documentation in ways that structurally disadvantaged brands without those resources. That weighting was not deliberate exclusion — it was an artefact of using the web as training material. But the effect is structural regardless of intent. And the practical response — earned media, entity data, original research — addresses it at the level it operates: the evidence base from which AI systems draw their knowledge.


The Close

Generic GEO vendors offer tactics. Citation audits. Source gap analysis. Content briefs for the current citation map.

The problem with that framing is that the citation map moves — 120% average source volatility in three months in late 2025 (Semrush AI Visibility Index, Geaney, Oct 2025) — while the underlying epistemic exclusion does not. Chasing this quarter’s top-cited sources does not change the structural fact that a brand with no pre-training signal is operating from a deficit that tactical content alone cannot close.

The bias is not in the algorithm. It is in the data that trained it. And data — accumulated evidence, earned recognition, consistent cross-source corroboration — can be changed. It just requires understanding what it is you are actually trying to change.


Sources: Coeckelbergh, Social Epistemology 39(1), 2025; Chen et al., arXiv:2601.16858 (University of Toronto, January 2026); Semrush AI visibility trend update, October 2025; earned media AI citation analysis, multiple sources.

Article: #0049
Date: 2026-07-16

Comments

Leave a Reply

Your email address will not be published. Required fields are marked *