Series: Confident and Wrong — Article 3 of 7
Intro:
In the first seven weeks of 2026, one in every 277 academic papers published contained a fabricated citation — a reference to a source that does not exist, generated by an AI system and presented as fact.
Two years earlier, the rate was one in 2,828.
That is not an improvement story. It is a sixfold increase. And it is accelerating.
Confident and Wrong — Article 3 of 7.
To read the whole series from the start click here
The Lancet Data: Scale and Trend
The finding comes from a peer-reviewed study published in The Lancet in May 2026, led by Maxim Topaz at Columbia University. The research team analysed more than 2 million papers containing 97 million citations. What they found was systematic, not incidental.
The numbers in sequence:
– 2023: 1 in 2,828 papers contained fabricated citations
– 2025: 1 in 458 papers
– Early 2026 (first seven weeks of the year): 1 in 277 papers
Across the full dataset: approximately 4,000 fabricated citations identified, spread across roughly 2,800 papers.
The trend line is steeper than the current rate implies. From 2023 to 2025 alone, the rate increased sixfold. The early 2026 figure continues that trajectory — it does not level off, and there is no mechanism in the data that would cause it to.
Misha Teplitskiy at the University of Michigan, commenting on the study, was direct: “This is one of the first papers telling us something about the quality of what’s being produced with LLMs, and it’s a signal of slop.”
The distribution of fabricated citations is not uniform. More than one-third originate from two large open-access publishers. Science journals with strong editorial oversight report no published fabricated citations. Publishers with weaker oversight — PLOS among them — acknowledge “numerous” unverifiable references. The correlation is clear: fabrication risk is inversely proportional to editorial rigour. Where the system is designed to catch errors, fewer fabrications survive. Where it is not, they accumulate.
Why This Rate Is a Floor, Not a Ceiling
The Lancet study measures academic publishing. That context matters for understanding what the numbers actually represent.
Academic papers are the most adversarial environment for AI-generated fabrication. Authors have professional and reputational incentives to check their own references. Editorial systems at the better journals are specifically designed to identify errors. Peer reviewers are domain experts who may recognise a fabricated citation by name. And the stakes — careers, institutional credibility, funding — are high enough that the detection pressure is real.
None of those conditions apply to commercial brand representation.
When an AI system describes your company’s founding date, your service offering, your partnership history, or your executive team, there is no editorial system checking the claim. There is no peer review. The person receiving the information is unlikely to verify it against primary sources. And the model generating it is the same model generating academic citations — it has no separate “commercial accuracy” layer, no distinct mechanism for verifying facts about brands that it would not also apply to academic literature.
The academic fabrication rate is therefore a lower-bound estimate for commercial entity hallucination — not an upper bound. The conditions that cause fabrication exist across all AI output. The conditions that catch fabrication are largely absent outside of high-oversight publishing contexts.
If one in 277 papers contains a fabricated citation in a context designed to catch fabrication, the implied rate in contexts that are not designed to catch fabrication is not lower. The data does not support optimism in that direction.
The Corroborating Evidence From a Different Domain
The Lancet study is not alone. It confirms a pattern that the CJR/Tow Center study at Columbia Journalism Review identified from a different angle in 2025.
That study tested eight generative AI tools against 1,600 prompts drawn from news publishers — evaluating whether AI correctly identified articles, publishers, and URLs. The headline finding: more than 60% of responses across all eight tools contained incorrect information. Error rates ranged from 37% at best to 94% at worst.
Two studies, different domains, different methodologies, different research teams, different publication years. Both arrive at the same structural conclusion: AI citation and reference failure is not an edge case, and it is not confined to low-quality tools. It is a systemic property of how current AI systems generate text.
The Deloitte finding adds a third dimension. In 2024, 38% of business executives reported making incorrect business decisions based on hallucinated AI outputs. That number does not come from a lab — it comes from operational deployment, at scale, by sophisticated organisations that were presumably paying attention. One in three executives making a wrong call because of fabricated AI information is not a user-education problem. It is a system-level failure rate.
The Mechanism That Makes It Worse
The acceleration from 2023 to 2026 is not explained by a single cause. But the primary driver is structural, and it is self-reinforcing.
AI systems are trained on text corpora assembled from the open web, academic publishing databases, and other large-scale text sources. As AI-generated content proliferates — in academic papers, on websites, in news articles — it is absorbed into the corpora that future models train on. Fabricated citations in published papers become part of the training data. The model that trained on clean data in 2022 was not exposed to AI-generated academic slop. The model training in 2026 is.
This is not a theoretical risk. The Lancet data shows the rate accelerating at the pace you would expect if fabricated content were being absorbed as legitimate training signal. Each generation of models learns, in part, from the outputs of previous models. Each generation produces a slightly higher baseline fabrication rate. The loop compounds.
The academic publishing world is beginning to respond — stronger editorial checks, citation verification tooling, mandatory AI disclosure requirements. But those responses operate on the current margin, not the accumulated training signal. The models already trained on pre-2026 academic literature carry whatever was in that corpus. They cannot be retroactively cleaned.
Article 2 in this series established that AI is trained to sound certain. This article establishes the rate at which that confident output is factually wrong. The combination is the problem. Fabrication alone would be detectable — a system that generated obvious nonsense with visible hesitation would be identifiable as unreliable. Fabrication delivered with the linguistic confidence of an accurate statement is something different. It passes human scrutiny at the first reading. It propagates.
What This Means for How AI Represents Your Brand
AI systems that fabricate citations in peer-reviewed science are the same systems writing about your products, your executives, your founding date, and your competitive positioning. The failure mode is not domain-specific.
The underlying cause — confident text generation without evidential grounding — is a property of how these systems work, not of the subject matter they are applied to. A model that invents a citation because it does not have the source but needs to produce fluent, confident output will apply exactly the same process when it does not have accurate information about your brand but is asked to describe it. The mechanism is identical.
This means that “AI is often wrong about things” is not the frame. The frame is: AI fabrication is a measurable, accelerating, systemic problem with an identified rate, a documented trend, and a structural mechanism that is making it worse. The one-in-277 figure is not a worst-case scenario. In the context of academic publishing — with all its oversight and incentives — it is the measured current rate.
For brands, the question this data raises is not “could AI get our information wrong?” The data has answered that. The question is: what does the AI have to draw on when it is generating text about us? A model working from rich, consistent, independently corroborated entity data will still sometimes fabricate. But it is less likely to fabricate about the things it has strong evidence for — because the fabrication mechanism kicks in when evidence is absent or thin.
The structural response to a system that confidently generates wrong information is not to hope the system improves. The Lancet trend runs directly against that hope. The structural response is to ensure that the system has enough accurate evidence about your entity that there is less gap for fabrication to fill.
The Direction of Travel Is Not Reassuring
The “AI is improving” response to fabrication data is persistent, but the Lancet study directly contradicts it on the specific question that matters here.
Overall AI capability is increasing. Benchmarks are improving. Some hallucination rates on standardised evaluations have declined for some models in some contexts. None of that reverses the finding that fabricated citations in academic publishing increased sixfold from 2023 to 2025 and continued accelerating into 2026.
Capability and accuracy are not the same thing. The Premium Model Paradox — described in Article 1 of this series — established that higher-tier models can produce more confidently wrong outputs, not fewer. The Lancet data confirms the trajectory at the publishing level: as AI use grows, as AI-generated text enters training corpora, as the feedback loop compounds, the fabrication rate in the corpus goes up, not down.
The question for brands is whether they want to wait for the trend to reverse — with no evidence it will — or whether they want to do something about their own entity evidence now, while the gap between well-evidenced and poorly-evidenced brands is still a gap rather than a cliff.
The rate is one in 277. It was one in 2,828 two years ago. The direction is unambiguous.
The Next Article in This Series
If fabrication is accelerating and confidence is not a reliable signal of accuracy, the next question is whether the traditional defence — ranking well in Google — still offers any protection. Article 4 shows that it does not, and quantifies exactly how much ground has been lost.
*Sources: Topaz et al., “Fabricated Citations in Scientific Literature,” The Lancet (May 2026); CJR / Tow Center for Digital Journalism, “We Compared Eight AI Search Engines. They’re All Bad at Citing News.” (2025); Deloitte, AI in the Enterprise (2024).*

Leave a Reply