The Ghost Stat Epidemic

Welcome to the cesspit of content marketing.

I have been watching this for two years. Every time I need a statistic for a client piece, I follow the source. And almost every time, the chain ends the same way. A broken link. A PDF with no author. A blog post citing another blog post citing a report that does not contain the number being attributed to it.

The statistic is everywhere. The evidence for it is nowhere.

I kept thinking it would get better. Better tools, more scrutiny, rising standards. It has not got better. It has got worse. Because now we have AI generating content at a scale nobody anticipated, trained on an internet that was already full of numbers nobody checked.

So I am done being polite about it.

This is not a niche problem. It is not a small agency problem. I am talking about statistics being confidently repeated in content published by some of the biggest names in B2B marketing. Organisations with proper budgets, dedicated content teams and no excuse for not knowing better. The ghost stat epidemic is not driven by junior freelancers cutting corners. It is driven by a culture that has decided speed matters more than accuracy, at every level.

I have been in sales and marketing for over twenty years. I have run agencies for fifteen. I have briefed writers, reviewed copy and signed off content across professional services, financial services, technology and more industries than I can count. And I am telling you directly: this has to stop.

Here is the problem, who is responsible and what good actually looks like.

The pattern always looks the same

Someone publishes a stat without a source. Maybe they genuinely forgot to link it. Maybe they could not find the original so they cited the article that cited it. Maybe they made it up.

A second writer finds it, uses it, publishes it. A third writer finds both and treats the repetition itself as confirmation. An AI tool gets trained on all three articles and starts surfacing the number as established fact. A fourth, fifth and sixth article repeat it. Within eighteen months the statistic appears on 200 websites and nobody can find where it came from.

This is not a hypothetical. I have documented it happening in real time.

The repetition of a statistic is not evidence that it is true. It is evidence that people stopped checking.

A case study in how this actually works

Take this claim: “75% of customers believe it takes too long to reach a live agent.”

It circulates constantly in customer service, CX and contact centre content. It sounds credible. It has a percentage. It has a named source attached to it: Harris Interactive.

Here is what happens when you actually follow it.

Provide Support publishes it in 2016. BarnRaisers repeats it in 2017. CCSI repeats it in 2018. Jitbit repeats it in 2019, names Harris Interactive. Call Experts repeats it in 2024. HelpScout publishes a listicle called “107 Customer Service Statistics” in 2025 which gets cited by emailanalytics.com as a source. DIDWW posts it on LinkedIn in March 2026 with a designed graphic and the words “DID you know?”

Whether any of these sites ever linked to the original study or not, the result today is the same. In 2026, every one of these pages repeats the statistic and provides no path to the research. Some cite Harris Interactive by name. None provide a working link to the study. If the research exists in the form being cited, it is not findable through any of the content repeating it.

That is a decade of repetition across major platforms, established publications and a technology company’s LinkedIn presence. And at the end of it: nothing. No methodology. No sample size. No date. No link.

DIDWW’s LinkedIn post, March 2026. A professionally designed graphic citing a statistic with no source, no date and no link to any research.

This is just the first page of Google results. The same statistic continues across multiple further pages, appearing on sites ranging from industry publications to software companies to academic-looking PDFs. Every result attributes it to Harris Interactive. None provide a working link to the original study.

The first page of Google results. DIDWW, emailanalytics.com, CCSI, Call Experts, Jitbit, Provide Support, a Quality Management Institute PDF, Zendesk, Microsoft, BarnRaisers. Every result repeats the same number.

emailanalytics.com attributes the stat to “Harris Interactive and HelpScout”. The link goes to a HelpScout listicle, not to any study.

callexperts.com names Harris Interactive as the source with no link. The article was published in 2024.

At every step in that chain, a writer or marketer made a choice. They chose to use the number without checking it. That choice is not a small editorial lapse. It is the thing that keeps the ghost stat alive.

Now here is where it gets genuinely interesting.

I searched for the stat on Google this week. The AI Overview confirmed it immediately: 75% of customers believe it takes too long to reach a live agent. But the AI Overview did not cite Harris Interactive. It cited the Nextiva Customer Patience Study and the Avaya Blog.

HelpScout’s “107 Customer Service Statistics” — a listicle published August 2025. This is what emailanalytics.com links to as its source. It is a list, not a study.

The stat has mutated. After a decade of being attributed to Harris Interactive, the AI has migrated it to new hosts. It has found fresh sources to cling to. The number remains. The attribution has changed. Nobody noticed.

And Harris Interactive? The company was acquired by Toluna in July 2014. It continued to trade as a Toluna sub-brand for years after that, with its own website and LinkedIn presence. That website is now being retired and redirected to Toluna. Which means that across the entire decade this statistic has been in circulation, the source was technically reachable. It had a website. People could have gone looking. Not one of the sites citing this statistic appears to have done so.

Nobody linked to the study in 2016. Nobody linked to it in 2019 or 2024. DIDWW did not link to it in March 2026. And now, as the website finally closes, the trail goes permanently cold.

A statistic attributed to a company that had a working website for a decade and nobody visited it to check the source. Now cited by Google’s AI Overview with a completely different attribution. This is not a fringe problem. This is the first page of a Google search. Scroll further and it keeps going.

The ones I keep seeing follow this same pattern: a percentage, a plausible-sounding attribution, and a source chain that collapses the moment anyone follows it back. I will document more specific examples in part two of this series. For now, take it from someone who has spent two years following these threads: if you have written B2B content in the last five years, you have almost certainly repeated at least one of them.

I am not saying these numbers are entirely fabricated. I am saying they have been stripped of context, detached from their sources and repeated so many times that nobody using them has any idea what the original research actually said, when it was conducted or whether it is still true.

AI is accelerating this. Dramatically.

Let me be clear about something: AI is not the villain here. The problem existed long before large language models. SEO culture, content velocity incentives and the absence of editorial standards in most marketing teams created the ghost stat epidemic years ago.

But AI has poured fuel on it.

Large language models train on the internet. The internet is full of content that repeats unverified statistics. When a model gets asked for evidence to support a claim, it surfaces the most frequently repeated numbers it has seen. It does not know whether those numbers trace to credible primary research. It knows they appear often.

The result is AI tools that confidently generate statistics with citations, where the citation exists but does not contain the number, or the citation does not exist at all, or the citation leads to a page that itself cites nothing.

In 2026, AI-generated content is being published at a scale that would have been unimaginable five years ago. If the training data is polluted with fabricated statistics, the output will repeat and amplify them. The problem does not stay contained to AI content. It contaminates the entire information environment.

I have started running a simple test. Take a statistic from a piece of AI-generated content. Ask the AI tool to provide the primary source. Then go and find it. In my experience, roughly half the time the source either does not exist or does not support the claim being made. The other half, the source is real but outdated, misquoted or stripped of the context that makes the number meaningful.

This is the information environment we are operating in. And most B2B content teams are not equipped to deal with it.

This is not just a quality problem. It is a commercial one.

I hear people treat this as a minor editorial issue. A bit sloppy, not the end of the world, everyone does it. I disagree completely.

Think about what content is actually supposed to do in B2B. It is supposed to build credibility. It is supposed to demonstrate expertise. It is supposed to give buyers enough confidence to move forward with you. Every stat you cite is a signal about how seriously you take the work.

When a decision-maker reads your thought leadership and the numbers look suspicious, they do not contact you to request a correction. They write you off. The whole piece loses credibility. The argument collapses.

If the stats are sloppy, the thinking probably is too.

There are harder consequences as well. The ASA in the UK requires that marketing claims can be substantiated. The FTC in the US demands that advertising claims are truthful and backed by evidence. Neither regulator requires you to hyperlink every statistic. But if your organisation uses fabricated data to influence purchasing decisions and someone decides to push back, you are exposed.

More practically: journalists will not cite you. AI systems are getting better at cross-referencing claims against primary sources, which means content built on invented data will increasingly be excluded from AI-generated answers. The brands that have invested in genuine, evidenced, verifiable content will appear. The ones relying on ghost stats will not.

And in a world where AI Overviews, ChatGPT and Perplexity are becoming the first stop for buyers doing research, being excluded from those answers is a serious commercial problem.

Humans created this problem. Humans are perpetuating it.

I want to be specific about who is responsible here, because the easy answer is to blame AI and move on.

The ghost stat problem started with humans. Lazy citation culture, pressure to publish fast and a collective decision to treat repetition as validation. AI did not create that. Humans did. And the humans doing the most damage are not junior writers at small agencies who do not know any better. They are experienced marketers at well-resourced organisations who absolutely should.

But AI has added a new failure mode on top of the human one. And it is different in kind, not just degree.

Human laziness repeats bad citations. AI generates new ones.

Large language models do not simply regurgitate what they have been trained on. They synthesise across sources and produce confident-sounding output that can combine, distort or fabricate attribution in ways that never existed in any of the source material. The technical term is hallucination. The practical result is a citation that looks authoritative and points nowhere real.

The example in this piece is a live demonstration of exactly that. The human citation chain attributed the 75% statistic to Harris Interactive. Consistently, for a decade. But when I searched for it this week, the Google AI Overview confirmed the statistic and attributed it to the Nextiva Customer Patience Study and the Avaya Blog. Neither of those sources appeared anywhere in the human citation chain I documented. The AI did not repeat the human error. It introduced a new one.

That is the problem with blaming AI and moving on. The machine is not just a mirror. It is generating fresh attribution errors on top of existing ones, at scale, in the answer positions that buyers now trust most.

So yes. Humans created this. Humans are perpetuating it. And AI is now industrialising it in ways that will take years to unpick.

The brands publishing ghost stats are not victims of a broken internet. They are contributors to one. Every unverified statistic that goes live with their name on it makes the information environment slightly worse for everyone.

What good looks like

I built Nurtura specifically because I spent fifteen years watching this problem from the inside. The agencies I ran, the content programmes I oversaw, the briefings I sat in. The pressure to ship fast and the absence of any real standard for what “good” meant.

So let me tell you what we actually require before a statistic appears in anything we produce.

  • We trace every number to the primary source. Not the article that quoted it. The original study, survey or dataset. If we cannot find the primary source, the number does not go in.
  • We check the source exists and says what it is claimed to say. This sounds basic. It eliminates a significant proportion of the statistics that circulate in B2B marketing content.
  • We record the date, the sample size and the methodology. A survey of 60 respondents in 2012 is not the same as a survey of 3,000 respondents in 2025. Context is not optional.
  • We do not dress opinion as research. If we believe something is true based on experience, we say so. We do not invent a statistic to make the belief look more credible.
  • If the research does not exist, we say so. “Despite the number widely cited online, no reputable primary study appears to support this claim” is a more credible sentence than a fabricated percentage.

This makes content slower to produce. Good. That is the point.

It also makes the content more valuable. Buyers notice the difference between content that is evidenced and content that is not. The ones making significant purchasing decisions notice it immediately. They are exactly the buyers B2B businesses want to reach.

The bigger picture

The information environment that B2B operates in is not someone else’s infrastructure. It is shared. Every fabricated statistic that enters the content ecosystem makes it slightly harder for buyers to trust the genuine research that follows. Every ghost stat that gets trained into an AI model degrades the quality of the answers that model produces.

This sounds abstract. It is not. It is the ground underneath the entire content marketing industry. If buyers cannot trust what they read, they discount it. If they discount it, content loses its commercial value. The brands that invested seriously in building genuine expertise through content lose their advantage.

I am not arguing for perfection. I am arguing for a standard. Check the source. Date the data. Attribute the research accurately. Say what you know and be honest about what you do not.

The scarcest commodity in B2B content right now is not data. It is data you can actually verify.

The brands that recognise this and act on it will build the most durable authority over the next decade. The ones that keep repeating statistics they cannot prove will find themselves increasingly excluded from the environments where their buyers now make decisions.

That is not a moral argument. It is a commercial one.

This is the first in a series on content quality, AI and editorial standards. More to come…

Chad Harwood-Jones

Founder & Managing Director, Nurtura

More Insights