The Fake Stat Problem: Why AI Invents Numbers, Quotes and Sources (and the 3 Tells)
Search for statistics about AI hallucinations in marketing and one figure comes up again and again: roughly two-thirds of marketers have encountered hallucinated content, and fewer than half have a formal verification process, attributed to a well-known university AI institute. It appears on agency blogs, on a content-tool vendor’s page, on a couple of Medium posts. It has the shape of a fact. When we tried to trace it to a report, a survey, a page on the institute’s site, we could not find one. Several sites cite each other. None cites a primary.
It may exist somewhere. But a statistic about hallucination that cannot be traced to its source, ranking on page one, is the whole problem in one paragraph. This post is about why language models produce numbers, quotes and citations that look like evidence and are not, how to spot them in seconds, why “I found it in two places” is not verification, and a routine for checking AI-drafted marketing copy that fits inside the review step you already have.
Stats, quotes and citations are one failure
It is tempting to treat these as three separate problems. A made-up percentage, a fabricated quote from a customer or an executive, a citation to a paper that does not exist. They are the same failure wearing different clothes.
A language model produces the most plausible continuation of the text so far. When the text so far is a marketing paragraph that needs support, the most plausible continuation is a sentence shaped like support: a number with a percent sign and a source, a quote with a name and a title, a citation with an author and a year. The model has seen millions of these shapes. It can complete the shape without any of the content behind it being real, because completing shapes is the only operation it has. The number is not retrieved from anywhere. It is generated to fit.
This is why the failure is so hard to catch by reading. The output is not sloppy. It is exactly as fluent and confident as the true sentences around it, because fluency and confidence are properties of the shape, not of the fact. The Neil Patel team’s 2026 hallucination study, which surveyed 565 marketers and ran 600 prompts across six models, is the only primary research we found on this specifically for marketing. Its headline findings: about 47 percent of marketers run into AI inaccuracies several times a week, about 37 percent have had hallucinated content go live, and more than 70 percent spend one to five hours a week fact-checking. Those are the numbers from a study you can open and read, which is the standard this post is about to argue for.
The three tells
You do not need a detector. Most fabricated evidence gives itself away with one of three tells, and you can learn to see them in the time it takes to read the sentence.
Tell 1: the tidy number. Real survey results are lumpy: 47.1 percent, 36.5 percent, a sample of 565. Fabricated ones are round: 70 percent of marketers, a 3x increase, a 40 percent lift. Round numbers exist in real research, so this tell alone is not proof. It is a reason to look. A model that is completing a shape reaches for the number that fits the rhythm of the sentence, and the rhythm prefers round.
Tell 2: the orphan citation. The author is real. The journal or outlet is real. The paper or article does not exist. This is the most convincing kind of fabrication because two-thirds of it checks out. Academic studies of model citations found this pattern repeatedly in 2023 and it has not gone away; a public database maintained by a legal researcher tracks well over a thousand court decisions in which lawyers filed briefs with citations to cases that were never decided. Lawyers, with sanctions on the line, kept doing it, because the citations looked right. Your nurture email is not going before a judge, but the mechanism is identical.
Tell 3: the quote with no primary. “As one CMO put it…” followed by something a CMO plausibly would put. A customer quote that sounds like every customer quote. An executive quote from a company that is real, saying something the company might say, that no transcript, interview, press release or on-the-record page contains. A quote is just a citation with a name instead of a title, and the model fabricates it the same way.
One tell means stop and check. Two means assume fabricated until proven otherwise.
Hallucinated numbers launder themselves. A model invents a figure. A blog publishes it. A second blog cites the first. A third cites both. Within a few months the figure has three sources, and none of them is a study. This is why the opening statistic in this post is so hard to kill and why “I found it in two places” is not verification. Only a primary counts: the survey, the dataset, the filing, the transcript, the page where the person said it.
Why the newer models have not fixed it
Every model release improves on this. Newer models hedge more, refuse more, and cite real things more often. None of them has a mechanism that ties a generated number to a source, because that is not what generation is. Three things make it persist.
First, the model rewards itself for confidence. Training that optimizes for answers people rate highly favors sentences that sound sure, and a sentence with a specific number sounds surer than one without. A 2025 study from Carnegie Mellon found that chatbots became more confident after performing badly, not less, which is the opposite of how a careful human calibrates.
Second, the model does not know what it does not know about you. Ask it for your company’s average response time, your pricing, your customer count, and it will produce a plausible one. Your own product facts are the category most likely to be wrong in your own copy, and the least likely to be caught, because the person reviewing assumes the model got them from somewhere.
Third, the failure hides in the middle. A 2025 audit by the Tow Center for Digital Journalism found AI search tools answered questions about news content incorrectly more than 60 percent of the time, and often with the same confident phrasing whether right or wrong. If the tools built specifically to cite sources get it wrong that often, a general-purpose model asked to support a marketing claim is not going to do better.
Verify by claim class
The reason fact-checking eats an afternoon is that people check line by line. The faster approach is to sort claims into classes and give each class exactly one check. Almost everything in a marketing draft falls into five.
| Claim class | Example | The one check |
|---|---|---|
| Numbers | ”47% of marketers…”, “3x faster” | Open the primary. If you cannot, cut the number and keep the point. |
| Names and titles | ”Jane Doe, VP Marketing at Northwind” | A current lookup. Titles change; models remember the old one. |
| Dates and versions | ”since the March 2026 update” | The changelog, the announcement, the calendar. |
| Product facts about you | Your pricing, tiers, features, response times | Your own facts sheet. Never the model’s memory of your website. |
| Quotes | Anything in quotation marks with a name | The transcript, the interview, the on-the-record page. No primary, no quote. |
Two rules make this fast. The point survives the number. If a stat cannot be verified in two minutes, delete the stat and keep the sentence; “most marketers spend hours a week fact-checking” is true and needs no citation, while “73 percent of marketers spend 4.2 hours a week” needs a study you do not have. Your own facts get a sheet. A one-page document of the numbers, names, prices and claims you are allowed to make, kept current, is the highest-value fact-checking asset a small team can own, because the model will otherwise fill that gap with the most plausible version of your company. The marketing knowledge base post covers how to build it.
Related: the claims that carry legal weight (performance results, testimonials, health and finance claims) have their own standard, which the FTC substantiation post covers. The FTC does not care that the model told you there was a study.
Where this lives in a workflow
A team of one to five cannot add an afternoon of checking to every piece. The check has to sit inside the review step that already exists, and it has to be a different pass from the drafting, because the model that wrote the sentence is the worst judge of it.
The shape that works: the draft is produced with the brand facts sheet attached as context, so the product-facts class is right at the source. A review pass, before a human reads for tone, flags every sentence in the five classes that lacks a source: every number, every name, every date, every quote, every claim about you that is not on the sheet. The human review then starts with a short list of flagged items instead of a wall of prose, and the review agent has already caught the round number with no citation and the customer quote nobody said. Nothing is scheduled until the flags are cleared.
That is how Marqeable’s content review works: drafts are written against your brand facts, the review pass flags numbers, quotes and claims it cannot trace to a source, and a person clears the list before the piece can move to ready. The model still drafts fast. It just cannot ship a statistic that nobody checked.
Frequently asked questions
Does ChatGPT make up statistics?
Yes, and so does every other model, because a statistic is a shape (number, percent sign, attribution) the model can complete whether or not the fact exists. Newer models do it less and hedge more; none ties a number to a source. Treat every number in a draft as unverified until you have seen the primary.
How can I tell if a statistic or citation is hallucinated?
Three tells: a suspiciously tidy number, an orphan citation (real author, real outlet, nonexistent piece), a quote with no locatable primary. One tell means check; two means assume fabricated.
Is finding the same stat on two websites enough verification?
No. Fabricated numbers launder themselves through blogs citing blogs. Verification means the primary source: the survey, the dataset, the filing, the transcript.
How do marketing teams check AI content without spending hours?
Sort claims into five classes (numbers, names, dates, product facts, quotes), give each one check, keep a facts sheet for your own claims, and run a review pass that flags unsourced items before a human reads for tone.
The bottom line
A language model completes shapes, and evidence has a shape. Numbers, quotes and citations come out fluent and confident whether or not anything stands behind them, and the ones that are wrong look exactly like the ones that are right. Learn the three tells, refuse anything short of a primary source, keep a facts sheet for your own claims, and put the check inside the review step as a separate pass from the drafting. The model’s speed is real. So is the afternoon it costs when a made-up number goes live under your name.
Marqeable runs your campaigns, answers every visitor, text, and email in seconds, and turns them into booked jobs and meetings - even at 9pm on a Saturday. We’re in private beta with a small early cohort. Get early access
