How AI Assistants Actually Choose Which Brands to Cite (and How Small Sites Win)
You have probably had this experience by now: you ask ChatGPT or Google to recommend software in your own category, and the answer cites a competitor’s blog, a Reddit thread, and some comparison site you have never heard of - while your site, the one that has ranked for these terms for years, is nowhere in the answer. Most of what gets written about this problem is statistics: adoption is up, click-through is down, panic accordingly. Almost nobody explains the mechanism - how an answer engine actually assembles a response and decides which three to six sources earn a citation. That mechanism is worth understanding, because it is genuinely different from ranking, and the difference is the best news small marketing teams have had in a decade.
The stakes are already visible in the numbers. In CommonMind’s State of AI Visibility in B2B SaaS, 93% of B2B SaaS marketers say AI search visibility is critically important - and only 14% have a mature strategy for it. And per a G2 finding collected in TripleDart’s AI SEO statistics roundup, 69% of buyers say an AI chatbot influenced which vendor they ultimately selected. The answers are being written whether you show up in them or not.
How does an AI assistant actually build an answer?
Start with the definitional part, because it explains everything downstream. Modern answer engines do not answer product and buying questions from memory. They use retrieval-augmented generation (RAG): when a question needs current or specific information, the system runs live searches against a web index, retrieves a set of candidate pages, extracts the passages that appear to answer the question, and hands those passages to the language model, which writes a synthesized answer and attaches citations to the sources it actually used. (Ahrefs has a good technical explainer of the RAG pipeline if you want the deeper version.)
Two consequences of that pipeline matter more than everything else written about AI search:
First, retrieval and citation are separate steps. Being retrieved gets your page into the candidate pool. Being cited requires that the model actually used your passage to write a sentence of the answer. A page can rank, get retrieved, and still contribute nothing citable.
Second, citation is binary. A classic search result page has ten positions and position seven still gets some clicks. An AI answer cites a handful of sources and ignores everyone else. You are in the answer or you do not exist for that question.
Do ChatGPT, AI Overviews, and Perplexity work the same way?
The architecture is similar; the details differ enough to matter.
| Engine | Where it retrieves from | How it selects and cites |
|---|---|---|
| Google AI Overviews / AI Mode | Google’s own search index, grounded via Gemini | Fans the query out into multiple sub-queries, retrieves for each, cites passages that support each part of the answer - often not the top-ranked pages |
| ChatGPT (with search) | Live web search backed by Bing’s index plus OpenAI’s crawling | Decides per-question whether to search, evaluates candidates for how directly they answer, cites a small set inline |
| Perplexity | Its own crawl and index; retrieval runs on nearly every query | Retrieval-first by design; cites multiple sources through the answer, one per claim that contributed unique information |
The query fan-out step deserves a sentence of its own, because it is the piece most teams miss. When someone asks Google’s AI a broad question like “best CRM for a 20-person SaaS company,” the system does not run one search. It decomposes the question into sub-queries - pricing, integrations, comparisons, complaints - retrieves for each, and assembles the answer from across all of them. Every sub-query is a separate chance to be a source. A page that thoroughly answers one narrow sub-question can get cited in answers to broad questions it could never have ranked for.
One more difference worth knowing: these systems re-run retrieval constantly, so the set of cited sources is far less stable than rankings ever were. Analyses collected by TripleDart found 40 to 60% of cited sources change month to month across Google AI Mode and ChatGPT. Unstable is bad if you are defending a position. It is very good if you are trying to take one.
What makes a page get selected as a citation?
Nobody outside these companies has the exact formula, and anyone selling you precise percentage weights is guessing. But the selection pressures are consistent across engines and well-documented in practice:
- Relevance to the specific sub-question. Not topical relevance to the general subject - a direct match between the question being retrieved for and the answer your passage gives.
- Extractability. The model needs a passage it can lift: a direct answer near a clear heading, a table, a numbered list. A 400-word wind-up before the point is a retrieval dead end.
- Entity and authority signals. Engines increasingly reason about entities (companies, products, people), not just keywords. Notably, the signal that correlates with citations is not the one classic SEO optimized for: the same TripleDart roundup reports brand web mentions correlate with AI citation rates at 0.664, roughly three times the backlink correlation of 0.218. Being talked about beats being linked to.
- Consensus. A claim that appears consistently across several independent sources is safer for the model to assert. Contradictory descriptions of your own product across your site, your review profiles, and your directory listings actively suppress you.
- Freshness. Live retrieval means engines can prefer current pages, and for anything with prices, versions, or “best of” framing, they visibly do.
Domain authority is not gone - it moved. A trusted domain still helps a page get into the retrieval pool. But at the citation step, the engine is choosing between extracted passages, and a precise answer from a small site regularly beats a vague page from a big one. That step barely existed in classic SEO. It is where small sites win.
Why do forums and comparison posts get cited so often?
Ask any AI assistant a buying question and count the citations: Reddit threads, review sites, “X vs Y” posts. This annoys vendors, but by the mechanism above it is exactly what you should expect. As Hootsuite’s guide to LLM visibility puts it, “articles, reviews, social posts, forum discussions, and public content all feed into the system” - and community and comparison content happens to be shaped precisely how the selection step wants it.
A forum answer says “we switched from X to Y because the reporting was better for a team our size.” That is a direct, extractable, experience-backed answer to the sub-query “X vs Y for small teams.” A comparison post has the table already built. Meanwhile the typical vendor page says the company “empowers modern teams to unlock growth” - a passage that answers no question anyone asked and gives the model nothing to lift. The engines are not biased toward forums. They are biased toward answers, and forums are where the answers are.
The practical read: you cannot out-rank Reddit for these queries, but you can be present where the engines retrieve - which means earning honest mentions in the communities your buyers actually ask, keeping review profiles current, and making sure comparison posts about your category include accurate facts about you.
The playbook: how to become the clearest complete answer
Everything above compresses into six moves. None of them require a big domain. All of them require discipline.
- One question per page (or per section). Retrieval happens at the sub-query level, so match it: pick the specific questions your buyers ask and give each a dedicated page or a dedicated H2. Question-phrased headings are not a gimmick; they are literal retrieval targets.
- Answer first, then elaborate. The first sentence under each heading should be a complete, standalone answer - the passage you want quoted. Context, caveats, and story come after. If a sentence could be lifted into an AI answer and still make sense, you wrote it correctly.
- Structure anything comparative or procedural. Tables for comparisons, numbered steps for processes, definitional paragraphs for concepts. Structured passages get extracted; walls of prose get skipped.
- Make your entity facts identical everywhere. One company description - what you do, for whom, in which category - used verbatim on your site, LinkedIn, review profiles, and directories. Consensus across sources is a selection signal; self-contradiction is a suppression signal. (Structured data and clean crawl access help here too; whether files like llms.txt do anything yet is a separate question we have tested.)
- Earn mentions where the engines retrieve. Given that mentions correlate with citations about three times as strongly as backlinks, the link-building hour is better spent getting genuinely useful answers, reviews, and comparisons about you onto the community and review sites already being cited in your category. We cover the full version of this in how to get your SaaS recommended by ChatGPT.
- Keep it visibly fresh. Show updated dates, and actually update the pages that answer time-sensitive questions - pricing, comparisons, “best” lists. With 40 to 60% of citations churning monthly, every refresh is a new roll on questions you currently lose.
Measure answers, not rankings. Your rank tracker cannot see any of this. Ask the major engines your ten most valuable buying questions monthly, record who gets cited and how you are described, and treat a wrong description as a bug to fix at its source. We compared the tooling options in our review of AI visibility tools, and the traffic side of the story in what AI Overviews are doing to B2B traffic.
Where Marqeable fits
One honest note on execution, because the playbook above is mostly a writing-discipline problem. Content created in Marqeable’s AI content studio starts from a brief - the question being answered, the audience, the key points - and gets reviewed for structure and clarity before it ships. Those are the same properties that make a page citable: a brief forces one clear question per piece, and review catches the vague openings and unsupported claims that retrieval skips over. We are in private beta with early customers; if you want the citable-by-default version of your content operation, get early access.
Frequently asked questions
How do AI Overviews choose sources?
Google AI Overviews use Gemini grounded in Google’s search index. The system fans your question out into multiple sub-queries, retrieves candidates for each, and cites the passages that most directly support each sentence of the answer - which is why cited pages are often not the top-ranked organic results.
How does ChatGPT pick sources?
When ChatGPT decides a question needs current information, it runs a live web search backed by Bing’s index and OpenAI’s own crawling, evaluates which candidate pages actually answer the question, and cites a small set inline. There is no position three; a source is cited or invisible.
Why does AI search cite Reddit and comparison sites so much?
Because their content is shaped the way selection works: direct, experience-backed, extractable answers to specific buying questions. Vendor pages that describe themselves in the abstract give the model nothing to lift, so it quotes the forum thread that answered plainly.
Can a small site really beat a big brand in AI citations?
Yes, at the citation step. Selection happens per passage, the clearest complete answer wins, and cited sources churn 40 to 60% month to month - so the position is winnable in a way a locked-in ranking never was.
The bottom line
AI assistants build answers by retrieving candidates, selecting the clearest extractable passages, and citing the few sources they actually used. Every step of that pipeline rewards precision over size: fan-out rewards pages that own one narrow question, extraction rewards answer-first structure, entity signals reward consistent facts and real mentions, and monthly citation churn means the winners are re-decided constantly. Big brands with vague pages are structurally disadvantaged here, and only 14% of B2B SaaS teams have a mature strategy for it. Be the clearest complete answer to the questions your buyers ask, everywhere the engines look, and you can win citations your domain rating says you have no business winning.
Marqeable runs your campaigns, answers every visitor, text, and email in seconds, and turns them into booked jobs and meetings - even at 9pm on a Saturday. We’re in private beta with a small early cohort. Get early access
