Large language models choose sources in two main ways. First, they rely on knowledge learned during training, which reflects how often and how consistently information appeared across the text they were trained on. Second, when connected to search, they retrieve web pages in real time, select the passages most relevant to the question, and synthesize an answer, often citing the pages they used. Sources are more likely to be retrieved and cited when they are accessible to the system’s crawlers, rank well in the underlying search index, answer the question clearly and specifically, and are corroborated by other trustworthy sources.
This guide is for marketers, business owners, and content teams who want to understand what actually happens when ChatGPT, Perplexity, Gemini, Copilot, or Google’s AI features decide which sources to use. I will explain the difference between training and retrieval, walk through the retrieval process step by step, describe the signals that influence source selection, cover how platforms differ, and turn all of that into practical actions. None of the major providers publish complete formulas, so I will be careful to separate what is documented or widely observed from speculation. This article is part of my series on generative engine optimization.
Two Ways an LLM “Knows” Things
Parametric knowledge from training
A large language model is trained on very large amounts of text, including publicly available web content, books, code, and licensed data. During training, it learns statistical patterns about language and facts. This knowledge is stored in the model’s parameters, which is why it is sometimes called parametric knowledge.
Parametric knowledge has three important characteristics. It has a cutoff date, after which the model knows nothing unless it retrieves new information. It tends to reflect what was common and consistent in the training data: well known brands and widely repeated facts are represented more strongly than obscure ones. And it does not come with citations; the model cannot point to which document taught it a particular fact.
Retrieved knowledge from search
Most AI assistants can now search the web. When a question needs current, specific, or verifiable information, the system retrieves documents and uses them to ground its answer. This approach is often called retrieval augmented generation. Retrieved knowledge is current, it can be cited, and it depends heavily on search indexes and on how well pages match the question.
For businesses, retrieval is where most of the practical opportunity lies. You cannot quickly change what a model learned during training, but you can influence whether your pages are retrieved and cited when the model searches.
| Aspect | Training (parametric) knowledge | Retrieval (search) knowledge |
|---|---|---|
| Source | Large training datasets | Live web search and indexes |
| Freshness | Fixed at a cutoff date | Current |
| Citations | Not available | Often shown |
| How to influence | Long term, consistent presence across the web | Accessible, relevant, clear, trusted pages |
| Speed of change | Slow, only with new model versions | Relatively fast, as pages are crawled and indexed |
How Retrieval and Source Selection Work
Each platform implements this differently, but the general process follows a recognizable sequence.
Step 1: Deciding whether to search
The system first decides whether it needs to search at all. Simple general knowledge questions may be answered from training alone. Questions about recent events, specific businesses, prices, comparisons, or anything the model is uncertain about are more likely to trigger a search. Some interfaces also let users request a search directly.
Step 2: Rewriting and expanding the query
The user’s question is often rewritten into one or more search queries. A single conversational question may be broken into several searches covering different aspects. Google describes this in AI Mode as query fan out: running multiple related searches across subtopics to assemble a fuller answer. This means your page may be retrieved for a subquery rather than the original question.
Step 3: Retrieving candidate documents
The system sends those queries to a search index. Depending on the platform, this may be its own index, a partner’s index, or a combination. Documents that rank well for the queries become candidates. This is the step where traditional SEO matters most: if a page is not indexed or does not rank for the relevant queries, it is unlikely to be considered.
Step 4: Reading and selecting passages
The system fetches or accesses the content of candidate pages and identifies the passages most relevant to the question. Clear headings, direct answers, and self contained statements make this easier. A page might rank well but contribute nothing if its useful information is buried, vague, or hard to separate from surrounding text.
Step 5: Reranking and filtering
Systems typically apply additional judgments about relevance, quality, and reliability. They may prefer sources that agree with other trusted sources, avoid low quality or spammy pages, and aim for some diversity so that one site does not dominate the answer.
Step 6: Synthesizing and citing
The model writes an answer using the selected passages and its own knowledge, and attaches citations to some statements. Not every source used is necessarily cited, and not every citation receives equal prominence. Brands may also be mentioned without a link, especially when the model draws on its training knowledge.
Signals That Influence Source Selection
No provider has published a complete list, but documentation, patents, and consistent observation point to several categories of signals.
Accessibility
The page must be reachable by the platform’s crawler and present in its index. Blocking search related AI crawlers, or accidentally blocking them through firewalls, removes a page from consideration. My guide to AI crawlers, robots.txt, and llms.txt covers this in detail.
Search relevance and ranking
Because retrieval often uses search indexes, pages that rank well for the relevant queries are more likely to be candidates. Strong conventional SEO remains the most direct influence on retrieval.
Specificity and clarity
Passages that answer the question directly and specifically are easier to use. A paragraph that clearly states a definition, a number, a step, or a comparison is more useful than one that talks around the topic. This is why answer-first content writing has become so valuable.
Authority and trust
Systems tend to favor sources with strong reputations in a subject: established publications, recognized experts, official sources, and sites with deep coverage of the topic. Building topical authority and earning links and mentions, as covered in my link building guide, strengthen these signals.
Corroboration
Claims supported by several independent sources are safer for an AI system to repeat. If many credible sites describe your company or recommend your product for a particular use, those consistent references make it more likely the system will include you.
Entity clarity
Systems need to know which entity a source is about. Clear, consistent information about your business across your site and external profiles reduces confusion, as explained in my guide to entity SEO and the Knowledge Graph.
Freshness
For time sensitive topics, current content is often preferred. Visible publication and update dates, and genuinely updated information, help when recency matters.
Format and structure
Structured content, such as descriptive headings, lists, tables, and FAQ style questions, helps systems locate relevant passages. Content locked in images, PDFs, or scripts that do not render is harder to use.
Content Types That Tend to Be Retrieved
Certain kinds of pages appear again and again as sources in AI answers, because they match the way people ask questions.
Comparison and “best of” pages
When someone asks for recommendations or compares options, systems often retrieve comparison articles, review roundups, and category guides. These pages collect several options in one place, which suits the task of summarizing choices. Being included fairly in credible comparisons, and publishing honest comparisons of your own, both help.
Clear definitions and explainers
Questions that begin with what or why frequently draw on pages that define a concept precisely and explain it plainly. Glossaries and explainer guides from recognized experts are common sources.
Step by step guides
How to questions draw on pages that present processes as clear, ordered steps with practical detail.
Pricing and specification pages
Questions about cost, features, and specifications draw on pages that state them explicitly. When a business hides this information, AI systems rely on third party estimates instead, which may be less accurate.
Review platforms and community discussions
For questions about reputation or real world experience, systems often use review sites, forums, and community threads. These are sources you cannot write yourself, which is why genuine customer reviews and participation in communities matter.
Why Third Party Sources Matter So Much
It is tempting to focus only on your own website, because it is the part you control. But for many commercial questions, AI answers lean heavily on third party sources: review platforms, industry publications, comparison sites, and discussions. From the system’s perspective, independent sources are often more reliable guides to which businesses are good than businesses’ own descriptions of themselves.
This does not make your website less important. It means your website and your external reputation need to tell the same story. When your site clearly explains what you do and for whom, and independent sources confirm it, AI systems have both the facts and the corroboration they need to include you confidently.
How Platforms Differ
The general process is similar across platforms, but the details vary.
- ChatGPT combines training knowledge with web search, using OpenAI’s own search crawler alongside third party search providers, and shows citations for search based answers. My guide on getting your brand mentioned in ChatGPT covers it specifically.
- Perplexity is built around search and shows sources prominently for almost every answer, which makes it especially useful for seeing which sources are trusted in a category.
- Google AI Overviews and AI Mode are grounded in Google Search, so ranking and eligibility in Google are the main factors. My guide to Google AI Overviews explains the specifics.
- Gemini can draw on Google Search for grounding, so Google search signals and entity information matter.
- Microsoft Copilot draws on Bing’s index, which makes Bing visibility important.
Because each platform uses different indexes and methods, the same question can produce different sources on different platforms, and even on the same platform at different times.
What This Means for Your Strategy
Win retrieval first
Most practical gains come from being retrieved. That means being indexed in Google and Bing, allowing search related AI crawlers, and ranking well for the questions and subquestions your buyers ask.
Make passages easy to use
Structure each page so every important question has a clear heading and a direct, specific answer. Assume any paragraph might be read in isolation.
Build corroboration beyond your site
What others say about you carries weight. Reviews, industry coverage, partner pages, comparison articles, and community discussions all contribute to the consistent picture AI systems look for.
Invest in long term representation
Training knowledge changes only with new model versions, but a consistent, widely referenced presence increases the chance that future models learn accurate information about you.
Measure across platforms and over time
Because answers vary, track a consistent set of prompts across several platforms and look for trends rather than reacting to individual responses. The relationship between search and AI focused disciplines is covered in SEO vs AEO vs GEO, and my AI visibility guide explains how to set up measurement.
Common Misconceptions
- “AI answers are random.” They vary, but source selection follows consistent patterns tied to relevance, accessibility, and trust.
- “SEO no longer matters.” Retrieval usually depends on search indexes, so SEO is more important, not less.
- “You can submit your site to ChatGPT.” There is no submission process for organic answers; visibility comes from crawling, indexing, and relevance.
- “Hidden instructions can make AI cite you.” Attempts to manipulate AI systems with hidden text are unreliable, can be detected, and damage trust.
- “One test tells you the truth.” Individual answers vary; patterns across repeated tests are what matter.
- “Blocking training crawlers removes you from AI search.” Training and search crawlers are often separate; blocking one does not necessarily affect the other.
Best Practices
- Ensure your pages are indexed in Google and Bing.
- Allow search related AI crawlers if AI visibility is a goal.
- Answer specific questions directly under descriptive headings.
- Cover topics comprehensively, including common follow up questions.
- Use tables and lists for comparisons and processes.
- Keep business information consistent across the web.
- Earn credible third party mentions and reviews.
- Show authors, expertise, and update dates.
- Track a consistent prompt set across platforms over time.
Practical Example: A Project Management Software Company
Consider a company selling project management software for construction firms. This is an illustrative scenario. When users asked AI assistants for “project management software for construction companies,” the answers named large general tools and a few construction specific competitors, but rarely this company.
Looking at the retrieval process explains why. The query would likely be expanded into subqueries such as “construction project management software features,” “best construction project management tools,” and “construction scheduling software comparison.” For those searches, the retrieved pages were comparison articles on software review sites, industry publications, and competitors’ detailed feature pages. The company’s own site had a general homepage and a features page written for “teams of all kinds,” with little construction specific language.
The strategy would address each stage. For retrieval, the company would create dedicated pages for construction use cases, such as scheduling subcontractors, tracking change orders, and managing site documentation, each optimized for the relevant searches. For passage selection, each page would open with a clear statement of what the software does for construction firms and include a comparison table of features relevant to construction workflows. For corroboration, the company would encourage construction customers to leave reviews on major software review platforms, pursue inclusion in credible comparison articles, and contribute expertise to construction technology publications.
The realistic outcome is that the company becomes a more frequent candidate for retrieval and a more credible option when systems compile recommendations, particularly for construction specific questions. It cannot guarantee mentions, but it aligns the company with the way AI systems actually find and choose sources.
Checklist: Becoming a Source AI Systems Choose
- Important pages indexed in Google and Bing.
- Search related AI crawlers allowed and not blocked by firewalls.
- Pages exist for the main questions and subquestions buyers ask.
- Each page answers its questions directly under descriptive headings.
- Comparisons and processes presented in tables and lists.
- Business and product information consistent across the web.
- Credible reviews, mentions, and comparison coverage pursued.
- Authors, expertise, and update dates visible.
- Prompt set tracked across ChatGPT, Perplexity, Gemini, Copilot, and Google AI features.
Frequently Asked Questions
Do LLMs always search the web before answering?
No. Many answers come from training knowledge alone. Systems are more likely to search for current, specific, or uncertain information, or when the user asks them to.
Why does the same question give different sources?
Query rewriting, retrieval, and generation involve variation, and indexes change over time. Different platforms also use different indexes. Look for patterns across repeated tests.
Can I see which pages an AI used?
Often partly. Search based answers usually show citations, although not every source used is necessarily cited. Training knowledge cannot be traced to specific pages.
Does ranking on Google guarantee AI citations?
No, but it helps considerably, especially for Google’s own AI features and for systems that rely on search indexes. Clear, specific passages and strong trust signals also matter.
Do AI systems favor big brands?
Well known brands have an advantage in training knowledge because they are mentioned more often. In retrieval, however, specific and clearly answered questions give smaller specialists a real chance, particularly for niche topics where large brands offer only general coverage.
How can a small business become a cited source?
Focus on specific questions within your specialty, answer them better than general sources, keep your business information consistent, and build reviews and local or industry mentions.
Will changes to my site affect AI answers quickly?
Changes can influence retrieval based answers after pages are recrawled and reindexed, which can take days to weeks. Changes to training knowledge only occur when new model versions are trained.
Conclusion
LLMs choose sources through a combination of what they learned in training and what they retrieve in real time. For most businesses, retrieval is where the practical opportunity lies: be accessible to the right crawlers, rank for the questions and subquestions your buyers ask, write passages that answer clearly and specifically, and build the corroboration and entity clarity that make your information trustworthy. Over the long term, a consistent, credible presence across the web also shapes what future models learn.
Understanding the mechanics removes much of the mystery. AI systems are not choosing sources at random; they are looking for accessible, relevant, clear, and trusted information. The businesses that provide it are the ones most likely to be chosen.
Want to Become a Source AI Systems Trust?
If competitors keep appearing in AI answers for your category, I can analyze which sources are being retrieved and cited, identify the gaps in your content and reputation, and build a plan to improve your chances of being chosen. Get in touch through DigitalKetan.com to discuss your AI visibility.
