How AI Search Finds and Selects Sources: What We Actually Know
Last reviewed: August 2026. Platform documentation in this area changes regularly. This guide distinguishes publicly documented behaviour from Echeva's interpretation, and is reviewed periodically as providers publish new information.
A lot of AI search advice describes retrieval as though it's one well-understood process that works the same way across every platform. It isn't. Each major provider documents its own systems differently, in varying levels of detail, and none of them publish a complete account of how sources are ultimately selected. This guide sets out what's actually confirmed by Google, OpenAI, and Microsoft, and separates that clearly from what's commonly claimed but not documented.
Why this distinction matters
Most AI search content blends confirmed platform behaviour with inferred or assumed mechanics, often without saying which is which. That's understandable, since businesses want clear, actionable answers, and "nobody fully knows" isn't a satisfying one. But treating inference as fact leads to confident-sounding advice that isn't actually grounded in anything a platform has said.
This guide is organised around a simple structure: what a platform has stated directly, and what remains undisclosed. Where we don't have a primary source for a claim, we say so, rather than presenting it as settled.
What Google confirms
Google published an official guide to optimising for its generative AI Search features in May 2026, and it's now the clearest primary source available on this topic. Google says its generative AI features, including AI Overviews and AI Mode, are rooted in its core Search ranking and quality systems, and specifically documents two techniques:
Retrieval-augmented generation (RAG), also called grounding, which Google describes as relying on its core Search ranking systems to retrieve relevant, up-to-date web pages from its Search index, then generating a response from the specific information in those pages, with clickable links back to the sources used.
Query fan-out, a set of concurrent, related queries generated by the model to gather more information and fetch additional relevant results addressing the user's original question. Google's own example: a query like "how to fix a lawn that's full of weeds" might fan out into related searches such as "best herbicides for lawns" or "how to prevent weeds in lawn."
Google is also explicit about what it doesn't require: no special AI-specific structured data, no llms.txt file, and no need to chunk content into small AI-oriented fragments. Google's crawler may discover an llms.txt file, but treats it like any other text file, with no special treatment.
What Google hasn't published is the complete weighting or decision process that determines which eligible pages ultimately become supporting links within a particular AI-generated response.
What OpenAI confirms
OpenAI states that when ChatGPT Search is used, ChatGPT can retrieve current information from the web and return linked sources, going beyond information available in its built-in model knowledge. It also confirms ChatGPT Search can rewrite a user's request into one or more targeted search queries when working with search providers, rather than always searching for the literal text a user typed.
OpenAI states that ranking in ChatGPT Search is based on a number of factors designed to help users find reliable, relevant information, while being explicit that there's no way to guarantee top placement. For a site's content to be discoverable and usable within ChatGPT Search summaries and snippets, OpenAI says site owners should allow its OAI-SearchBot crawler and ensure hosting or CDN configuration permits traffic from its published IP ranges.
What OpenAI hasn't published is the specific set of factors its ranking is based on, or how they're weighted relative to each other.
What Microsoft confirms
Microsoft documents that Copilot Studio agents can access the web via Bing APIs, through specific URLs, open web search, or custom search configurations, and that enabling web search causes an agent to generate a search query in response to a user's question. Microsoft separately documents Grounding with Bing Search, used within Microsoft Foundry, as a tool that lets AI agents incorporate real-time public web data; its documented process moves through query formulation, search execution, information synthesis, and source attribution. See our guide on why Bing matters for AI search for a closer look at Bing's role across Microsoft's ecosystem.
What Microsoft hasn't published in the same level of public detail is the complete ranking or selection logic behind which Bing results ultimately contribute to a particular Bing-grounded Copilot Studio or Microsoft Foundry agent response. It's also worth being careful not to assume every product carrying the "Copilot" name shares identical architecture; the documentation above specifically covers Copilot Studio and Foundry, not every Microsoft product using the Copilot name.
What is publicly confirmed, by platform
| Platform | What is publicly confirmed | What remains undisclosed |
|---|---|---|
| Google Search | Core Search ranking systems, retrieval-augmented generation, query fan-out, use of the Search index, clickable supporting links | Complete source-selection and weighting logic |
| ChatGPT Search | Query rewriting, use of search providers, live web retrieval, source links, ranking based on multiple factors | The specific ranking factors and how they're weighted |
| Microsoft Copilot Studio / Foundry | Query generation, Bing-based retrieval, result synthesis, source attribution | Complete result-selection and weighting logic |
What the documentation lets us say with confidence
Accessibility matters for live web retrieval. Google requires a page to be indexed and eligible to appear with a snippet in order to appear as a supporting link within its generative Search features. OpenAI says allowing OAI-SearchBot is important for ChatGPT Search inclusion. Bing-grounded Microsoft experiences draw on public content indexed by Bing.
Relevance matters, but each platform evaluates it differently. Google explicitly describes retrieving relevant pages through its Search ranking systems; OpenAI says ChatGPT Search ranking seeks reliable, relevant information; Microsoft says Bing returns relevant results to its grounded agents. None of the three publishes one universal definition or weighting of relevance.
There's no documented universal AI citation formula. Things frequently discussed in AEO content, precise weighting for corroboration, page structure, word count, freshness, or particular content formats, shouldn't be presented as universal platform ranking factors unless the relevant provider actually documents them. Google is explicit, for example, that content doesn't need to be rewritten in a special format or chunked for its generative AI features.
Model knowledge and live web retrieval are different
Live web retrieval and a model's existing knowledge are different sources of information. ChatGPT Search allows ChatGPT to retrieve current information from the web, going beyond information available in its built-in model knowledge. When a response isn't grounded through a live web search, businesses shouldn't assume that changing a webpage will immediately change what the model says about them.
Exactly what information contributes to a particular non-search response can depend on the specific product and context, so it's more accurate to distinguish live retrieval from non-live responses than to treat every answer as coming exclusively from either "the web" or "training data."
A useful high-level model, not a universal architecture
Across the documented, web-grounded systems above, a useful conceptual pattern emerges: question → query formulation → search or retrieval → returned information → synthesis → source attribution.
Google documents RAG and query fan-out, OpenAI documents iterative targeted searches, and Microsoft explicitly documents query formulation, retrieval, synthesis, and attribution as distinct stages. But these platforms don't implement this pattern identically, and none has published a complete, end-to-end account of the final selection and weighting involved. This should be treated as a simplified model for understanding retrieval-grounded AI search, not as evidence of a single underlying "AI search algorithm" that can be reverse-engineered.
Common mistakes
Presenting a single unified "AI ranking algorithm." No such thing has been documented publicly, and treating retrieval and generation as identical across Google, OpenAI, and Microsoft misrepresents what's actually known.
Confusing retrieval with training data. Advice that assumes every AI response involves a live web search overlooks that some responses draw on different sources of information depending on the product and context.
Citing unverified "ranking factors" as fact. A great deal of published content lists specific weighted factors as though they're confirmed, when in most cases they're inferred from observation rather than stated by the platform itself.
Chasing tactics a platform has explicitly said don't matter. Google specifically states that llms.txt files, AI-specific structured data, and rewriting content into small AI-oriented chunks aren't required for its generative AI features. Following this kind of advice anyway wastes effort that could go toward things platforms actually document.
Mass-producing content to chase fan-out queries. Google's own guidance warns that creating separate near-duplicate pages primarily to manipulate rankings or generative AI responses violates its scaled content abuse policies.
What this means for an individual business
This guide covers platform-level mechanics. For a more practical look at what these mechanics mean for why one specific business might appear in AI search while another doesn't, see How Businesses Appear in ChatGPT and AI Search.
Sources
- Google Search Central — Optimizing your website for generative AI features on Google Search
- Google Search Central — AI features in Search
- OpenAI — ChatGPT Search
- Microsoft Learn — Data, privacy, and security for web search in Copilot Studio
- Microsoft Learn — Use Grounding with Bing Search tools with the agents API
Related
- The Complete Guide to Answer Engine Optimisation (AEO) (guide)
- How Businesses Appear in ChatGPT and AI Search (guide)
- Bing and AI Search: Why Microsoft's Search Engine Still Matters (guide)
- Entity SEO: What Businesses Need to Get Right (guide)
- How do AI platforms choose which websites to recommend? (FAQ)
- How does ChatGPT decide what businesses to mention? (FAQ)