The direct answer is that Google AI Overviews chooses sources by running your query through a customized Gemini model, breaking it into smaller sub-questions, and then pulling short passages from pages that answer those sub-questions clearly and credibly. Relevance to the exact sub-question comes first. Trust in the page and the domain comes second. Everything else is a tiebreaker. If you understand how Google AI Overviews chooses sources at that level, most of the tactics stop feeling like guesswork and start looking like a checklist.
The short version: relevance first, then trust
An AI Overview is not a ranked list of ten blue links. It is a synthesized answer stitched together from a handful of passages, each attributed to a source. Google is not asking which page deserves to rank number one. It is asking a narrower question: which passage best answers this specific slice of what the user wants, from a source I can stand behind. That distinction changes the whole game. A page that ranks fourth or fifth for the broad keyword can still be the single best answer to one precise sub-question, and that is often enough to earn a citation.

Relevance to the sub-question is the gate. Trust is what gets you through it once you qualify. Pages that try to be everything to everyone tend to answer no single sub-question cleanly, which is why thin, broad content struggles here even when it ranks acceptably in classic search. The system rewards precision more than reach.
It helps to be precise about what a source even is in this context. Google is not endorsing your whole page when it cites you, and it is not ranking your domain against a competitor’s. It is selecting one specific passage, a sentence or two, to support one specific part of the answer, and attaching your name to that fragment. A single page can supply several passages to one Overview, or a single passage and nothing else, or contribute to the answer without being visibly cited at all. Once you see citation as passage-level rather than page-level, the whole optimization problem sharpens: you are not trying to make a great page, you are trying to make a page full of individually quotable passages.
That reframing also explains why length alone does nothing. A two-thousand-word page that never states a clean answer offers Google no passage to lift, while a shorter page that answers three questions plainly offers three. The system is indifferent to how comprehensive you feel your page is. It cares only whether, somewhere on the page, there is a passage that answers the sub-question better than every other candidate it retrieved. Everything useful you can do to understand how Google AI Overviews chooses sources follows from taking that passage-level view seriously.
Query fan-out is the mechanism most people miss
The piece that trips up most marketers is fan-out. When someone asks a question, Google does not run one search. It generates several related queries behind the scenes, some narrower, some adjacent, and gathers candidate passages for each. The Overview you see is assembled from the winners across all of those hidden searches. This is why a page can appear in an Overview for a term it does not visibly rank for: it won one of the sub-queries, not the headline one.
For anyone trying to understand how Google AI Overviews chooses sources, fan-out is the single most useful concept. It means you are not competing for one keyword. You are competing for the full set of questions a topic implies. A page structured to answer one narrow question in one clean section, then another in the next, gives Google more surfaces to pull from. A page that buries its answers inside long unstructured prose gives Google fewer clean passages to lift, and loses to a competitor that made extraction easy.
The seven signals, ranked by weight
Passage-level relevance sits at the top. Google is matching a specific chunk of your page against a specific sub-question, so the tighter that chunk answers it, the better. Second is a direct-answer structure, meaning the answer appears near the top of a section rather than after four paragraphs of throat-clearing. Third is domain-level trust, the accumulated sense that your site knows this subject, built from topical depth and citations from other credible sites.

Fourth is freshness, which matters more for topics that change and barely at all for stable ones. Fifth is corroboration: Google prefers claims that agree with what other trusted sources say, so a wild outlier claim rarely gets quoted even when it ranks. Sixth is clarity of entity, meaning Google can tell who published the page and connect it to a known organization or author. Seventh is technical accessibility, the unglamorous requirement that the page is indexed, fast, and not blocking snippets. Miss the seventh and the other six never get a chance.
Why does Google demote some pages it clearly indexed?
Plenty of pages rank fine in ordinary search and never appear in a single Overview. The usual reason is that they answer the broad query acceptably but never answer any sub-question cleanly. Google indexed them, considered them, and found nothing quotable. The fix is rarely more content. It is better-structured content: shorter sections, each opening with the answer, each mapped to a question a real person would ask.
The second common reason is a trust ceiling. A page can be perfectly clear and still get skipped because the domain has not earned enough credibility on that topic for Google to attach its name to the passage. An Overview citation is a public endorsement, and the system is conservative about whose words it repeats. Thin topical authority is a slower fix than structure, but it is the one that separates sites that get cited occasionally from sites that get cited constantly.
The Overview Eligibility Ladder
Here is a model worth keeping. Call it the Overview Eligibility Ladder, four rungs a page climbs before it can be quoted. The bottom rung is indexable and accessible: crawlable, fast, snippets allowed. The second rung is relevant at the passage level: at least one section answers a real sub-question directly. The third rung is credible: the domain has enough topical trust for Google to repeat the claim. The top rung is corroborated: the claim aligns with what other trusted sources say, so quoting it is low risk.
Most pages that fail to get cited are stuck on rung two. They are accessible and sometimes credible, but no single passage answers a sub-question cleanly enough to lift. Diagnosing which rung a page is stuck on tells you exactly what to fix, instead of guessing. A page stuck on rung one needs a technical fix. A page stuck on rung three needs authority work. Treating every citation miss as the same problem is why so much AEO effort gets wasted on the wrong lever.
Does structured data change your odds?
Structured data comes up constantly in AEO conversations, usually with more hope attached to it than it deserves. Schema markup does not force an AI Overview citation, and any tool promising that it does is selling you something. What structured data actually does is remove ambiguity, helping Google understand what your page is, who wrote it, and how its pieces relate. On a page that already answers a question clearly, that clarity reinforces the signals that make you citable. On a vague page, no amount of markup rescues it.
The place structured data earns its keep is entity understanding. Marking up your organization, your authors, and your content types helps Google connect your page to a known entity it can trust, which supports the credibility rung of the eligibility ladder. This matters more the less established your brand is, because a newer site benefits from every signal that says who it is and what it knows. For a page competing to be quoted, being unmistakably identifiable is a quiet advantage over an equally good page whose ownership is murky.
Treat schema, then, as reinforcement rather than a lever you pull for its own sake. Get the content clear and the answer direct first, because that is what determines whether a passage is quotable at all. Then add structured data to make sure Google reads the page the way you intend and attributes it to the entity you have built. In that order, markup helps. Reversed, with markup standing in for clarity you never provided, it does nothing, which is why so many schema-heavy pages still never appear in an Overview.
The topical-authority flywheel
The signal that separates occasionally-cited sites from constantly-cited ones is topical authority, and it behaves like a flywheel. When you cover a subject deeply, answering not just the head question but the whole fan-out of sub-questions around it, Google starts to see your domain as a genuine authority on that topic. That perception makes it more willing to quote you, which brings more visibility, which tends to attract more corroboration and links, which deepens the authority further. Each turn of the wheel makes the next citation easier to earn.
The implication is that scattering your effort across many unrelated topics is the slow path, while concentrating it on a focused set of subjects is the fast one. A brand that publishes deeply on three tightly-related topics builds real authority in all three and gets cited across the questions they imply. A brand that publishes one page each on thirty unrelated topics builds authority in none and struggles to get cited anywhere, because Google never accumulates enough signal to trust it on any single subject. Depth compounds. Breadth without depth dissipates.
This is why the smartest AEO strategy often looks narrow from the outside. Picking a lane and owning its full question space, rather than chasing every keyword that shows volume, is what spins the flywheel. Once it is turning, understanding how Google AI Overviews chooses sources becomes almost academic, because your domain has become the obvious source for its topic, and the citations start arriving as a byproduct of authority you have already built rather than a prize you fight for one page at a time.
Where to start this week
Pick your five highest-value questions and check whether each has a page that answers it in the first two sentences of a section. Most brands find the answer is buried or missing. Rewrite those sections so the answer leads and the explanation follows. Then confirm the pages are indexed and snippets are allowed, because the most common silent failure is a page Google would happily quote but technically cannot. Fix eligibility first, structure second, authority third, and you will move more pages into Overviews than any amount of extra publishing would. Knowing how Google AI Overviews chooses sources is only useful if you turn it into that kind of specific, boring, repeatable work.