Welcome back to our GEO series!
In the first article, we explored why rankings alone no longer explain visibility in AI-powered search. This time, we take a closer look at candidate page selection—a concept we use at Lumar to help explain why some pages become eligible for retrieval while others don’t.
We’ll explore why candidate page selection matters for GEO, how it differs from traditional rankings, and what influences whether a page becomes eligible for retrieval.
Short on time? Here are the key takeaways…
- Candidate selection determines eligibility, not ranking. Before AI systems retrieve content, they first decide which pages are suitable candidates for a specific query.
- Page-level signals come first. Lexical signals, semantic relevance, authority and freshness all help determine whether a page is considered for retrieval before individual passages are evaluated.
- Candidate selection is query-dependent. A page may be a strong candidate for one search but never be considered for another, depending on the user’s intent and the signals the query prioritizes.
- Traditional SEO tools don’t measure this stage. Rankings and traffic show how a page performs once it’s visible, but they can’t tell you whether it was ever considered as source material for an AI-generated response.
- GEO starts with page eligibility. Clear topical focus, well-aligned titles and headings, and strong credibility all help establish whether a page can become a candidate. Optimizing individual passages comes later.
Candidate selection isn’t the same as ranking
Traditional SEO is built around rankings. Once a page is indexed, the goal is to improve its position in search results. Rankings fluctuate over time, but every eligible page is competing for visibility.
Candidate page selection is different. Instead of determining where a page appears, it determines whether that page is considered as a potential source for a particular query.
From what we’ve seen, this stage has traditionally been largely invisible because standard SEO tools don’t measure it. Rankings, keyword positions and traffic show how a page performs in search results, but they can’t tell you whether it was ever considered for an AI-generated response.
That’s why strong rankings don’t always translate into AI visibility. A page may rank highly in traditional search yet never contribute to an AI-generated response if it isn’t selected as a candidate.
Why candidate selection exists
If candidate selection determines which pages are eligible for retrieval, the next question is why AI systems need this extra stage at all.
To generate a coherent response, AI systems need an efficient way to identify useful source material. Evaluating every indexed page or every possible passage for every query wouldn’t be practical, so candidate selection narrows the search space to a smaller pool of trustworthy pages.
That also explains why page-level signals matter so much. Before individual passages can contribute to a response, the page itself needs to establish that it’s relevant to the query and appropriate to retrieve from.
What influences candidate page selection?
Candidate page selection appears to be influenced by a combination of lexical, semantic, authority and freshness signals. Their relative importance changes depending on the query. A breaking news search, for example, is likely to place greater emphasis on freshness, while a medical query may depend more heavily on authority and trust.
Lexical signals
Despite advances in semantic search, traditional keyword signals still play an important role. Exact and near-exact matches in titles, H1s, headings and anchor text provide strong topical cues, helping search systems identify what a page is primarily about.
Semantic relevance
AI systems also appear to evaluate whether a page is genuinely about the topic being searched for, rather than simply containing matching keywords. Pages that explore a subject in depth are more likely to become candidates than those that only mention it in passing.
At this stage, the focus is on establishing broad topical relevance before identifying the strongest individual answer.
Learn more about semantic relevance for GEO/AEO
Authority and trust
Authority also appears to be important during candidate selection, particularly for YMYL (Your Money or Your Life) topics where inaccurate information could have real-world consequences, such as medical, financial or legal advice.
In our experience, official documentation, government guidance, recognized organizations and established publishers are more likely to be considered for these types of queries. For broader informational searches, authority is balanced alongside topical relevance and content quality.
Freshness
Freshness doesn’t influence candidate page selection equally for every query. Evergreen content may remain relevant for years, while news, product updates and other rapidly changing topics often favor more recent sources.
Candidate selection depends on the query
Whether a page becomes a candidate depends on the query being asked. A page may be considered a strong source for one search, but never enter the candidate pool for another.
For example, a general automotive blog might be a suitable source for a query such as “how oil changes work.” A much more specific search, such as “Honda Civic oil drain torque,” is more likely to favor manufacturer documentation or specialist technical resources, where accuracy and precision carry greater weight.
Why traditional SEO tools don’t measure candidate selection
Traditional SEO tools are built around the assumption that indexed pages are eligible to compete for visibility. From there, they measure performance through rankings, keyword positions, backlinks and traffic.
Candidate selection introduces an earlier stage that sits outside those metrics. A page may be fully indexed and perform well in traditional search, yet never become a candidate for a particular AI-generated response.
As a result, existing SEO tools can’t tell you whether a page entered the candidate pool, why it may have been excluded, or which signals influenced that decision. Understanding AI visibility will require ways of measuring performance beyond rankings alone.
What this means for GEO
Candidate selection starts at the page level, so optimization should too. A page with a clear, well-defined topic is easier for AI systems to interpret than one trying to cover multiple unrelated subjects. Aligning titles, H1s and headings with the queries people actually use to search helps reinforce that topical focus.
Once the page’s intent is clear, credibility becomes just as important. For queries where trust carries greater weight, the quality and authority of the source may influence whether the page is considered for retrieval at all.
Only then does it make sense to optimize individual passages. If the page never becomes a candidate, even the strongest content is unlikely to contribute to an AI-generated response.
What comes next
Candidate selection determines which pages are eligible for retrieval. The next stage is deciding which parts of those pages are actually used to generate a response.
In the next article, we’ll look at:
- How AI systems evaluate semantic relevance once a page becomes a candidate
- Why the same content can be interpreted differently at different stages of retrieval, and
- How chunk-level retrieval ultimately determines which passages contribute to an AI-generated response.