Answer Engine Optimization Challenges: Why Rankings Stopped Predicting Citations


Catalin Avatar

|

Updated:

Every marketing team we talk to has the same question: why does a page that ranks in the top three get ignored by ChatGPT?

The honest answer is that ranking and citation are now two separate races, judged by different referees, scored on different sheets.

Answer engine optimization sounds like a rebrand of SEO. It is not. It is the same goal reached through a different mechanism, and that mechanism breaks most of the assumptions that made SEO manageable: stable rankings, visible queries, click attribution, and one dominant engine to satisfy.

We have run content properties since 2007 and currently operate more than ten publishing sites.

That gives us something most agencies discussing AEO do not have: a live feed of how citation behavior actually shifts, on our own domains, week after week. What follows is not a theory of AEO. It is the list of problems that show up repeatedly when a startup tries to build visibility inside AI answers, and what we do about each one.

Key Takeaways

  • Ranking no longer predicts citation. Ahrefs found roughly 11% to 12% overlap between the URLs AI assistants cite and the top ten results for the same query.
  • Each engine is effectively a different search engine. Cross-platform studies repeatedly land near 11% domain overlap between ChatGPT and Perplexity.
  • Answers are probabilistic, so there is no position to hold. The same prompt, run repeatedly, returns different brands in different orders.
  • Measurement is improving but still partial. Google’s Search Console generative AI report shows impressions, not clicks, and never shows you the prompt.
  • Most citation real estate belongs to somebody else. Third-party sources, not your own domain, carry the majority of brand citations.
  • The fix is structural, not tactical. Entity clarity, extractable answers, and third-party presence beat volume every time.
FigureWhat it measuresSource
~11%Average overlap between AI citations and the search top tenAhrefs, 15,000 queries
76% to 38%Drop in AI Overview citations coming from top ten ranked pages, mid 2025 to early 2026Ahrefs
~68%Share of US Google searches ending without a click, early in the yearSparkToro
0 clicksClick data included in Google’s new generative AI performance reportGoogle Search Central

Why AEO is harder than SEO ever was

SEO was difficult, but it was legible. You could see the ranking, the query, the click, and the competitor sitting above you. Every one of those four feedback loops is either weakened or gone in AI search.

An answer engine does not rank documents for a user to choose from. It retrieves candidate sources, extracts claims from them, synthesises a single response, and attaches citations to justify it.

Your page is not competing for a slot. It is competing to be the source a model finds most convenient to quote, and convenience here is a technical property, not a marketing one.

That distinction explains almost every frustration below.

What SEO taught youWhat AI search actually does
Rank the page, earn the clickCite the claim, often without any click
One engine sets the standardFive or more engines with different indexes and preferences
Position is stable and trackableOutput is sampled from a probability distribution
Query data appears in Search ConsolePrompts are private and far longer than keywords
Authority is built with backlinksAuthority is inferred from mentions, consensus, and entity clarity
Your domain is the destinationThird-party sources carry most of your brand citations

01. Rankings no longer predict citations

This is the challenge that catches teams off guard, because it invalidates the one dashboard they trust. Ahrefs ran 15,000 long tail queries through Google and Bing, then asked the same questions to ChatGPT, Gemini, Copilot, and Perplexity.

The average overlap between what the assistants cited and what ranked in the top ten was around 11%. Perplexity was the outlier at roughly one in three, and it runs its own crawler and index.

Google’s own AI surfaces are drifting the same way. Ahrefs research across a large keyword sample found the share of AI Overview citations drawn from top ten ranked pages fell from roughly 76% to roughly 38% within about eight months.

Other trackers put the figure lower still. The direction is not in dispute even where the exact number is.

The practical consequence: a page one ranking is now a weak signal of AI visibility, not a guarantee of it. You can dominate a keyword and be structurally absent from the answer that keyword now triggers.

What we do about it

  • Audit citation presence and ranking presence as two separate inventories, never one blended report.
  • Prioritize pages that rank well but are never cited. These are the cheapest wins, because the authority already exists and only the structure is failing.
  • Rebuild those pages so the answer to the core question appears early, self-contained, and quotable rather than buried under a narrative build-up.

02.Every answer engine is a different search engine

Teams tend to optimize for whichever assistant their CEO uses. That is a strategy for winning one fifth of the surface area. Independent studies keep converging on roughly 11% domain overlap between ChatGPT and Perplexity, and the pattern holds across other engine pairings.

Google’s own AI Overviews and AI Mode cite substantially different URLs even when the answers agree.

The reason is architectural. Engines differ in which index they retrieve from, how aggressively they weight freshness, how many sources they pull per answer, and what source types they trust.

Assistants built on Bing behave like Bing. Assistants running proprietary crawlers behave like nothing else. Some lean heavily on brand-owned pages, others on third-party directories and community discussion.

The trap

Single-platform tracking gives you a confident-looking dashboard that describes a fraction of reality. You can be dominant in one engine and completely absent from another, and a one-platform reading will never surface that gap.

What we do about it

  • Pick two or three engines that map to real buyer behavior in the category, then measure each separately rather than averaging them into a vanity score.
  • Treat the shared requirements (clear entity definition, extractable claims, credible sourcing) as the base layer that transfers across every engine.
  • Treat platform-specific patterns as a thin layer on top, and expect to revisit it often.

03.There is no position to hold, only a probability to raise

SEO rewards a stable artefact: the ranking. AI answers are generated by sampling, which means identical prompts return different brand lists, in different orders, with different numbers of results.

SparkToro’s research into AI brand recommendations documented exactly this inconsistency across repeated runs.

The citation layer is even more volatile than the recommendation layer. Analysis published by Growth Memo with AirOps, covering hundreds of thousands of prompt and page pairs, found that running the same ChatGPT prompt three times leaves only a small fraction of the original citations intact. A single run is a sample of one, and reporting it as a result is closer to superstition than measurement.

This does not make AEO unmeasurable. It makes it statistical rather than positional. Weather forecasting is probabilistic too, and nobody argues it is therefore pointless.

What we do about it

  • Run every tracked prompt multiple times on a fixed schedule and report a frequency, not a snapshot.
  • Report inclusion rate (how often the brand appears across runs) separately from ordering, because the two move independently and inclusion is the one you can influence.
  • Set change thresholds in advance so nobody panics over movement that sits inside normal sampling noise.

04.You cannot see the prompt

Keyword research worked because queries were short, repetitive, and visible. Prompts are none of those things. They are several times longer than search keywords, phrased conversationally, and shaped by whatever the assistant already knows about the user.

Two people asking the same underlying question rarely type the same string, and no platform exposes the prompt to you.

Recent work adds another wrinkle: persona conditioning. The recommendation set can shift based on what the model infers about who is asking, which means your visibility varies by audience segment in a way you cannot observe directly.

What we do about it

  • Build prompt sets around buying intents rather than phrasings. One intent, several natural variants, tracked as a group.
  • Mine real language from sales calls, support tickets, and community threads instead of keyword tools, since that is closer to how people actually prompt.
  • Cover the full question tree around a topic, not just the highest-volume node, because the assistant is matching intent rather than string.

05.Attribution breaks down at exactly the wrong moment

Zero-click behavior was already climbing before generative answers arrived. SparkToro’s clickstream analysis put the share of US Google searches ending without a click at roughly 68% early in the year, and the rate is markedly higher on queries where an AI answer appears.

The visibility that replaces those clicks is real, but it lands in places your analytics cannot see.

A prospect reads your comparison table inside an AI answer, forms a view of your product, and arrives three weeks later as direct traffic or a branded search. The influence happened where you have no instrumentation.

Reporting improved in June, when Google launched dedicated generative AI performance reports in Search Console. They separate impressions inside AI Overviews, AI Mode, and Discover from standard organic data for the first time. They also stop well short of closing the gap.

What the new Google report gives you, and what it withholds

You get: impressions, pages, countries, devices, and dates for your appearances in Google’s AI features.

You do not get: clicks, click-through rate, prompt data, or position. Access rolled out gradually, and none of it covers ChatGPT, Perplexity, or Claude. Google shipped an opt-out toggle alongside it, which turns exposure in AI answers into a deliberate editorial decision rather than a default.

What we do about it

  • Instrument AI referral traffic as its own channel and judge it on conversion quality, not volume. The volume is small and the intent is unusually high.
  • Track branded search and direct traffic as downstream indicators of AI-assisted discovery, and read them alongside citation frequency.
  • Set expectations with leadership early: the reporting will stay incomplete, so the operating question is whether citation share is rising, not whether every touch can be attributed.

06.Most of the citation surface is not yours to win

Look at what actually gets cited in a category and a pattern appears fast. A large share of the most-cited pages sit on sources no amount of content strategy will get you into: encyclopedia entries, government and academic domains, app stores, and major news outlets.

The rest is dominated by third-party comparison sites, review platforms, directories, and community discussion.

AirOps research indicates brands are cited through third-party sources several times more often than through their own domain. That is uncomfortable for teams whose entire AEO plan is publishing more blog posts on their own site.

It is also the most actionable finding on this list, once you accept it. Citations beat mentions, and other people’s pages are where most of your citations live.

What we do about it

  • Map the specific third-party pages that already get cited for your priority intents, then work to be accurately represented on them.
  • Prioritize inclusion in category roundups, comparison pages, and review platforms over another owned blog post covering the same topic.
  • Publish original data, because proprietary statistics get quoted by the third parties that engines then cite. The Princeton and IIT Delhi research on generative engine optimization found that adding statistics and citations from credible sources measurably raised a source’s share of the generated answer.

07.Your entity is defined by everyone except you

Before an engine can recommend you, it has to resolve what you are: the category you belong to, who you serve, what you compete with, and whether the various references to your name across the web describe the same organization.

Get that wrong and you are not losing a ranking battle, you are absent from the consideration set entirely.

Ahrefs’ analysis across tens of thousands of brands found web mentions correlated with AI citation rates far more strongly than backlinks did. Not links. Mentions. Being talked about consistently, in language that resolves cleanly to one entity.

Startups suffer here disproportionately, because their positioning changes faster than the web’s description of them updates. Repositioning your homepage does not reposition your entity.

What we do about it

  • Force one consistent category description across the site, schema markup, funding databases, review profiles, social bios, and executive presence.
  • Name the competitive set explicitly on your own pages, because engines learn category membership partly from how brands are grouped.
  • Treat entity consistency as an ongoing maintenance job rather than a one-off audit, especially after a pivot or a rebrand.

08.Models remember an older version of your company

Retrieval softens this problem but does not remove it. When an engine answers from training data rather than live retrieval, it can reproduce pricing you retired, features you replaced, or positioning you abandoned two funding rounds ago.

For fast-moving startups the gap between reality and the model’s default understanding can be substantial.

The failure is quiet. Nobody files a support ticket to tell you an assistant quoted your old plan tiers to a prospect who then chose a competitor.

What we do about it

  • Run factual accuracy prompts alongside visibility prompts. Ask each engine what your product costs and who it is for, and log the wrong answers.
  • Publish clearly dated, unambiguous canonical facts on pricing, positioning, and integrations, so retrieval has something better to find than a cached third-party summary.
  • Correct the third-party pages carrying stale information, since those are frequently the retrieved source rather than your own site.

09.The tooling sells more certainty than it has

The AI visibility tooling market matured quickly and confidence outran methodology.

Dashboards report a single visibility score without disclosing how many runs produced it, which engines were sampled, which geography was used, or what the confidence interval looks like. Two tools measuring the same brand in the same week routinely disagree.

The same inflation affects the advice layer. Tactics circulate as settled fact long before anyone has tested them at scale. Nobody can guarantee placement in an AI answer, and any vendor promising it is selling something other than AEO.

A worked example

Take the llms.txt file. It is cheap to add and harmless to maintain, and it may matter for agentic browsing later. But it is not currently a citation-driving tactic, and Google has been clear it does not use it for search. Treating it as low-cost insurance is reasonable. Treating it as an AEO strategy is not.

What we do about it

  • Ask any vendor for run counts, sampling method, and engine coverage before trusting a number in a board deck.
  • Validate tool output against manual spot checks on prompts that matter commercially.
  • Separate tested tactics from plausible ones in your roadmap, and size the budget accordingly.

The nine challenges, and where to start on each

ChallengeRoot causeFirst move
Rankings do not predict citationsRetrieval differs from rankingAudit ranked-but-never-cited pages first
Engine fragmentationDifferent indexes and source preferencesMeasure two or three engines separately
Output volatilityAnswers are sampled, not rankedMultiple runs, report inclusion rate
Invisible promptsLong, conversational, private queriesTrack intents, not phrasings
Broken attributionZero-click answers and delayed visitsSeparate AI referral channel plus branded search
Third-party citation dominanceEngines prefer independent sourcesMap and earn placement on cited pages
Weak entity resolutionInconsistent descriptions across the webStandardize category language everywhere
Stale model knowledgeTraining lag and cached summariesRun factual accuracy prompts monthly
Overconfident toolingUndisclosed methodologyDemand run counts and sampling detail

What to measure instead of rankings

Most reporting problems in AEO come from importing SEO metrics wholesale. These four hold up better under scrutiny.

MetricWhat it tells youWhere it comes from
Inclusion rateHow often you appear across repeated runs of a tracked intentPrompt tracking with a fixed run schedule
Citation shareYour share of cited sources versus named competitorsPrompt tracking plus manual verification
AI feature impressionsWhether Google’s AI surfaces are showing your pages at allSearch Console generative AI report
AI referral qualityWhether the small volume that arrives actually convertsAnalytics with AI sources isolated as a channel

The honest summary

AEO is not harder than SEO because the tactics are more complex. Structuring a page for extraction is simpler than most technical SEO work. It is harder because the feedback loops are worse.

You are optimizing a probabilistic system, across several engines that disagree with each other, using measurement that omits the most important variables, for outcomes that often arrive without a click.

That is precisely why the opportunity is still open. Most competitors are running SEO reports and calling it AEO.

The teams building durable AI visibility are doing unglamorous work: fixing entity consistency, restructuring pages so answers are extractable, earning presence on the third-party sources engines actually trust, and measuring frequency instead of position.

None of that is a growth hack. All of it compounds.

Want to know where you actually stand?

We map citation presence across the engines that matter for your category, identify the gaps that are costing you consideration, and build the structural fixes that close them.

Become the source AI systems cite, rather than the brand they occasionally mention.

Frequently Asked Questions

Is AEO replacing SEO?

No. AEO builds on SEO rather than replacing it. Crawlability, indexation, site performance, and topical authority still determine whether your content is available to be retrieved in the first place. What changes is the layer above that: content has to be structured so a model can extract a self-contained claim from it, and your brand has to be described consistently enough for an engine to resolve what you are. Teams that abandon SEO fundamentals in favour of AEO tactics usually get worse at both.

Why does my competitor appear in ChatGPT when we outrank them on Google?

Usually one of three reasons. They are mentioned more often across third-party sources that the engine trusts, their content answers the question in an extractable format while yours buries it, or their entity is described more consistently across the web. Ranking position and citation likelihood are only loosely correlated, so outranking a competitor tells you very little about relative AI visibility.

How long does AEO take to show results?

Structural fixes on pages that already rank tend to move fastest, sometimes within weeks, because the authority already exists and only the format was blocking extraction. Entity consistency work and third-party presence take longer, typically a few months, because they depend on external sources updating. Anyone quoting a fixed timeline is guessing, and anyone guaranteeing placement is selling something unreliable.

Should I block AI crawlers to protect my content?

It is a genuine trade-off rather than an obvious call. Blocking protects content from being summarized without attribution, but it also removes you from the answers your buyers are reading. Google now offers a toggle to exclude content from AI features without affecting organic rankings, which makes the decision more granular than it used to be. For most startups seeking visibility rather than ad revenue, exclusion costs more than it protects.

Do I need a separate AEO tool, or is Search Console enough?

Search Console’s generative AI report covers Google’s AI surfaces only, and it reports impressions rather than clicks or prompts. If a meaningful share of your buyers research in ChatGPT, Perplexity, or Claude, that report will not see them. A dedicated prompt tracking setup is worth it once AI-influenced discovery affects real pipeline, provided you interrogate the vendor’s methodology before trusting the numbers.

Does llms.txt help with AI citations?

Not meaningfully, based on current evidence. Google has stated it does not use llms.txt for search, and there is no reliable evidence it drives citations in the major assistants today. It is cheap to implement and may become relevant as agentic browsing matures, so treat it as low-cost insurance rather than a visibility tactic, and do not let it displace structural work.

Sources referenced: Ahrefs AI search overlap study (15,000 queries); Ahrefs AI Overview citation analysis; SparkToro research on AI recommendation consistency and zero-click search; Google Search Central announcement on Search Generative AI performance reports; Growth Memo and AirOps prompt tracking analysis; Princeton and IIT Delhi generative engine optimization research; AirOps State of AI Search research. Figures reflect the most recent published analyses available at the time of writing and should be re-verified before reuse, since citation behavior shifts quickly.

More Articles

  • 1 minute

    Answer Engine Optimization Challenges: Why Rankings Stopped Predicting Citations

    Every marketing team we talk to has the same question: why does a page that ranks in the top three get ignored by ChatGPT? The honest answer is that ranking and citation are now two separate races, judged by different referees, scored on different sheets. Answer engine optimization sounds like a rebrand of SEO. It…

    Catalin Avatar
  • 1 minute

    Ecommerce AEO Strategy: How to Get Your Products Recommended by AI

    The shelf moved. For twenty years, ecommerce visibility meant winning a ranked list of ten blue links and then converting the click. Now a growing share of product discovery happens inside an answer: a shopper describes a problem, an AI system assembles a shortlist, and three or four brands get named. Everyone else is invisible,…

    Catalin Avatar
  • What 65 Million AI Crawler Requests Reveal About AI Visibility (and the 11 Things to Do About It)
    35 minutes

    What 65 Million AI Crawler Requests Reveal About AI Visibility (and the 11 Things to Do About It)

    We logged 65,116,344 requests from seven AI crawlers across 38 of our own websites between October 2025 and July 2026. The headline result upends the popular story. On a normal content site, OpenAI now reads pages into live ChatGPT answers roughly as often as it crawls them for training. Most published AI crawler statistics are…

    Borja Avatar