AEO and GEO still matter — even as the tools built to measure them are failing. The measurement is broken. The work is not. Those are two separate claims, and the evidence supports both simultaneously.

The AI visibility industry has a measurement problem, and everyone is starting to notice. In the last few months, a wave of studies from Semrush, SparkToro, Conductor, and Unusual.ai have all landed on the same uncomfortable finding: the tools brands are paying to track their AI visibility are measuring noise as much as signal. Rankings flip run to run. The same prompt phrased two ways surfaces different brands. Persona context reshapes recommendations. Buyer intent changes everything.

The temptation, watching this unfold, is to conclude that AEO and GEO were a mistake. That "marketing to AI" is a category error, or at best a rebranded version of SEO that inherits all of SEO's problems and adds new ones. Eli Schwartz, on LinkedIn, put it bluntly: GEO "was never alive." His argument has traction because it feels true when you stare at a Profound dashboard that shows your brand at rank 3 on Monday and rank 7 on Wednesday with no content changes in between.

I want to argue the opposite, but carefully. The measurement is broken. The work is not. Those are two separate claims, and the evidence supports both simultaneously.

What Does the New AI Visibility Research Actually Show?

The latest independent research converges on one finding: single-prompt AI visibility tracking measures an artifact of its own design more than it measures real market position.

Start with the paraphrase problem, because it is the most damaging finding for the current AEO tracking industry. In May 2026, researchers at Unusual.ai ran roughly 12,000 API calls across OpenAI and Anthropic models to measure how much a brand recommendation depends on the exact wording of a prompt versus the underlying buyer intent. The results are stark. Two cosmetic rewordings of the same question — "best CRM" versus "top CRM" — produced recommendation sets with a Jaccard similarity of 0.288. Adding a constraint, "CRM for a SaaS startup under 50 people," dropped it to 0.135. The same-prompt rerun baseline, meanwhile, sat at 0.50 to 0.61.

Translated: the prompt string, not the buyer intent, is the dominant input to which brands surface. As the paper puts it, "the prompt string is not a benign label for an intent; it is the dominant determinant of the answer." A tracking tool that issues one canonical prompt per category is measuring a metric whose largest source of variance is the tracker's own phrasing choice.

Rand Fishkin's SparkToro study, run in late 2025 with 2,961 volunteer-driven runs across ChatGPT, Claude, and Google AI, arrived at compatible numbers from a different angle. Fewer than 1 in 100 responses produced the same brand list. Fewer than 1 in 1,000 produced the same list in the same order. Real human users, prompted to ask about the same topic in their own words, generated prompts with a semantic similarity of 0.081. Users don't ask the same question twice. And yet Fishkin, who started the research as a committed skeptic hoping to disprove AI visibility tracking entirely, ended up partially reversing his position. Visibility percentage across many runs, he concluded, is "probably a reasonable metric." Bose and Sony showed up in 55 to 77 percent of 994 headphone responses despite the prompt chaos. There is signal underneath the noise. It just isn't where the trackers are looking.

Conductor's 14,000-run study across 10 industries, 7 intent types, 4 LLMs, and 5 personas added a third dimension. Intent type, not industry size, predicts consistency. Purchase intent produces 40 percent brand overlap between runs — essentially a coin flip. Comparison intent reaches 63 percent with 91 percent lead-brand stability. The ranking of intent types from least to most consistent held across every LLM and every industry they tested.

Then the persona layer. Unusual.ai's persona conditioning paper, also from May 2026, measured what happens when you prefix the same prompt with different buyer contexts. The recommendation set moves by a Jaccard delta of 0.12 to 0.20, a magnitude comparable to switching AI providers entirely. L3 mid-market brands swap up to 75 percent of their recommendation slots as persona changes. L1 category leaders barely move.

Four independent studies. Four different methodologies. All converging on the same finding: single-prompt AI visibility tracking measures an artifact of its own design more than it measures your brand's position in the market.

Why Did AI Visibility Metrics Break?

The old game was different, and understanding why matters. Google's algorithm was stable enough that a tracking tool could report your rank against a keyword and everyone agreed what the number meant. One player, one system, one ranked list. When your rank moved, something had changed.

The new system violates every one of those assumptions. There are multiple AI engines — ChatGPT, Claude, Perplexity, Gemini, Copilot — each with different retrieval substrates, different priors, different training data cutoffs, different reasoning modes. As Chen et al. (2025) documented, AI search services differ significantly from each other in domain diversity, freshness, cross-language stability, and phrasing sensitivity; Conductor tested four of them and found meaningfully different behavioral profiles. None of these engines disclose who they recommended to whom, or why. The generation process is stochastic. And the input space has expanded: the "query" is no longer a keyword but a natural-language question conditioned on the user's persona, prior turns, and account context, most of which the tracking tool cannot see.

The Unusual.ai prominence audit, a 37,000-run study across 215 prompts and 533 brands, adds one more layer. AI recommendation failure modes are not uniform. They differ sharply by where a brand sits on the prominence ladder. L1 category leaders like Salesforce and HubSpot surface in 77 percent of relevant queries but win the recommendation slot only 25 to 41 percent of the time. Their bottleneck is positioning, not discoverability. L4 and L5 long-tail brands never surface in 48 to 52 percent of runs. Their bottleneck is pure Stage-1 invisibility. L3 mid-market brands fail at every stage simultaneously.

A single tracking metric cannot capture this. A brand at L1 and a brand at L4 could show identical "visibility scores" in a monitoring dashboard while facing completely different problems that require completely different investments.

Why AEO and GEO Work Is Still Valid

Here is what the same body of research says about the supply side of the equation — where AEO and GEO work actually operates.

The original GEO paper from Princeton and IIT Delhi, published at KDD 2024, ran 10,000 queries and found that content-side interventions produce measurable visibility lifts. Adding quotations lifted position-adjusted word count visibility by 41 percent. Adding statistics, 31 percent. Citing sources, 27 percent. Keyword stuffing, the dominant traditional SEO tactic, produced a negative 8 percent effect. Critically, lower-ranked pages benefited far more than top-ranked pages: citing sources delivered a 115 percent lift to rank-5 pages versus a negative 30 percent effect on rank-1 pages.

Chen et al. (2025) documented what they called a "systematic and overwhelming bias" in AI search toward earned media — third-party, authoritative sources — over brand-owned and social content. This is not the same signal as Google ranking. As Profound's team observed in their platform data, there is "surprisingly small overlap between Google and ChatGPT" citation patterns.

Semrush's study of 50,000 brands across 1,094 U.S. ChatGPT categories produced perhaps the most strategically important number in the entire dataset: 53.7 percent of categories have no clear brand leader. In the highest-demand topics, accounting for 98 percent of AI search volume, only 11.3 percent have a clear owner. And branded search volume, not organic traffic or domain authority, was the only traditional SEO metric that correlated meaningfully with AI category ownership.

Put those findings side by side. Content interventions produce real, measurable lifts. Earned media is dramatically over-weighted by AI retrieval. Most categories are still up for grabs. Once a brand achieves clear category ownership, it retains first place in over 90 percent of month-over-month comparisons. And the traditional SEO signals many brands are still optimizing for are not the signals that predict AI visibility.

The work is valid. It just isn't the work the trackers are measuring.

What Effective AEO/GEO Strategy Actually Looks Like Now

The prominence-stratified audit gives us a concrete framework. If you're an L1 category leader, you appear in nearly every relevant search but win the recommendation slot less than half the time. Your marginal investment should go to differentiation content, direct comparisons against named peers, and consistency across the authority sources AI retrieves from. Discoverability is solved; positioning is the lever.

If you're L4 or L5, roughly half your peers never surface in 37,000 runs. The problem is that AI cannot find you at all. 60 percent of L4 brands that do surface appear only via external retrieval systems like Exa or Brave, never via the provider's native web search. Investment goes to authority-list seeding, third-party comparison articles, and canonical content on platforms neural retrieval systems reach. Positioning matters, but it's upstream-bottlenecked by the retrieval step.

L3 mid-market brands face all three failure modes at once. 12 percent never surface at all, Stage-4 conversion drops to 34 to 40 percent, and persona effects peak at 75 percent recommendation swap rates. Single-channel investment at L3 is bounded by the other two channels. The work compounds only when applied together.

Above the tier-specific work, some prescriptions apply universally. Chen et al.'s finding that earned media dominates suggests that the highest-leverage single investment for most brands is not on-site content but off-site authority: getting mentioned in the third-party sources AI retrieves from. The strategy, as the Profound team frames it, is to become the trusted answer to questions wherever those questions are asked.

Conductor's intent taxonomy gives a second layer of prioritization. Comparison-intent queries produce 91 percent lead-brand stability once a brand is established as the answer. Purchase-intent queries are the least consistent (40 percent overlap) and the most contestable. Education queries produce zero-brand responses 45 to 72 percent of the time depending on the LLM, but they build the foundational authority that drives visibility everywhere else. A brand can't compete on every intent type at once. The intent-stratified data tells you which fights are winnable at your current position.

The Schwartz Counter-Argument: Where Does It Break Down?

Eli Schwartz's core claim deserves a direct response. His argument is that AI visibility is perfectly correlated with traditional search visibility, so GEO is "SEO with a value-added tax." If true, the entire AEO industry is redundant.

The empirical record doesn't support this. Semrush found that Authority Score correlates with AI category ownership at 52.5 percent, essentially chance. Organic traffic performs no better at 48.4 percent. Only branded search volume (55.7 percent) shows a statistically meaningful relationship, and even that is weak. Only 21 percent of the most-cited domains in a ChatGPT category are also the most-mentioned brands. The two signals measure different things.

The Unusual.ai prominence audit is the more direct refutation: 48 to 52 percent of long-tail brands sourced from established authority lists — Y Combinator batches, national business registries, Crunchbase mid-stage filters — never surface in 37,000 AI runs. If Schwartz's correlation held, this couldn't happen. AI retrieval systems carry different priors than Google's ranking algorithm, weight earned media differently, and consolidate around a narrower set of authority sources. As Mostafa ElBermawy noted in the LinkedIn thread pushing back on Schwartz, LLMs are "making decisions and taking actions on our behalf," not just retrieving links. The functional relationship between brand and AI is categorically different from brand and search engine.

Schwartz is right that SEO fundamentals matter. Crawlability, structured data, E-E-A-T signals, authority — all of these carry over. What he misses is that they are necessary but not sufficient. Making a brand extractable, comparable, and citable in AI-mediated recommendation contexts requires its own strategic frame.

How Should Brands Measure AI Visibility When Prompt Tracking Fails?

If prompt-by-prompt tracking is broken, what should brands measure instead?

The Unusual.ai paraphrase paper is explicit on this. Downstream-of-mention metrics — conversion, qualified pipeline, revenue attributable to AI-surface exposure — aggregate over the buyer's natural paraphrase distribution by construction. Each buyer issues their own phrasing. The paraphrase-choice artifact that dominates single-prompt tracking dissolves at the outcome layer.

The NerdWallet case study reported by CXL is instructive: 35 percent revenue growth despite a 20 percent site traffic decline. The company optimized to be cited as the answer source across AI platforms rather than to drive clicks. Traffic and revenue decoupled. Revenue, not traffic, is the metric that reflects whether AI-mediated visibility is working.

Below that, if you must track something at the mention layer, visibility percentage across many prompts is Fishkin's least-bad option. Not rank, which is meaningless. Not any single prompt, which inherits the paraphrase artifact. Aggregate visibility across a distribution of buyer-realistic phrasings, tracked over months rather than weeks, with confidence intervals that acknowledge how noisy the underlying signal is.

The Semrush concept of topic-level visibility, measured across a full buyer journey (definition, comparison, alternative, use case, purchase), is a stronger unit than any single prompt. Conductor's intent-stratified reporting adds another useful axis. Neither is a silver bullet. Both are meaningfully better than what the current generation of trackers reports.

The AEO/GEO Window Is Open — and It Will Close

Category positions in AI search are being claimed right now, by brands doing the work right now. The window is open. It won't stay that way.

One finding from the Semrush data deserves its own moment. Only 15.2 percent of ChatGPT categories have a clear brand leader today. In the categories that account for 98 percent of AI search volume, only 11.3 percent have a clear owner. Once a brand achieves clear ownership, it retains first place in over 90 percent of month-over-month comparisons.

The competitive dynamic is winner-take-most, but 84.8 percent of the race hasn't been run yet. Brands that invest in AEO/GEO fundamentals now — earned media presence, third-party authority, structured extractable content, tier-appropriate positioning — are competing for durable positions in a market where the leaders haven't been selected. Brands that wait for the metrics to become reliable will be waiting while those positions get claimed.

The measurement infrastructure will improve. Someone will figure out how to sample the paraphrase distribution efficiently, integrate persona conditioning, and produce trend data that isn't dominated by tracker-choice artifact. That's a solvable problem, and the academic literature is starting to solve it. But the strategic window is not the measurement window. Category ownership is being established now, on the current generation of models, by brands doing the work now.

The uncomfortable truth for the AEO/GEO industry is that the trackers many brands are paying for are unreliable in ways their vendors haven't disclosed. The equally uncomfortable truth for the skeptics is that the underlying strategic imperative is real, urgent, and empirically supported. Content interventions work. Earned media dominates. Prominence tier dictates strategy. Persona segments recommendations. Intent type predicts consistency. Every one of these is actionable. None of them requires you to trust a single-prompt visibility score.

What's changed is not whether marketing to AI matters. What's changed is that the shortcut of watching a dashboard has stopped working, and the actual work — which was always the work — is what remains. Positioning against named competitors. Getting cited in the sources AI retrieves from. Building segment-specific content for the personas your buyers actually inhabit. Winning on the intent types where your category is contestable. Measuring outcomes at the revenue layer, not the mention layer.

The metrics are broken. The work is not. At B2X Marketing, that's the distinction we keep coming back to — and anyone selling you the opposite combination is selling you something you shouldn't buy.

Sources