How Do I Tell?
“Different signals than Google SEO.” Is that true?
A vendor pitch landed in front of us with a four-part claim about how LLMs assess businesses. Three of the four parts are doing very little work. The fourth is the one that matters, and it is buried in the middle.
“LLMs evaluate businesses on structured data (JSON-LD), llms.txt files, citation patterns, and content comprehension. Different signals than Google SEO.”
Terms with a dotted underline have a definition attached — hover on desktop, tap on mobile.
Mostly false as stated. One of the four signals holds up under causal testing. The concluding sentence — the part the pitch is really selling — is where it breaks.
Structured data (JSON-LD)
The correlation is real and the causation is not. Ahrefs tracked 1,885 pages that added JSON-LD between August 2025 and March 2026, matched them against 4,000 control pages, and measured citation change on either side of the addition. The result: no meaningful uplift on Google AI Overviews, Google AI Mode, or ChatGPT.
A companion experiment from searchVIU tested what five major systems actually read when fetching a page live. ChatGPT, Claude, Perplexity, Gemini, and AI Mode all extracted only the visible HTML. JSON-LD, hidden Microdata, and hidden RDFa were ignored.
The widely repeated statistic — that AI-cited pages carry schema roughly three times as often as uncited pages — is confounded. Sites that implement structured data also do technical SEO, publish authoritative content, and maintain their pages. Strip out the schema and the rest of those signals still get the page cited.
Schema still earns its place, but through an indirect route: it helps map your business to a Knowledge Graph entity, which strengthens conventional authority signals, which feeds citation eligibility. That is an SEO pathway. It is not a separate AI pathway.
llms.txt
No major engine has formally committed to consuming the format. Adoption sat at 8.7% of the world's top 1,000 sites as of June 2026 — and adoption is not the same thing as effect.
On effect, the sharpest finding comes from SE Ranking, which ran a machine-learning model to test whether the presence of an llms.txt file predicts citation frequency. Removing the llms.txt variable from the model improved its accuracy. The file was contributing noise, not signal.
Publish one if it costs you an hour. Do not pay for it as a strategy, and be alert when a pitch leads with it.
Citation patterns
This one is real, and it is the strongest item on the list. Ahrefs' analysis across 75,000 brands found YouTube mentions correlating with AI visibility at 0.737, branded web mentions at 0.664, and branded anchor text at 0.527 — against backlinks at 0.218.
Read that ordering carefully. The things that predict whether an AI engine talks about you are mentions of your name across surfaces you do not own. Where you get discussed matters more than what you publish on your own domain.
Content comprehension
Vague as written, but pointing the right way. Content organized to answer a specific question — clear headers, a direct answer near the top, self-contained passages, cited primary data — is what these systems select when they summarize. This is closer to good editing than to a technical discipline, which is why it rarely appears on a vendor's invoice.
“Different signals than Google SEO”
This is the load-bearing sentence, because it is the one that implies you need a separate service. It is half true at best, and for some industries it is close to wrong.
The half that is true: the overlap has genuinely loosened. The share of Google AI Overview citations coming from pages in the organic top ten fell from roughly 76% to 38%. A first-page ranking is no longer a complete AI visibility strategy.
The half that is not: when sites lost organic visibility in early 2026, their AI citations declined almost in lockstep. The engines are still drawing from search indexes. Organic collapse and citation collapse move together.
And the part that matters most if you are a professional practice: YMYL sectors — health, finance, legal — show the highest citation-to-top-ten overlap, in the range of 68–75%. The reason is straightforward: the trust signals that drive conventional ranking in those categories are the same trust signals that drive citation selection.
Our ICP sits squarely inside that band. Boutique attorneys, fee-only RIAs, CPAs, and real estate teams are the practices for which “these are different signals” is least accurate. If a vendor tells a family law attorney that AI search runs on a separate signal set, the vendor is describing an average that does not include that attorney.
The tell
Notice the ordering. The pitch opens with the two items that survive causal testing least well and buries the one factor that actually predicts citation in third position.
That ordering is not an accident of drafting. Schema and llms.txt are billable — they are discrete technical deliverables that can be installed, invoiced, and screenshotted. Earning mentions across YouTube, Reddit, LinkedIn, and the trade press is slow, is partly outside anyone's control, and does not photograph well in a monthly report. Pitches tend to lead with what is easy to sell, not with what is most likely to work.
Four questions for the next vendor who says this
- Which engine, specifically? ChatGPT, Perplexity, Gemini, and AI Overviews retrieve differently. A claim that does not name an engine has not been tested against one.
- Is that correlation or causation? If the answer cites how often cited pages have schema, that is correlation. Ask what happened when someone added it.
- What does the evidence say for my vertical? Averages across all industries do not describe a regulated professional practice.
- What is the measurement plan? A baseline before the work, a defined window after, and the same queries run both times. No baseline, no claim.
Evidence base: Ahrefs schema causal study (May 2026, 1,885 treated pages / 4,000 controls); searchVIU live-retrieval extraction test; Ahrefs brand visibility correlation study (75,000 brands); Ahrefs AI Overview citation study (863,000 keywords); SE Ranking llms.txt modeling; BrightEdge sector-level citation overlap analysis.
Findings above are reported as directional where they rest on correlation and verified where a controlled or matched test exists. Where the two disagree, we say so rather than choosing the more flattering number.