Short answer: AI engines cite pages that already rank, put the answer in the first third of the page, were published or updated recently, include a comparison table, quote a source or a statistic, and belong to a brand other websites already mention. Structured data did not change citation rates in the one controlled test anyone has run; llms.txt is almost never fetched. What structured data does do is keep the engines from confusing you with someone else, which matters more the more locations you have. And the part most companies skip is measurement: without crawler logs and referrer tracking you cannot tell whether any of this is working. That measurement, plus the entity layer, is what our MawGeo product is; the content advice below is what we tell every client to do themselves.

What the studies agree on

Three large 2026 analyses, run by companies with different things to sell, reached compatible conclusions.

Signal Finding Source
Ranking Position 1 is cited 55% of the time, position 9 34%; cited pages average rank 4.24 TripleDart, 11,002 pages
Answer placement 64% of citations come from the first 40% of the page; median cited passage sits 27% down TripleDart
Freshness 77% of cited pages were published or updated in 2024 or later; 2018 pages cited 26% vs 50–54% for 2023+ TripleDart
Comparison tables Present on 23% of cited pages vs 17.8% of skipped: the only on-page feature that survived statistical controls TripleDart
Length Cited pages were shorter on average (2,924 words vs 3,055) TripleDart
Brand mentions Correlate 0.664 with AI Overview visibility; backlinks only 0.218 Ahrefs, 75,000 brands
Quotations and statistics +41% and +33% generative-engine visibility; keyword stuffing −9% Princeton GEO-bench
Schema (FAQ) 19.5% on cited pages vs 19.3% on skipped: no difference TripleDart

Sources: TripleDart, Ahrefs brand study, Aggarwal et al., KDD 2024. The TripleDart data is B2B SaaS, so treat the exact percentages as directional for other sectors; the direction has been replicated.

Two additions worth knowing. Ahrefs found that by March 2026 only 38% of AI Overview citations came from the top ten results, down from about 76% the previous summer, with YouTube the most-cited domain (Ahrefs); ranking helps, but it is no longer the whole game. And Semrush’s study of 600,000 ChatGPT citations found 85% of categories have no clear "owner," and that the most-cited domain was also the most-mentioned brand only 21% of the time (Semrush). Most niches are still open.

The schema question, honestly

Here is the study that made us re-read our own product page. In May 2026 Ahrefs tracked 1,885 pages that added JSON-LD structured data against 4,000 matched controls over 30 days. Google AI Overview citations fell 4.6% (statistically significant); AI Mode and ChatGPT moved +2.4% and +2.2%, indistinguishable from noise (Ahrefs). Their explanation for why schema correlates with citations (cited pages are nearly three times as likely to have it) is the obvious one: "the sites that add structured data tend to also invest in technical SEO, publish authoritative content, build links, maintain their pages, and rank well." A separate experiment cited in the piece found that ChatGPT, Claude, Perplexity and Gemini all read only the rendered HTML during a live fetch and ignored JSON-LD entirely.

Caveats, which Ahrefs states itself: every test page was already heavily cited, the window was 30 days, and the study cannot say whether schema helps pages that are not yet visible. A separate correlational study found schema pages 2.3× more likely to be cited (The Stacc), which is consistent with Ahrefs’ correlation and does not contradict its controlled result.

So what is structured data for? Entity clarity and the knowledge graph. Google’s own documentation is where Organization and LocalBusiness markup earn their keep: rich results, the Knowledge Panel, and, for anyone with more than one location, making sure every branch is understood as a branch of the same company rather than a competitor with a similar name (Google: LocalBusiness, Google: Organization). The pattern that works is one Organization node at brand level and one LocalBusiness node per location page, linked back to the parent, with the same name, address and phone everywhere, including your Google Business Profile, which requires exactly that consistency (Google Business Profile guidelines). That is a prerequisite for being recommended correctly, not a lever for being recommended more often. We should say so plainly, and now we do.

llms.txt, briefly

Skip the debate: 28% of 137,000 domains surveyed by Ahrefs publish an llms.txt; 97% of those files received zero requests in May 2026, and of the requests that did arrive, 1.1% came from AI retrieval bots (Ahrefs). Google’s John Mueller: "no AI system currently uses llms.txt," restated in June 2026 as "purely speculative for now" (SEJ). We ship one because it costs nothing and some tools read it. We do not count it as work.

Which crawlers to allow

This one matters and is widely misunderstood. The AI companies run separate bots for training and for search, and blocking the first does not remove you from the second (OpenAI, Anthropic, Perplexity, Google):

You want to Allow Blocking is your call
Appear in ChatGPT search answers OAI-SearchBot GPTBot (training)
Appear in Claude’s web answers Claude-SearchBot ClaudeBot (training)
Appear in Perplexity PerplexityBot (does not train models) —
Appear in Google AI Overviews Googlebot (there is no separate opt-out without leaving Search) Google-Extended (Gemini training and grounding only)

The user-triggered fetchers (ChatGPT-User, Perplexity-User, Claude-User) generally ignore robots.txt because a person asked for the page, and they are the most interesting line in your logs: each hit means a human saw your page inside an answer at that moment.

Is any of this worth it yet?

Volume: small. AI assistants send roughly 2.6% of what organic search sends, across a 95-site benchmark (SEO Works). Quality: higher, but the headline figure is one company’s data. Ahrefs found AI search was 0.5% of its visits and 12.1% of signups, about 23× organic’s conversion rate (Ahrefs); the broader benchmark found AI sessions converting at 7.18% against 6.08% for organic, with B2B services at 1.21× and B2B products below parity. Growth: real. Similarweb counted 9.5 billion monthly visits to generative-AI platforms in mid-2026, up 70% year on year, and the share of ChatGPT answers carrying an external link rose from 1.6% to 6.8% in a year (Similarweb). For B2B companies, the homepage took 31.9% of LLM citations and 81.7% of AI-attributed leads in one study (TripleDart), which argues for getting the homepage’s first screen right before anything else.

What to actually do

  1. Put the answer first. Every page that targets a question should answer it in the first 150 words, then elaborate. This post does.
  2. Add a comparison table wherever a comparison is honest. It is the one on-page feature with evidence behind it.
  3. Quote sources and cite numbers. Named statistics and quotations lifted visibility by a third or more in the Princeton benchmark; vague claims did nothing.
  4. Refresh dated pages. Three-quarters of cited pages carry a date within two years. Update the content, not just the date.
  5. Get mentioned elsewhere. Brand mentions on other sites correlate three times more strongly with AI visibility than backlinks. Partner pages, directories, press, podcasts, community posts.
  6. Fix your entity. Organization and LocalBusiness markup, consistent name/address/phone, one node per location. Not for citations; for being the right company when you are cited.
  7. Allow the search bots listed above, whatever you decide about training bots.
  8. Measure it. Log every AI user agent server-side, verify against the published IP ranges (guide), and build a GA4 channel for referrers matching chatgpt.com|perplexity.ai|claude.ai|gemini.google.com|copilot.microsoft.com (guide). Google AI Mode strips referrers and shows as direct; accept the gap.

Items 6, 7 and 8 are what MawGeo installs and runs for a site; items 1 to 5 are editorial work, which is why our Granite writes answer-first posts with tables, sources and a refresh schedule. If you want to see where a site stands today, the free audit checks the crawler rules, the entity markup and whether anything in the logs says an AI has ever sent you a visitor.