Two different kinds of trust signal get lumped together constantly in AEO advice, and they don't carry equal weight: what's actually written on the page, and what's true about the domain publishing it. This page covers the first kind - specificity, entity clarity, freshness, and how the content itself earns credibility. A companion page covers the second kind - backlinks, brand reputation, and domain-level reputation. One honest note before any of it: this is a genuinely mixed evidence base, not a settled one, and the biggest research finding below is a peer-reviewed study contradicting a popular claim rather than confirming it.
Does Adding Statistics and Citations Actually Improve Citation Odds?
The evidence is real but genuinely mixed, not a settled "yes." A 2024 benchmark found that rewriting content to add quotations, statistics, or source citations increased a source's share of cited text in its testbed - a GPT-3.5-based setup measuring a fixed set of five Google-ranked sources. That result gets cited constantly as proof that adding statistics works. What usually goes unmentioned: a peer-reviewed NeurIPS 2025 benchmark tested closely related rewrites across six domains and four newer models (GPT-4o-mini, Claude 3.5 Haiku, o3, o4-mini) and found most rewrites did not reliably improve citation rank - and specifically, a statistics-focused rewrite significantly decreased ranking in 19 of 24 tested settings. Citation-only rewrites were negligible; quote-based rewrites gave minimal improvement. The researchers' own conclusion: these methods are "not only largely ineffective but can actually have the opposite expected effect."
Both studies are real and both are worth knowing - they just don't agree, and pretending otherwise misrepresents the evidence. The honest takeaway isn't "add statistics" or "don't add statistics." It's that there's no universal rewrite recipe that reliably works across models and domains. The more durable, defensible reason to write with specific, accurate evidence is editorial, not algorithmic: it helps a human reader verify your claim. Whether it moves a specific engine's citation behavior is something worth testing on your own content and queries, not assuming from either study.
Clear Entities - What's Actually Documented vs. What's Inferred
Worth separating a documented fact from a reasonable-sounding inference. Google's Article structured-data documentation does recommend an author URL or a sameAs link specifically so its Search systems can better distinguish who wrote something - solid, real guidance for how Google's Search systems parse authorship. What that documentation does not do is report an experiment showing higher citation rates in AI Overviews, ChatGPT, Gemini, Claude, or Perplexity as a result. Entity clarity as a general content principle is reasonable advice on its own terms. But treat it as good practice for how machines parse your content, not as a demonstrated AI-citation lever - the source proves a narrower claim than that.
Freshness - a Real but Narrower Finding Than "Update More"
The freshness evidence is real but narrower than a simple "keep updating" rule. In a controlled test comparing two synthetic sources, a passage carrying a recent (2026) timestamp was preferred over one with an old (2019) or no timestamp - a real, if narrow, finding from the same peer-reviewed research already cited on this page. That test changed the displayed date on injected content; it did not test whether a substantive real-world content update changes citation behavior, which is a different and unverified claim. Separately, Authoritas found that Google AI Overview source lists carry high weighted volatility - the vendor's own published figures are 0.68 over roughly the first eight weeks and 0.73 over the following thirteen, in a US desktop sample of AI-Overview-triggering queries, which it summarizes publicly as "about 70% of pages changing." That shows AI Overview citations move a great deal - it does not show that refreshing any specific page caused or prevented that movement, since the study is observational, not a controlled test of updating. The defensible guidance: update content when the underlying facts actually change, and don't change a displayed date just to appear fresh - Google's own guidance explicitly warns against exactly that.
Does E-E-A-T Actually Matter for AI Citation?
Indirectly, and Google has said this plainly in its own words: "While E-E-A-T itself isn't a specific ranking factor, using a mix of factors that can identify content with good E-E-A-T is useful." Google's automated systems use a mix of factors that can help identify content demonstrating experience, expertise, authoritativeness, and trust - with trust explicitly called the most important of the four, and content not required to demonstrate all of them. Google's Search Quality Rater Guidelines describe the same framework in more depth, as a tool for human raters evaluating page quality - not a formula those raters directly control in live rankings. Use E-E-A-T as a mental checklist for what credible content looks like. Don't treat it, or any individual component of it, as a scored AI-citation metric - no platform has disclosed one.
Original Research and Data
This deserves a more careful claim than "the strongest evidence shows statistics work" - neither controlled study actually isolates "original research" or "first-party data" as its own tested variable. The 2024 benchmark tested an LLM-generated rewrite that added statistics; it didn't verify the numbers were original to the publisher. The SIGIR study tested explicit price, specification, and evidence content in synthetic product-review passages. And as covered above, the peer-reviewed NeurIPS benchmark found comparable statistics rewrites often hurt performance rather than helping. A single vendor case study (Go Fish Digital) reports gains in AI-referral traffic and conversions after publishing several fact-dense pages - directionally encouraging, but it changed prompt mapping, page count, structure, and sourcing simultaneously with no way to isolate which change did what, and it measured traffic and conversions, not citation probability specifically. The honest framing: original research and real data are worth publishing because they give readers and other sites something distinctive and verifiable to reference - a genuinely useful publishing strategy worth testing on your own content - not a proven, isolated citation factor with controlled evidence behind it.
Do AI Engines Prefer First-Party or Third-Party Sources?
This is the section most worth getting precise about, because three studies that sound like they're measuring the same thing are actually measuring three different things. AirOps reports that 85% of attributed brand mentions in its commercial-intent sample came from external domains, 13.2% from the brand's own domain, with the rest uncited. A separate, much larger cross-market preprint (167,551 URL citations across 128 brands, 12 markets, 13 languages) reports a similar 85.7% external share - but it's counting URL citations under its own ownership classification, not attributed brand mentions, and its exact ownership methodology (specifically, how it decides a domain "belongs" to a brand) isn't fully detailed in the materials the authors released publicly. The two figures point the same direction and are worth citing together, but they're not a numerical replication of each other - different units, different prompt sets, different ownership definitions. Worth disclosing plainly: this preprint's author is affiliated with a commercial AI brand-intelligence company that supplied the underlying data, a real conflict of interest even though the study itself is large and substantial.
Yext's local/entity-query research reports 86% of citations from "brand-managed" sources - and this is where a clean "third-party dominates reputation queries, brand-owned dominates local queries" story breaks down. Yext's own breakdown is 44% first-party websites, 42% third-party listings platforms, and 8% reviews and social - all three counted as "brand-managed" because the brand can influence or claim them, not because the brand owns the domain. Read plainly, that means even in Yext's own local-query dataset, a genuine majority of citations (50% or more) come from platforms the brand doesn't own outright - websites like review sites, directories, and listings platforms - just ones the brand can actively manage its presence on. That's a meaningfully different, more useful takeaway than "brand-owned wins locally, third-party wins for reputation." The practical conclusion isn't a binary at all: it's that the mix of source types depends on both the query type and what "brand-controlled" actually means for that query - worth auditing directly for your own brand and query set rather than assuming either pattern applies. It's also worth knowing, as a general caution, that AirOps, Yext, and the cross-market preprint's supplying company all sell products in exactly this measurement market - a reason to weigh the underlying methodology carefully rather than the headline number alone, for any of them.
What to Actually Do
Write specific, sourced claims because they help real readers verify what you're saying - not because one study proves it drives citations, since another comparably rigorous study found the opposite. Make entities and authorship unambiguous, understanding that this helps Google's Search systems parse your content correctly, which is real and separate from an unproven citation-lift claim. Update content when the facts genuinely change, not on a manufactured schedule. Treat E-E-A-T as a mental model for credible writing, not a metric being scored. Publish original data when you actually have it, as a strategy worth testing rather than a guaranteed lever. And before leaning on any single vendor's "X% of citations come from Y" statistic, check what it's actually counting - attributed mentions, URL citations, and "brand-managed" sources are three different measurements that don't translate cleanly into each other.
← Back to the AEO fundamentals pillar
Does domain authority actually matter for AI citations? →
How to structure content so AI engines can actually quote it →
Sources: The finding that statistics/quotation/citation rewrites increased cited-text share in a fixed-source testbed is from Aggarwal et al., "GEO: Generative Engine Optimization" (ACM KDD 2024) - verified directly, including the specific GPT-3.5, five-source experimental setup. The directly opposing finding - that comparable rewrites across six domains and four newer models mostly failed to improve citation rank, with statistics rewrites specifically hurting it in 19 of 24 settings - is from Puerto, Gubri, Green, Oh, and Yun, "C-SEO Bench: Does Conversational SEO Work?" (NeurIPS 2025, Datasets & Benchmarks Track - confirmed via the official NeurIPS proceedings listing, not just an arXiv posting). The timestamp-preference and completeness findings are from Vishwakarma, Kumar, and Jamidar, "What Gets Cited: Competitive GEO in AI Answer Engines" (confirmed as genuine ACM SIGIR 2026 proceedings, 49th International ACM SIGIR Conference) - cited here with its two-document synthetic testbed scope stated directly rather than generalized. Entity-disambiguation guidance is from Google's Article structured-data documentation. Citation-list volatility data is from Authoritas' volatility research - verified directly, scoped here to Google AI Overviews and its US desktop sample specifically. The E-E-A-T quote is from Google's "Creating helpful, reliable, people-first content" documentation - verified verbatim - supplemented by Google's Search Quality Rater Guidelines (dated September 11, 2025) for the fuller rater-facing framework. The vendor case study is Go Fish Digital's GEO case study, cited with its limitations stated directly. The attributed-brand-mention figure (85% external) is from AirOps' "Third-Party Sources Drive 85% of Brand Discovery" report, already cited on the pillar. The URL-citation figure (85.7% external) is from Żatuchin's cross-market preprint and its released dataset (167,551 citations, 128 brands) - verified directly at both, the author's commercial affiliation with the data-supplying company disclosed above, and the exact ownership-classification methodology noted as not being fully detailed in the public release. Yext's local/entity-query breakdown (44% first-party websites, 42% listings, 8% reviews/social) is from Yext's own research release, verified directly and corrected here from an earlier draft's oversimplified "86% brand-owned" characterization.
About the author
Zarko Zivkovic is the founder of CoreAEX, building technical SEO, AEO, and AI-visibility systems for B2B SaaS companies. Connect on LinkedIn.