// finding 1
Cited pages cluster at 2,000 to 3,500 words
Word counts were measured on rendered HTML, not source files. Source measurement roughly doubles the figure, and we corrected that error mid-study rather than publishing it.
// what cited pages share
Six structural findings
Lists are the strongest signal
A median of 110 list items per cited page, far above what a typical service page carries. A list item is a self-contained retrievable chunk; a paragraph inside a div is not.
Length peaks in the middle
25 of 58 pages sat between 2,000 and 3,500 words. Only 3 exceeded 6,000. The advice to publish very long pages for AI visibility is not supported by what is cited.
FAQ schema is not required
FAQPage markup appeared on 41% of cited pages while a visible FAQ section appeared on 72%. The visible section is nearly twice as common as the markup.
Listicles are a minority
Only 18% carried a numbered listicle headline, despite the format being widely recommended for AI visibility. Use it where a topic enumerates, not by default.
One H1, always
55 of 58 had exactly one H1 and none had zero. This is basic hygiene rather than an insight, but it is the most consistent trait in the set.
Freshness is common, not universal
dateModified was present on 56%. Worth having and cheap to add, but a clear majority rather than a requirement.
// finding 2
Domain authority did not predict citation
We pulled referring-domain counts for the agencies most frequently named by ChatGPT and Perplexity across nine buyer prompts. If authority drove citation, the ordering should track. It did not.
| Domain | Referring domains | LLM mentions | Read |
|---|---|---|---|
| netalico.com | 557 | 17 | Small profile, heavily cited |
| swankyagency.com | ~1,000 | 18 | Mid profile, heavily cited |
| elogic.co | ~1,300 | 19 | Mid profile, heavily cited |
| outerboxdesign.com | 4,718 | barely present | Largest profile, near-absent |
This does not mean links stopped mattering. Authority still governs organic ranking, and our own site sits at 53 referring domains, which is precisely why we rank poorly on competitive terms. The narrower claim is the interesting one: in this sample, authority and citation came apart.
// finding 3
Five widely repeated claims the data contradicts
| Common claim | What we measured |
|---|---|
| "Publish 5,000+ word pages to win AI visibility." | Median cited page: 2,813 words. Only 3 of 58 exceeded 6,000. |
| "FAQPage schema is a Tier 1 AI citation factor." | Present on 41% of cited pages. Google states no special schema is needed. |
| "llms.txt is a technical necessity." | Google says no AI system currently uses it. Ahrefs found 97% of llms.txt files got zero requests across 137,000 domains in May 2026. |
| "Every page needs a listicle." | Listicle headlines appeared on 18% of cited pages. |
| "Pages under 125 words convert 15% higher" (attributed to Unbounce). | Not in that report. We could not locate this figure in any primary source. Treat as fabricated. |
// method
How we ran it
01
Capture
12 commercial e-commerce queries run through Google with AI Overview loading enabled, yielding 232 citation URLs.
02
Filter
Publisher, marketplace and platform-vendor domains excluded, so the sample reflects pages a services company can realistically compete with.
03
Fetch
The 60 most-cited remaining URLs requested with a real browser user-agent. 58 returned successfully.
04
Measure
Rendered HTML only: script, style, noscript and SVG stripped before counting words, headings, list items, tables and schema types.
05
Publish
Scripts and raw JSON committed to the repository so any figure here can be re-derived or challenged.
// limitations
What this study cannot tell you
Stated in full, because a study that buries its limitations is marketing rather than research.
- Sample is 58 pages in a single vertical: e-commerce and agency commercial queries.
- Google AI Overviews only. ChatGPT and Claude citation URLs were not captured on the analysis run, so their page structures are unmeasured.
- Capped at the 60 most-cited URLs, so long-tail cited pages are excluded.
- Publisher, marketplace and vendor domains were deliberately excluded, which shifts the medians.
- Correlation, not causation. These pages are cited and share these traits. Nothing here proves the traits caused the citation.
// what Google says
The primary sources agree with the data
Google states directly that there are no additional requirements to appear in AI Overviews or AI Mode and no special schema.org structured data to add. OpenAI documents that OAI-SearchBot governs ChatGPT search inclusion, separately from the GPTBot training crawler. Anthropic documents Claude-SearchBot and Claude-User as its retrieval crawlers, with ClaudeBot used for training. Our structural findings are consistent with all three: this is ordinary excellent SEO plus extractability, not a separate discipline.
STUDY FAQ
Questions about the method and the findings
Topics
Can’t find your answer?
Talk to the founderThe findings
What do pages cited by Google AI Overviews have in common?
Across 58 cited pages the median carried 2,813 words, 11 H2 sections, and 110 list items. 72% had a visible FAQ section and 95% had exactly one H1. None had zero H1. The single most distinctive trait was list density: a median of 110 list items per page.
How long should a page be to get cited by AI?
The median cited page was 2,813 words. Only 3 of 58 exceeded 6,000 words, while 25 sat between 2,000 and 3,500. Longer is not better. The common advice to publish 5,000-plus-word pages for AI visibility is not supported by what is actually being cited.
Is FAQPage schema required for AI citation?
No. Only 41% of the cited pages had FAQPage schema, while 72% had a visible FAQ section. The visible section is far more common than the markup. Google states directly that no special schema is needed for AI features. Ship FAQ schema because it is cheap and honest, not because it buys citations.
Does domain authority predict AI citation?
Not in our data. Measuring referring domains against LLM mention counts across nine buyer prompts, netalico.com with 557 referring domains was cited 17 times, while outerboxdesign.com with 4,718 referring domains barely appeared on the same prompts. Authority buys organic position. It did not buy citation here.
Do listicles get cited more by AI?
Not for commercial queries. Only 18% of the cited pages had a numbered listicle headline. The listicle format is widely recommended for AI visibility and it was a clear minority pattern in this sample. Use it where the topic genuinely enumerates, not as a default.
Do comparison tables help with AI citation?
32% of the cited pages contained a table. That makes tables common but not required. They are worth adding where the topic genuinely compares things, since a table is an unusually clean extractable unit, but the data does not support treating them as mandatory.
Method
How was this study conducted?
We ran 12 commercial e-commerce queries through Google with AI Overview loading enabled, captured 232 AI Overview citation URLs, excluded publisher and marketplace domains so the sample reflected pages a services company could realistically compete with, then fetched the 58 most-cited remaining pages and measured their rendered HTML.
What exactly was measured on each page?
Rendered word count after stripping script, style, noscript and SVG elements; heading counts by level; list items; tables; ordered and unordered lists; JSON-LD schema types; presence of a visible FAQ; listicle headline patterns; and whether dateModified was present. Everything was measured on live rendered HTML, never on source files.
Why measure rendered HTML rather than the source?
Because source files include markup, styles, and structured data that inflate word counts by roughly double. We made that exact mistake mid-study and corrected it: a page measuring 6,883 "words" in source was 3,904 rendered. Any study comparing content length has to state which it measured.
Limitations
What are the limitations of this study?
Four worth stating. The sample is 58 pages in one vertical, e-commerce and agency commercial queries. It reflects Google AI Overviews only, not ChatGPT or Claude. It is capped at the 60 most-cited URLs, so long-tail cited pages are excluded. And it is correlation: these pages are cited and share these traits, which does not prove the traits caused the citation.
Why is ChatGPT and Claude data missing from the structural analysis?
Our capture of citation URLs from ChatGPT and Perplexity answers returned no URLs on the run used for the structural analysis, so those pages could not be fetched and measured. The structural findings are Google AI Overview only. We have separate LLM mention-rate data, but not the page structures behind it.
Could the sample be biased toward certain page types?
Yes, deliberately and disclosed. We excluded publisher, marketplace, and social domains such as Forbes, Reddit, Clutch and the platform vendors themselves, because the question we were answering was what a services company can compete with. A study including those would produce different medians.
Applying it
What should we actually change based on this?
Three things with the clearest support. Raise list density, since the median cited page carries 110 list items and most service pages carry far fewer. Aim for 2,500 to 3,500 words rather than padding higher. Add a visible FAQ section, which 72% of cited pages had. Everything else in the data is weaker.
Does this mean backlinks do not matter?
No. Authority still governs organic ranking position, and our own site sits at 53 referring domains against a competitor median of 1,799, which is exactly why we rank poorly on competitive terms. The finding is narrower: authority did not predict AI citation in this sample. The two are different games with different inputs.
What does Google itself say about optimising for AI Overviews?
Google states there are no additional requirements to appear in AI Overviews or AI Mode and no special schema.org structured data to add. Its guidance is ordinary SEO fundamentals: allow crawling, keep important content in text, use internal links, and ensure structured data matches visible content. Our data is consistent with that.
Is llms.txt worth implementing?
Not as a visibility lever. Google has stated no AI system currently uses llms.txt, and an Ahrefs analysis of 137,000 domains found 97% of llms.txt files received zero requests in May 2026. It is cheap and harmless to publish, but it is not a ranking or citation factor.
Use the data, challenge the data
The scripts and raw JSON behind every figure are committed in our repository. If you reproduce this and get a different answer, we would genuinely like to know. If you want help applying it to your own commerce site, we build B2B commerce, replatforming and headless storefronts.