A wave of advice about AI search citation has been circulating, much of it repeated often enough to acquire the social weight of consensus. Some of it has evidence behind it. Some of it is a statement dressed as proof. Philipp Götza’s analysis of three recommendations — llms.txt, schema markup and content freshness — offers a method for telling the difference.
Table of contents
- What this is about
- What the source says
- What the data shows
- What it does not mean
- What to do
- The three claims at a glance
- FAQ
- Key takeaways
Key takeaways
- llms.txt sits at “statement” on the evidence ladder — crawlers fetch the file, but no data shows it changes citation rate.
- Schema markup correlates with AI visibility in some studies, but rival explanations exist and recent experiments found no direct mechanism.
- Content freshness has the strongest evidence base of the three, supported by research from Ahrefs, Generative Pulse, Seer Interactive, and a peer-reviewed paper — though that paper has methodological limits.
- A meta-analysis circulating in newsletters misrepresented the 2023 GEO study, misdated it as 2024, and described a weighted score as a correlation coefficient.
- Götza’s ladder — statement, fact, data, evidence, proof — is a usable filter for any AI search claim you encounter.
What this is about
Three recommendations dominate current advice about AI citation: create an llms.txt file, add schema markup, and publish fresh content. The analysis measures each against a five-rung ladder from statement to proof. Two fail well before the top rung. One has genuine support, with caveats. Site owners need the method as much as the verdicts.
What the source says
Philipp The piece is built around the ladder of misinference, drawn from Alex Edmans’ book May Contain Lies, and it is worth reading alongside the analysis itself. A claim climbs from statement to fact to data to evidence to proof. Most AI search advice, he argues, stalls somewhere in the middle and gets shared as if it reached the top.
He opens by identifying three reasons practitioners accept weak advice: genuine ignorance, cognitive bias (confirmation bias in particular), and what he terms amathia: the voluntary refusal to know better. Black-and-white thinking compounds all three: “backlinks are always good”, “Reddit is always important for AI search”, “blocking AI bots is always stupid.” None of those absolutes survives contact with a specific site and a specific prompt set.
Before reaching the three recommendations, it turns to a circulating meta-analysis, left unnamed. His specific objections: it misdates the GEO study as 2024 rather than 2023, describes a weighted score as a correlation coefficient, and claims the GEO study “confirms” that schema markup, lists, and FAQ blocks significantly improve AI inclusion; a conclusion the study does not make. Reported sample sizes are inaccurate, and one source is described as multiple layers of hearsay. Edmans is quoted directly on this pattern: that loudly proclaiming findings as groundbreaking is itself a signal they may not be.
On llms.txt: the file is a 2024 proposal that spread through influencer amplification rather than platform adoption. Google crawls and indexes these files, and that is a fact. Evidence that they change citation rate does not exist. A small cited experiment found no impact from either llms.txt or .md files on AI citations. The original proposal also recommended appending .md to page URLs to serve markdown versions, which would create internal duplication and inflate crawl volume. The only case where the file serves a clear purpose is a complex API that AI agents can navigate.
On schema markup: the mechanism breaks down at two points. During training, HTML is stripped before text reaches the model, and tokenisation during pretraining would destroy coherent markup anyway. During grounding, the retrieval step where a chatbot fetches live pages to support an answer, there is no evidence in the analysis that AI chatbots read schema. Correlation studies show sites with schema have better AI visibility, but he names several rival explanations. Recent experiments, including one he tested in Perplexity Comet, showed the tool hallucinated schema that was not on the page. Schema is still recommended, for a different reason. It supports rich results in traditional search, and those rankings are a signal answer engines use during fan-out queries.
On content freshness: foundation models have a training cutoff around end of 2022. Platforms including OpenAI, Anthropic, and Perplexity use freshness as a signal when deciding whether to trigger web search. Research from Ahrefs, Generative Pulse, and Seer Interactive supports the hypothesis that recently updated content gets cited more often. A scientific paper adds further support, though Götza flags three limits: it used API results rather than the user interface, it asked the model to rerank rather than observing how reranking actually works, and date injection was artificial enough to exaggerate the effect. He still rates freshness as the most defensible of the three recommendations.
What the data shows
The freshness research Götza cites, Ahrefs, Generative Pulse, Seer Interactive, and the peer-reviewed paper, is not fully described in the source. None of the studies is quoted with a prompt count, a repetition number, a stochastic variation note, or a date window. The analysis itself flags the API-versus-interface gap and the artificial date injection in the scientific paper.
Following the standard this site applies to third-party research: the directional finding (fresher content correlates with higher citation rates) is worth acting on. The specific percentages from those studies are not quoted here because the method points needed to treat them as precise figures are absent from the source.
The llms.txt experiment mentioned there is described only as “a small experiment” with no sample size, prompt set, or repetition count. It is consistent with the absence of any positive evidence, but it is not itself proof of no effect.
The meta-analysis critique is the most methodologically specific part of the source. It identifies a concrete error, a weighted score labelled as a correlation coefficient, and a concrete misrepresentation of the 2023 GEO study. Those are checkable claims, not impressions.
What it does not mean
This is where most of the AI search advice ecosystem goes wrong, so it’s worth being direct.
“Schema markup correlates with AI citations” does not mean schema markup causes AI citations. The analysis names this explicitly. Sites that implement schema well tend to be sites that also have strong technical foundations, clear entity signals, and higher search rankings. Any of those factors could explain the correlation. The mechanism, schema markup being read by a retrieval tool during grounding, has not been demonstrated. Götza tested it directly in Perplexity Comet and found the tool hallucinated markup that wasn’t there. Correlation studies with no mechanism are not evidence of a lever you can pull. The same caution applies to vendor conversion data.
“Freshness matters” does not mean updating a date stamp is enough. The point is explicit: Google stores up to 20 historical versions of a page and can detect date manipulation without substantive content changes. The recommendation is to update content, not metadata.
“llms.txt is indexed by Google” does not mean it influences AI citations. Indexing and citation influence are separate things. The file being crawlable is a fact, and the Web Almanac data shows how many sites now carry one. The file being useful for citation is a claim without supporting data.
A meta-analysis that synthesises studies is not automatically more reliable than a single study. The critique of the unnamed meta-analysis makes this case precisely. Combining misread studies, misdated sources, and mislabelled statistics produces a more impressive-looking document, not a more accurate one. The volume of sources is not a substitute for reading them correctly.
Authority is not accuracy. The analysis closes with this point, and it applies directly to how AI search advice spreads. A recommendation repeated in enough newsletters acquires the social weight of consensus. That weight is not evidence. The ladder runs from statement to proof regardless of who is making the claim or how many times it has been forwarded.
One more thing worth naming: the advice to “use AI to summarise research before acting on it” has a specific problem here. Götza notes that brief-summary prompts increase hallucination rates, and that source material can lend false credibility to a response. Using an answer engine to evaluate answer-engine research is a loop with no exit.
What to do
Ordered by effort, lowest first.
1. Apply the ladder before you act. When you encounter an AI search recommendation, ask where it sits: statement, fact, data, evidence, or proof. Most current advice peaks at “fact”, something measurable exists, without climbing to evidence of a causal mechanism. That is enough to inform a hypothesis, not enough to justify a content or technical overhaul.
2. Keep content current where freshness is query-relevant. Update the substance of pages that cover topics where recency matters, pricing, product comparisons, regulatory status, anything date-sensitive. Update on-page dates, schema markup dates, and sitemap lastmod consistently, and only when the content has actually changed.
3. Maintain schema markup for rich result eligibility. Use supported types with all relevant properties. The direct mechanism for AI citation grounding is unproven, but schema supports the search rankings that answer engines use as a retrieval signal. It is worth doing for that reason, not because it speaks directly to a retrieval tool.
4. Monitor llms.txt adoption in crawler logs, but don’t create the file yet. Check log files to see how crawl volume to llms.txt changes over time, and review quarterly whether OpenAI, Anthropic, or Google have formally announced support. If OpenAI, Anthropic, or Google formally announce support with published specification, create the file to that specification. Until then, the cost is low and the benefit is undemonstrated.
The three claims at a glance
Sources: Götza’s analysis; method limitations noted above apply to all freshness figures.
| Recommendation | Evidence level (Götza’s ladder) | Known mechanism | Götza’s verdict |
|---|---|---|---|
| llms.txt | Statement | None demonstrated | Skip for now; revisit quarterly |
| Schema markup | Correlation data; no causal mechanism | Not demonstrated for grounding | Keep for rich results and SEO signals |
| Fresh content | Data + supporting research (method caveats) | Freshness signal used by OpenAI, Anthropic, Perplexity to trigger web search | Act on it |
| Meta-analysis (unnamed) | Fails at data level | Weighted score mislabelled; GEO study misdated and misrepresented | Disregard |
The through-line is uncomfortable but useful: two of the three recommendations everyone repeats have no demonstrated mechanism behind them, and the third works for reasons that have nothing to do with the tactic itself. If you do one thing after reading this, ask of the next AI-search claim you meet where it sits on the ladder — and refuse to act until it reaches evidence.
FAQ
Does blocking AI crawlers hurt citation rates?
The analysis treats this as black-and-white thinking. Blocking all AI crawlers is not always wrong, and allowing all of them is not always right. For some business models, paywalled content, proprietary data, competitive research, blocking specific bots is a defensible choice. The relevant question is which crawler, at what access level, for which content. Different crawlers do different jobs, and a single blanket rule ignores that.
If the 2023 GEO study doesn’t confirm that schema improves AI inclusion, what does it actually say?
Götza’s specific objection is that the unnamed meta-analysis claimed the GEO study “confirms” that schema markup, lists, and FAQ blocks significantly improve inclusion in AI responses, and that a review of the study shows it makes no such claim. The source does not quote the GEO study’s actual findings directly, so the full scope of what it does conclude is not available here. What is clear is that a claim of confirmation requires the cited study to have tested that specific thing and found that specific result. Götza’s point is that it did not.
How should I evaluate freshness research given the methodological limits Götza identifies?
Three limits apply to the scientific paper he cites: API results differ from interface results because of system prompts and API settings; asking a model to rerank is not how reranking actually operates in production; and the artificial date injection may have exaggerated the effect size. For the industry research from Ahrefs, Generative Pulse, and Seer Interactive, the source does not provide prompt counts, repetition numbers, or stochastic variation data, so precise percentages from those studies should be treated as directional rather than exact. The convergence of multiple independent sources pointing the same direction is meaningful. The specific magnitudes are not reliable enough to use as targets.
What is the ladder of misinference, in one paragraph?
It is a five-rung scale borrowed from Alex Edmans: statement, fact, data, evidence, proof. A statement is an assertion. A fact is something measurable. Data is a measurement taken. Evidence links the measurement to the claim. Proof establishes the mechanism. Most AI search advice stalls between fact and data, then gets shared as though it reached the top rung.
Is there any case where creating an llms.txt file makes sense?
One: a complex API that AI agents are expected to navigate. Outside that case the analysis finds no demonstrated effect on citation, and the original proposal’s suggestion to serve markdown copies at .md URLs would create internal duplication and add crawl volume for no established gain.



