Seer Interactive tracked 1,562 branded prompts for five weeks and reports that the six engines refused to answer 31.2% of the time, while answering wrongly only 2.8%. We checked whether SEO brands are shutting the crawlers out themselves.
Table of contents
- What this is about
- What the source says
- What the data shows
- What it does not mean
- What to do
- The numbers
- FAQ
Key takeaways
- Seer Interactive reports that across six engines, branded prompts were declined 31.2% of the time and answered incorrectly 2.8% of the time. Silence is the dominant failure mode, not error.
- Comparative prompts of the “X versus Y” kind were declined 80.3% of the time, against roughly 10% for direct questions about the brand.
- The study discloses five of the seven method points we require. It does not say whether responses were collected through the interface or the API, and it does not state a region, so we report these as the company’s figures rather than as established rates.
- We checked robots.txt on 21 SEO and search publications: 18 leave every AI crawler open, one blocks GPTBot alone, one blocks seven crawlers. Industry silence is not caused by blocking.
- A refusal cannot be corrected the way an error can. There is nothing to dispute, and no page to fix.
What this is about
Seer Interactive published a study of what six AI engines say about Seer itself, after finding that models were repeating a claim about staff turnover sourced from a single third-party page. The useful finding is not the error rate. It is how often the engines said nothing at all.
What the source says
The study came out of a specific incident. Seer asked models what people say about working with Seer Interactive, and most answered with a claim about high account manager turnover — traced back to one third-party page. That prompted a wider check.
The method snapshot the company publishes is unusually detailed for this kind of research: six engines named (AI Mode, AI Overviews, ChatGPT, Claude, Gemini and Perplexity), 1,562 prompts tracked, collected daily from 1 May to 9 June 2026, producing 28,123 responses, scored against 87 brand attributes in six categories covering company facts, leadership, services, and positioning.
Responses were sorted into three buckets rather than two: accurate, inaccurate, and non-answer. That third bucket is the reason the study is worth reading. Most brand-monitoring work counts what the model got wrong; this one counts what the model declined to say.
The headline pair, in the company’s own framing: when the engines answered, they were right 95.9% of the time, and the average inaccurate response rate was 2.8%. But they declined to answer 31.2% of the time.
The split by prompt type is sharper. Direct questions about the brand were answered accurately 90.4% of the time, indirect ones 88.7%. Comparative prompts, the “which is better, X or Y” questions a buyer actually asks, collapsed to 18.8% accuracy, with 80.3% of responses classed as non-answers.
Per-engine, the company reports AI Mode with the highest overall accuracy at 74.2% and Claude the lowest at 52.6%. Filter Claude down to the prompts it actually answered and the figure becomes 98%. The engine that refuses most is also the engine that is most reliable when it speaks.
What the data shows
Before the numbers, the method check we apply to any vendor research. Of the seven points we require (model, interface or API, collection date, region, prompt set, number of repeats, and stochastic variability) this study discloses five.
It names the engines. It gives the window, 1 May to 9 June 2026. It states the prompt count, 1,562, and the response count, 28,123. Daily collection over five weeks covers repetition, which is more than most studies of this kind offer.
Two points are missing. The study does not say whether responses were collected through each product’s interface or through an API, which matters because the two can return different things for the same prompt. And it states no region, which matters more than it sounds: AI Overviews and AI Mode vary by market, and a US-only sample says little about the same brand elsewhere. Day-to-day variance is not quantified either, though daily collection means the underlying data exists.
So the figures above are what Seer Interactive measured about Seer Interactive, reported by the company. They are a strong signal and a reasonable starting hypothesis. They are not a rate you can apply to your own brand.
Our own check. If engines decline a third of branded prompts, one obvious explanation is that they cannot read the brand’s own site. We tested that on the industry itself: we fetched robots.txt from the 21 SEO and search publications in our source list on 10 September 2026 and parsed the rules for ten AI crawler tokens.
Twenty of the 21 serve a robots.txt. Eighteen of those leave every AI crawler open — no disallow for GPTBot, ClaudeBot, OAI-SearchBot, PerplexityBot, Google-Extended, Applebot-Extended, CCBot or Bytespider. One, Moz, blocks GPTBot and nothing else, which is the precise “no training, still cite me” position the tokens actually allow. One, SE Ranking, blocks seven of the ten.
The industry is not shutting the door. Whatever makes an engine decline to answer, it is not that the brand’s own pages are unreachable.
What it does not mean
A 31% non-answer rate is not a 31% problem for your brand. It is one company’s measurement of one company’s name, run against a prompt set that company designed around its own 87 attributes. A brand with a shorter history, a more common name, or fewer third-party mentions would produce different numbers in both directions.
A non-answer is not evidence that the engine is being careful. It is tempting to read the Claude figures, refusing most and 98% accurate when it speaks, as a model exercising judgment about what it knows. The data supports no such claim. Refusal rates reflect product policy, retrieval coverage and prompt phrasing, and none of the six vendors documents how that decision is made. Treat the pattern as observed behaviour, not as a disclosed mechanism.
The comparative-prompt collapse is not a ranking problem. An 80.3% non-answer rate on “X versus Y” questions looks like an absence of visibility, but the more likely cause is that engines avoid adjudicating between named competitors at all. That is a policy boundary, not a gap you can fill with better content about yourself. Publishing your own comparison page does not make a model willing to pick a winner.
And blocking is not the lever here. Our robots.txt check makes that concrete: 18 of 20 sites are fully open to AI crawlers and the engines still decline. Access was never the binding constraint.
What to do
First, find out whether you are being declined or described. Ask each engine the three question shapes separately: a direct question about your company, an indirect one about the problem you solve, and a comparison against a named competitor. Record which produce an answer at all. That distinction is the whole finding, and no tool is needed to see it.
Second, treat a non-answer as a content gap, not a reputation problem. An inaccurate claim has a source you can find and challenge. A refusal has nothing to point at. Where an engine says nothing about, say, where you operate or who runs you, the practical fix is to state that plainly on a page of your own, in text, in one place. That is also what makes a passage liftable in the first place.
Third, check your third-party surface before your own site. The Seer incident started with one page nobody at the company controlled. Search for your brand plus the words a buyer would use (turnover, pricing, complaints) and see what a model would find. Fixing your own page does nothing about the page that is actually being cited.
Fourth, do not build a programme on someone else’s percentage. Ours included. Run the three prompt shapes monthly, log what came back, and use your own numbers — citation counts on their own move less than teams expect, and the same caution applies to refusal counts.
The numbers
| Measure | Reported figure | Source |
|---|---|---|
| Branded prompts declined | 31.2% | Seer Interactive, six engines |
| Inaccurate responses | 2.8% | same |
| Accuracy when the engine answered | 95.9% | same |
| Direct prompts, accuracy | 90.4% | same |
| Indirect prompts, accuracy | 88.7% | same |
| Comparative prompts, accuracy | 18.8% | same |
| Comparative prompts, non-answers | 80.3% | same |
| Highest overall accuracy, AI Mode | 74.2% | same |
| Lowest overall accuracy, Claude | 52.6% | same |
| Claude, answered prompts only | 98% | same |
| Prompts tracked / responses scored | 1,562 / 28,123 | same |
Seer Interactive, 1 May to 9 June 2026, collected daily, 87 brand attributes. The study does not state whether responses came from the interface or an API, and states no region; we report these as the company’s figures rather than as general rates.
| Our check: AI crawlers in robots.txt | Sites |
|---|---|
| Serve a robots.txt | 20 of 21 |
| Leave every AI crawler open | 18 |
| Block GPTBot only | 1 (Moz) |
| Block seven of ten tokens | 1 (SE Ranking) |
The 21 SEO and search publications in our source list, fetched 10 September 2026, robots.txt parsed at the domain root for ten crawler tokens. Directives were read per user-agent block, falling back to the wildcard block.
FAQ
Is a non-answer worse than a wrong answer?
For a brand, usually yes, because there is nothing to act on. An inaccurate claim has a source: you can find the page, check it, and ask for a correction, as Seer did after tracing a turnover claim to a single third-party page. A refusal gives you no page, no claim and no counterparty. It also does not show up in any monitoring that counts only mentions and sentiment, so it can persist unnoticed.
Should I block AI crawlers if the engines get things wrong about me?
Blocking removes your own words from the pool and leaves whatever third parties wrote. In our check of 21 SEO publications, 18 leave every AI crawler open and only one blocks more than a single token — the industry that studies this most closely has largely decided against it. If your concern is training rather than citation, the tokens separate those two things, and blocking the training crawler alone is the narrower move.
Can I use the 31.2% figure as a benchmark?
No. It is one company measuring its own brand with a prompt set built from its own 87 attributes, and the study does not state whether it used the interface or an API, or which region it ran in. Both change what an engine returns. Run the same three prompt shapes for your own brand and compare against yourself over time; that number is worth something, a borrowed one is not.



