7 October 2026

The web is readable to AI. It mostly doesn't say who wrote it.

We scanned 500 hosts drawn at random from the top 10,000 most-visited domains. Models can fetch and parse these pages without much trouble. What they cannot do, on four sites out of five, is work out who published the page or when.

That matters, because a model deciding whether to cite a source needs to know what the source is.

This study replaces an earlier one published on 30 September. That version contained a bug in our own scorer which invented its main finding. The correction is at the bottom of this page.

Method, and what it can and cannot say

Sample. 500 hosts drawn at random from the Tranco daily top 10,000, list of 5 October 2026. Tranco aggregates several popularity rankings and resists manipulation better than any single list.

Three categories, not two. Tranco ranks domains, not websites. 97 of our 500 are infrastructure: CDN hostnames, ad servers, API endpoints, names that don't resolve to a public address. They are not pages anyone reads, and counting them as failures would inflate our numbers.

Scored286
Websites we could not read117
Not websites97
Total drawn500

Our denominator is 403 websites. Every percentage below is against that, or against the 286 we could score โ€” each one says which.

One request per host, no JavaScript. This is a real limitation for single-page applications, and we would expect it to understate schema detection somewhat.

Evidence stored. Unlike our September run, we kept the served HTML, the robots.txt body, and every HTTP status. Any check here can be replayed without crawling again. One host (yoomoney.ru) contained a null byte that Postgres rejects; it was stripped and the row flagged.

Raw counts are at the bottom. Tell us if we got something wrong.

1. Readable, but anonymous

Split the score into its five dimensions and the gap is not where we expected:

Brand clarity

Mean
68.0
Median
67

AI accessibility

Mean
66.9
Median
75

Content structure

Mean
61.3
Median
75

Authority signals

Mean
42.9
Median
33

Schema coverage

Mean
41.3
Median
33

Access is not the problem. Crawlers are allowed, sitemaps resolve, pages parse. The two weak dimensions are both about provenance.

The individual checks make it concrete:

  • 81.8% give no author signal
  • 66.8% carry no publication date
  • 64.3% declare no Organization schema
  • 67.5% declare no WebSite schema

A page with no author, no date and no publisher entity is readable but unattributable. A model can summarise it. It has much less reason to name it.

2. The door is now identity-based

117 of 403 websites โ€” 29% โ€” would not let our crawler read them.

HTTP error57
robots.txt disallowed us27
Connection or other26
Certificate7

And within those HTTP errors, the distribution is lopsided:

40346
4045
4292
4012
4001
5221

Forty-six outright refusals. This is bot protection declining an agent it does not recognise.

Be careful what this means. It does not mean 29% of the web blocks AI. GPTBot announces itself and is widely allowlisted; our scanner is not. What it shows is that access is granted by identity, not by behaviour. A named, known crawler gets in. An unknown one does not โ€” even when it reads one page, slowly, and obeys robots.txt.

That is a structural fact about the 2026 web, and it has a consequence: a model can only cite what its own crawler is permitted to fetch, and permission now depends on being on a list.

3. Being popular buys you nothing

Correlation between score and rank in the top 10,000: โˆ’0.066. There is no relationship.

1โ€“1,000

Scored
16
Median
60

1,001โ€“2,500

Scored
56
Median
59

2,501โ€“5,000

Scored
73
Median
60

5,001โ€“10,000

Scored
141
Median
59

Four bands, four medians within one point of each other. The most-visited sites on the internet score the same as those ranked ten thousandth.

Nobody has a head start here. That window will not stay open.

4. The distribution

Minimum18
25th percentile49
Median59
75th percentile70.75
Maximum95
Mean58.7
10โ€“192
20โ€“298
30โ€“3928
40โ€“4940
50โ€“5969
60โ€“6961
70โ€“7949
80โ€“8928
90โ€“991

One site reached 90. None reached 100.

5. llms.txt: 19%

Of the 403 completed websites, 76 served a /llms.txt โ€” 18.9%. Restricting to the 337 that returned any status at all, it is 22.6%.

20076
404213
40341
no status66
other 4xx/5xx7

llms.txt is a young convention and we are not claiming every site needs one. But it remains the clearest single marker of whether anyone at an organisation has thought about this at all.

6. Which AI crawlers get blocked

Across the 298 websites whose robots.txt we read:

CCBot

Blocked by name
36
Blocked by User-agent: *
10
Allowed
252
Total blocked
15.4%

Bytespider

Blocked by name
33
Blocked by User-agent: *
10
Allowed
255
Total blocked
14.4%

GPTBot

Blocked by name
31
Blocked by User-agent: *
10
Allowed
257
Total blocked
13.8%

ClaudeBot

Blocked by name
27
Blocked by User-agent: *
10
Allowed
261
Total blocked
12.4%

Google-Extended

Blocked by name
21
Blocked by User-agent: *
10
Allowed
267
Total blocked
10.4%

PerplexityBot

Blocked by name
20
Blocked by User-agent: *
9
Allowed
269
Total blocked
9.7%

CCBot โ€” the Common Crawl collector, the oldest of the six and the one most associated with training data โ€” is blocked most. PerplexityBot, the newest and the one most associated with answering live questions, is blocked least.

The ordering tracks how much press each crawler received during the 2023โ€“2024 scraping backlash more closely than it tracks what each one actually does today. These rules were written as a reaction, and most have not been revisited.

7. The eighteen checks

Failure rate across the 286 scored sites:

Software/product schema*

Fail
279
%
97.6%

FAQ content present

Fail
263
%
92.0%

Author signal

Fail
234
%
81.8%

llms.txt accessible

Fail
214
%
74.8%

WebSite schema

Fail
193
%
67.5%

Publication date

Fail
191
%
66.8%

Organization schema

Fail
184
%
64.3%

Unique H1

Fail
153
%
53.5%

Any schema present

Fail
127
%
44.4%

Sitemap accessible

Fail
113
%
39.5%

Clear positioning

Fail
109
%
38.1%

Heading hierarchy

Fail
102
%
35.7%

Lists present

Fail
68
%
23.8%

External links

Fail
64
%
22.4%

GPTBot allowed

Fail
29
%
10.1%

ClaudeBot allowed

Fail
23
%
8.0%

Brand in title

Fail
13
%
4.5%

Paragraph length

Fail
10
%
3.5%

* Software schema is reported but excluded from the score. It asks a homepage to declare itself SoftwareApplication, which is reasonable for a SaaS and meaningless for a newspaper or a shop. Keeping it would have penalised almost every site for not being software.

Note the bottom of the table. Brand in title, paragraph length, heading hierarchy โ€” the classic on-page items two decades of SEO advice drilled in โ€” are where sites do well.

The top of the table is everything a model needs to attribute a claim to a source.

What we take from this

The problem moved. A year ago the question was whether models could reach your pages. On this sample they can: AI accessibility scores 66.9, sitemaps resolve, crawlers are allowed. The question now is whether a model can tell your page apart from anyone else's, and name you for it.

Attribution is cheap to fix and nobody is doing it. An author field, a publication date, an Organization block. These are a morning's work, not a quarter's. 82% of the top 10,000 have not done the first one.

Access is becoming a membership question. 46 outright refusals to an unknown crawler, against 41 named blocks of GPTBot. More sites refuse an agent they don't recognise than deliberately block the one everybody has heard of. If being fetchable depends on being on an allowlist, the set of sources a model can cite narrows to those already known.

Raw data

  • Sample: 500 hosts, random draw from Tranco daily top 10,000, list of 5 October 2026
  • Scanned: 5โ€“7 October 2026
  • Scored 286 ยท unreadable websites 117 ยท not websites 97
  • Denominator for site-level rates: 403 websites
  • Score: min 18, p25 49, median 59, p75 70.75, max 95, mean 58.72
  • Dimensions (mean): brand clarity 68.03, AI accessibility 66.87, content structure 61.28, authority signals 42.94, schema coverage 41.30
  • /llms.txt: 76 of 403 returned 200 (18.9%); 76 of 337 measured (22.6%)
  • robots.txt parsed on 298 websites
  • Pearson correlation score ร— rank: โˆ’0.0655 (n = 286)
  • Served HTML, robots.txt bodies and HTTP statuses retained for replay

Correction history

5 October 2026. An earlier version of this study, published 30 September, reported that the web was "built for humans and badly built for models", on the strength of an AI accessibility score of 30.3 against brand clarity of 66.6.

That finding came from a bug in our scorer. Two checks asked whether robots.txt named GPTBot and ClaudeBot, and counted an unnamed crawler as blocked. An unnamed crawler inherits User-agent: *; not being named means allowed. We were penalising roughly 85% of sites for something that was not happening, and the same page reported both 16.4% blocking GPTBot and 84.8% failing the "GPTBot allowed" check without us noticing the contradiction.

Two further checks were stricter than the specification they implemented. organization_schema required the literal type Organization and rejected NewsMediaOrganization, Corporation and LocalBusiness, all valid subtypes. software_schema required a homepage to be a software listing.

All three are fixed. This page is a fresh sample measured with the corrected scorer, not a patch of the old numbers.