# What actually reads this site Most of this site's traffic is machines. I show which ones, what they ask for, and how easily a chart about them can be poisoned, with every figure's window and method beside it. > Source: https://eduarddziak.com/visibility/ > Window: 2026-08-13 to 2026-08-23, 11 days > Updated: 2026-08-25 > Data: /data/visibility.json and /data/visibility.csv The primary window excludes 2026-08-22, the day one actor spent one hour claiming other operators' crawler names. Every figure names its own window, and Cloudflare-derived figures are estimates under sampling. ## Summary | Measure | Value | Window | | --- | --- | --- | | AI share of page views, with 2026-08-22 | 16.4% | 11 days to 23 Aug 2026 | | AI share of page views, without 2026-08-22 | 3.5% | 10 days to 23 Aug 2026 | | 404 share of page-shaped requests, without 2026-08-22 | 50.5% | 10 days to 23 Aug 2026 | | Impressions with a named query | 0.8% | 21 settled days to 18 Aug 2026 | The two AI-share figures are one finding split in two on purpose, and neither may be quoted without the other. One hour multiplied the apparent AI share 4.7 times. ## Where page views come from | Bucket | Share without 2026-08-22 | Share with 2026-08-22 | | --- | --- | --- | | AI crawlers | 3.5% | 16.4% | | Search engine crawlers | 7.6% | 6.1% | | SEO tool crawlers | 9.4% | 6.9% | | Other recognised bots and tools | 2.1% | 1.6% | | Unrecognised clients, not verifiable as people | 77.5% | 68.9% | The last row is everything my classification list does not recognise. It is never presented as people, and the section "What the unrecognised traffic claims to be" is the published justification. ## The hour that poisoned the chart On 2026-08-22, 557 requests arrived claiming 7 crawler names that belong to 5 operators, OpenAI, Perplexity, Google, Anthropic, Amazon. 539 of them landed in a single hour, every request came from one country, and 95.0% were answered 404. On the evidence I keep, spoofing is the near-certain reading and cannot be proven, because proving it needs the IP addresses behind those requests, and I refuse to hold them. No operator named here is accused of anything. The names were the costume, and an operator whose name was worn is the party imitated, not the party acting. | Measure | Value | | --- | --- | | Crawler names claimed | ChatGPT-User, GPTBot, OAI-SearchBot, PerplexityBot, Google-Extended, ClaudeBot, Amazonbot | | Operators whose names were used | 5 | | Requests across those names | 557 | | Landed in the single peak hour | 539 | | Answered 404 | 95.0% | | Countries of origin | 1 | ### Checked against the operators' published ranges, 22 August 2026 Claiming ChatGPT-User, 0 of 140 requests came from OpenAI's published ranges, and all 140 came from a single address. Googlebot, the same day, 40 of 40 came from Google's published ranges, from 8 addresses. | Crawler claimed | Operator | Requests | Verified | Distinct addresses | Verdict | | --- | --- | --- | --- | --- | --- | | Googlebot | Google | 40 | 40 | 8 | Verified | | bingbot | Microsoft | 8 | 8 | 8 | Verified | | GoogleOther | Google | 4 | 4 | 2 | Verified | | OAI-SearchBot | OpenAI | 48 | 3 | 4 | 3 of 48 verified | | GPTBot | OpenAI | 49 | 1 | 2 | 1 of 49 verified | | ChatGPT-User | OpenAI | 140 | 0 | 1 | Not verified | | PerplexityBot | Perplexity | 46 | 0 | 1 | Not verified | | Amazonbot | Amazon | 185 | Unverifiable | Unverifiable | Unverifiable. This operator publishes no ranges, so no check is possible. | | Google-Extended | Google | 46 | Not a user agent | 1 | Not a user agent. Google documents Google-Extended as a robots.txt control token with no HTTP user agent string, so a request carrying it in its user agent did not come from Google. ([Google's crawler documentation](https://developers.google.com/search/docs/crawling-indexing/google-common-crawlers)) | | ClaudeBot | Anthropic | 43 | Unverifiable | Unverifiable | Unverifiable. This operator publishes no ranges, so no check is possible. | Counts here are requests, not page views: verification filters to eyeball traffic but applies neither the page-view rule nor the own-tooling exclusion. Per the review, verdicts publish from this day only: other covered days carry partial failures whose addresses are discarded by design and can never be re-examined. Two different claims sit in that table, and I keep them apart. That the failing requests came from outside the ranges their operators publish is measured, and can be rechecked against a snapshot of the ranges taken the same day. That somebody wore those names as a costume is a reading, and it stays a reading whatever later data shows. ## What the unrecognised traffic claims to be Unrecognised means my classification list matched no known client, and that is all it means. It is not evidence of a person. 84.3% of unrecognised page views claim to be Chrome. Of those Chrome claims, 43.9% name major versions released before 2024. Chrome 120 went stable on 2023-12-05 and the next major not until 2024-01-23, so a claim of major 120 or older is a claim of a pre-2024 browser. A client claiming Chrome 78, released 2019-10-22, appeared on 11 of 11 observed days. | Claimed major | Stable release date | Date source | Share of Chrome-claiming page views | Note | | --- | --- | --- | --- | --- | | Chrome 131 | 2024-11-12 | [Chromium Dash release schedule](https://chromiumdash.appspot.com/fetch_milestone_schedule?mstone=131) | 39.8% | Mostly the 2026-08-22 and 2026-08-23 bursts | | Chrome 78 | 2019-10-22 | [Chromium Dash release schedule](https://chromiumdash.appspot.com/fetch_milestone_schedule?mstone=78) | 28.0% | Present on 11 of 11 observed days | | Chrome 89 | 2021-03-02 | [Chromium Dash release schedule](https://chromiumdash.appspot.com/fetch_milestone_schedule?mstone=89) | 8.9% | | | Chrome 151 | 2026-07-28 | [Chromium Dash release schedule](https://chromiumdash.appspot.com/fetch_milestone_schedule?mstone=151) | 4.9% | | ## Most page-shaped requests ask for pages this site never served A 404 on this site returns a real HTML page, so my page-view rule counts a page-shaped request answered 404 as a page view. Most of the 404 numerator comes from unrecognised clients, so this is a claim about request behaviour, never about who is behind it. | Window | Share answered 404 | | --- | --- | | without 2026-08-22 | 50.5% | | all eleven days | 62.5% | | without 2026-08-22 and 2026-08-23 | 41.0% | ## Who actually reads the markdown mirror Shares describe observed readers, not appetite, because a client that never discovered the mirror could not have read it. The AI slice rests on a small number of fetches and is held at first-signals strength. For ClaudeBot, the markdown mirror made up 45.7% of its own content fetches in this window. | Bucket | Share of markdown fetches | | --- | --- | | SEO tool crawlers | 50.0% | | Unrecognised clients, not verifiable as people | 18.1% | | Search engine crawlers | 16.2% | | AI crawlers | 15.3% | | Other recognised bots | 0.5% | ## Robots.txt etiquette, per crawler Both columns divide a crawler's robots.txt fetches in this window, first by the HTML pages it took, then by its content fetches with the markdown mirror counted as content. Every name is self-identified, and etiquette observed is not obedience proven. | Crawler | Robots.txt fetches per HTML page taken (10 days, without 2026-08-22) | Per content fetch, markdown counted as content | | --- | --- | --- | | Applebot | 4.0 | 0.8 | | ClaudeBot | 3.7 | 2.0 | | OAI-SearchBot | 1.5 | 1.4 | | Googlebot | 0.7 | 0.6 | | bingbot | 0.1 | 0.1 | | PerplexityBot | 0.1 | 0.1 | | GoogleOther | No robots.txt fetch observed in 10 days. | No robots.txt fetch observed in 10 days. | | GPTBot | No robots.txt fetch observed in 10 days. | No robots.txt fetch observed in 10 days. | | ChatGPT-User | No robots.txt fetch observed in 10 days. | No robots.txt fetch observed in 10 days. | | Amazonbot | No robots.txt fetch observed in 10 days. | No robots.txt fetch observed in 10 days. | | YandexBot | No robots.txt fetch observed in 10 days. | No robots.txt fetch observed in 10 days. | | CCBot | No robots.txt fetch observed in 10 days. | No robots.txt fetch observed in 10 days. | | Baiduspider | No robots.txt fetch observed in 10 days. | No robots.txt fetch observed in 10 days. | ## What Google shows, and what it withholds Across 21 settled days, Google named the search query behind 2 of 245 impressions, which is 0.8%. Query rows are a floor, never a total: the gap is queries Google declined to name. The withholding is a known mechanism rather than something odd about this site, and [Ahrefs measured it at 46.77%](https://ahrefs.com/blog/gsc-anonymized-queries/) across the sites it studied. | Window end | Settled days | Impressions with a named query | All impressions | Share | | --- | --- | --- | --- | --- | | 2026-08-18 | 21 | 2 | 245 | 0.8% | ## First signals, and what I am holding back First signals are observations too few to claim a pattern, published with their size stated so nobody mistakes them for one. The findings below are withheld on purpose, each with the condition its figure waits on, and a row leaves this list only through a review round. | Register id | Held finding | Publishes when | | --- | --- | --- | | VIS-010 | llms.txt before and after | Three to four weeks of after-data; the before-baseline and go-live date are already recorded. | | VIS-008 | Crawl versus credit, Google edition | All joint Search Console days settled. | | VIS-009 | AI training / search / user-fetch split | AI page views in the hundreds; one crawler session currently moves the split by 24 points. | ## How this is measured Two archives sit behind this page, and they answer different questions. The traffic archive holds sampled Cloudflare analytics for this site, and answers what asked for what. The verification archive holds daily counts from checking crawler-name claims against published address ranges, and answers whether a name's addresses matched what its operator publishes. Google Search Console is the third source, and it reports what Google's index did with the site afterwards. Four rules shape every traffic figure. I count eyeball traffic only, which is Cloudflare's own term for requests arriving from outside its network. A page view is a page-shaped path answered 2xx or 404, because a 404 on this site returns a real HTML page, and counting successes only would miss most of what actually happens here. My own fetches and this project's tooling are excluded before anything is counted. And nothing here is a tally. Every Cloudflare-derived figure is an estimate: the dataset is sampled, with intervals up to 3.3 observed. Everything the classification list does not match lands in one bucket, labelled "Unrecognised clients, not verifiable as people". That bucket is never presented as human traffic, in any chart, table or download, and the section on what the unrecognised traffic claims to be shows why. Every client name on this page is the name the client gave itself. A name is not proof of who sent the request, and no operator is accused of anything here. Since the verification collector first ran on 24 August 2026, every request claiming a known crawler's name is compared, as it is collected, against the address ranges that crawler's operator publishes. The comparison happens in memory, only counts are stored, and the addresses are discarded. The ranges are snapshotted the same day, so any verdict can be rechecked against exactly what the operator published then. Five verdict words appear on this page, and they are not interchangeable. Verified means the requests came from inside the operator's published ranges. Not verified means the operator publishes ranges and the requests came from outside them. Unverifiable means the operator publishes no ranges at all, so no check is possible. That is a fact about the operator's documentation, never a verdict on the crawler, and it is why Amazon's Amazonbot and Anthropic's ClaudeBot can show no number. Not covered means the day sits outside verification's reach. Not a user agent means the name is documented as never being an HTTP user agent at all, so no address range could redeem it. Verification cannot prove intent. It shows that an address sat inside or outside a published range, which is measured. Calling the rest a costume is a reading, and I keep the two apart everywhere on this page. The check also reaches back only seven days from its first run, so days before that reach are unverified permanently. No refresh may add a verdict to an earlier day, and no refresh may turn the reading into a proven claim, whatever later days show. Some things are deliberately not collected or published anywhere on this site. IP addresses. Query strings. Referrers. Bot scores. Raw request paths. Raw user-agent strings. A count of distinct addresses is a count, and it is the only address-shaped thing this page will ever show. Figures are shares and ratios rather than absolute visitor counts, because an absolute count invites judging the site's size instead of the finding. An event count appears only as method context inside one named episode, and always with its denominator and its window in the same block. Search Console figures count settled days only. Google keeps revising its most recent days, so an unsettled day is excluded entirely rather than shown provisionally. Days that look odd have public explanations. 2026-08-22 carries the wave hour. 2026-08-23 carried a second burst of requests for pages that never existed, which is why the 404 table names a window without both days. And on 2026-08-24 I changed the site's robots policy, added a tripwire behind it with a block for Bytespider, and took /llms.txt live, so the launch window ends the day before and the etiquette table is a before-baseline. 11 days of detailed data, 13 August 2026 to 23 August 2026. 21 settled Search Console days. Built 25 August 2026. The figures come from this site's own Cloudflare zone analytics and from Google Search Console, read through collectors I run myself. Chrome release dates come from the Chromium Dash schedule service.