Markdown

Visibility

What actually reads this site

Most of this site's traffic is machines, and I can show which ones, what they ask for, and how easily a chart about them can be poisoned. Where an operator publishes the addresses its crawler uses, I check the claim against them, so some of what follows is not just the name a client gave itself. Every figure below carries its own observation window and its method beside it, and each one can be recomputed from the archive it came from.

One day is treated specially throughout. On 2026-08-22, one actor spent one hour claiming crawler names that belong to 5 other operators, so the primary window excludes that day, and every strip below says either all 11 days or without 2026-08-22. The hour that poisoned the chart tells the full story.

The first two cells are one finding split in two on purpose, and neither may be quoted without the other. The dashed border marks the reading a naive chart would have published.

The split, with and without one bad hour

11 days of detailed data · to 23 Aug 2026 · sampled, estimates not talliesShare of page views. A page view is a page-shaped path answered 2xx or 404, eyeball traffic only, this project's own tooling excluded.

The five rows split every page view by what the client claimed to be. Each row shows the share without 2026-08-22 first, because that is the primary window, and the share with it underneath, dashed, as a warning rather than a finding. The AI row is the live demonstration, because one bad hour moves it from 3.5%to 16.4%. The last row is everything my classification list does not recognise, and I never present it as people.

Where page views come from, both windows.
BucketShare without 2026-08-22Share with 2026-08-22
AI crawlers3.5%16.4%
Search engine crawlers7.6%6.1%
SEO tool crawlers9.4%6.9%
Other recognised bots and tools2.1%1.6%
Unrecognised clients, not verifiable as people why this is not called users77.5%68.9%

The same figures are published as JSON and CSV, so you can check any number on this page against them.

The hour that poisoned the chart

11 days of detailed data · to 23 Aug 2026 · sampled, estimates not talliesShare of page views. A page view is a page-shaped path answered 2xx or 404, eyeball traffic only, this project's own tooling excluded.

On 2026-08-22, 557 requests arrived claiming 7 crawler names that belong to 5 operators, being OpenAI, Perplexity, Google, Anthropic, Amazon. 539 of them landed in a single hour, every request came from one country, and 95.0% were answered 404, meaning they asked for pages this site has never served.

The obvious question is who sent them. On the evidence I keep, spoofing is the near-certain reading and cannot be proven, because proving it needs the IP addresses behind those requests, and I refuse to hold them. No operator named here is accused of anything. The names were the costume, and an operator whose name was worn is the party imitated, not the party acting.

That single hour multiplied the apparent AI share 4.7 times. Any chart that trusts user-agent strings can be moved the same way, by one actor, in one hour.

The wave, in figures. Counts here are event counts inside one named episode.
MeasureValue
Crawler names claimedChatGPT-User, GPTBot, OAI-SearchBot, PerplexityBot, Google-Extended, ClaudeBot, Amazonbot
Operators whose names were used5
Requests across those names557 requests
Landed in the single peak hour539 requests
Answered 40495.0%
Countries of origin1

Checked against the operators' published ranges, 22 August 2026

2026-08-22 · verification archive · requests, not page viewsCounts here are requests, not page views: verification filters to eyeball traffic but applies neither the page-view rule nor the own-tooling exclusion.

Claiming ChatGPT-User0 of 140 requests came from OpenAI's published ranges. All 140 came from a single address.
Googlebot, the same day40 of 40 came from Google's published ranges, from 8 addresses.
Every crawler name in that day's verification file, sorted by verified requests.
Crawler claimedOperatorRequestsVerifiedDistinct addressesVerdict
GooglebotGoogle40408Verified
bingbotMicrosoft888Verified
GoogleOtherGoogle442Verified
OAI-SearchBotOpenAI48343 of 48 verified
GPTBotOpenAI49121 of 49 verified
ChatGPT-UserOpenAI14001Not verified
PerplexityBotPerplexity4601Not verified
AmazonbotAmazon185Unverifiable. This operator publishes no ranges, so no check is possible.
Google-ExtendedGoogle46Not a user agent1Not a user agent. Google documents Google-Extended as a robots.txt control token with no HTTP user agent string, so a request carrying it in its user agent did not come from Google. Google's crawler documentation
ClaudeBotAnthropic43Unverifiable. This operator publishes no ranges, so no check is possible.

Two different claims sit in that table, and I keep them apart. That the failing requests came from outside the ranges their operators publish is measured, and you can recheck it against a snapshot of the ranges taken that same day. That somebody wore those names as a costume is a reading, and it stays a reading whatever later data shows. The check itself was working, because OpenAI's own range file verified other requests on the same day it verified 0 of 140 for ChatGPT-User. A range file that verifies nothing cannot be told apart from a broken lookup, and one that verifies some requests and not others can.

Per the review, verdicts publish from this day only: other covered days carry partial failures whose addresses are discarded by design and can never be re-examined.

Verification first ran on 24 August 2026 and could reach back only seven days, so earlier days, including the start of the traffic window on 13 August 2026, are unverified and will stay that way. No later refresh can change that.

What the unrecognised traffic claims to be

11 days of detailed data · to 23 Aug 2026 · sampled, estimates not talliesClaimed Chrome major versions among unrecognised page views. Each bar's denominator is the Chrome-claiming slice, named beside the chart.

Unrecognised means my classification list matched no known client, and that is all it means. It is not evidence of a person, which is why the largest slice of this site's traffic is labelled "Unrecognised clients, not verifiable as people" and never anything warmer. This section is what that label links to.

84.3% of unrecognised page views claim to be Chrome. Of those Chrome claims, 43.9% name major versions released before 2024. The two figures have different denominators, so they stay in prose and never share an axis. Chrome 120 went stable on 5 December 2023 and the next major not until 23 January 2024, so a claim of major 120 or older is a claim of a pre-2024 browser.

A client claiming Chrome 78, released 22 October 2019, appeared on 11 of 11 observed days.

Release dates come from the Chromium Dash release schedule, the schedule service the Chromium project itself runs, and each date in the table links to the exact record. The dashed bar is a warning, because that spike is largely two burst days, not a finding about ordinary traffic.

What the unrecognised bucket claims to run.
Claimed majorStable releaseDate sourceShare of Chrome-claiming page viewsNote
Chrome 13112 Nov 2024Chromium Dash release schedule39.8%Mostly the 2026-08-22 and 2026-08-23 bursts
Chrome 7822 Oct 2019Chromium Dash release schedule28.0%Present on 11 of 11 observed days
Chrome 892 Mar 2021Chromium Dash release schedule8.9%
Chrome 15128 Jul 2026Chromium Dash release schedule4.9%

Most page-shaped requests ask for pages this site never served

11 days of detailed data · to 23 Aug 2026 · sampled, estimates not talliesShare of page-shaped requests answered 404, three windows, each named beside its figure.

A 404 on this site returns a real HTML page, so my page-view rule counts a page-shaped request answered 404 as a page view. The rule was set by checking it against Cloudflare's own page-view total for the same days, because counting successes only would miss most of what actually asks for pages here. Most of the 404 numerator comes from unrecognised clients, so this is a claim about request behaviour, never about who is behind it.

Both bars are honest findings with stated windows, so neither is dashed. 2026-08-23 carried a second, smaller burst of requests for pages that never existed, which is why the table also names the share without both days.

Requests for pages that never existed, three windows.
WindowShare answered 404
without 2026-08-2250.5%
all eleven days62.5%
without 2026-08-22 and 2026-08-2341.0%

Who actually reads the markdown mirror

10 days, without 2026-08-22 · to 23 Aug 2026 · sampled, estimates not talliesShare of markdown mirror fetches by classified client bucket. My own fetches are excluded.

Every page here has a markdown twin at the same address ending in .md, published for machines. Three limits sit on this chart, in the open. The AI slice rests on a small number of fetches and I hold it at first-signals strength. A client that never discovered the mirror could not have read it, so shares describe observed readers, not appetite. And my own fetches are excluded before anything is counted.

ClaudeBot is the wrinkle worth naming. In this window the markdown mirror made up 45.7% of its own content fetches, so nearly half of what it took from this site was the mirror rather than the HTML.

Who fetches the markdown mirror.
BucketShare of markdown fetches
SEO tool crawlers50.0%
Unrecognised clients, not verifiable as people why this is not called users18.1%
Search engine crawlers16.2%
AI crawlers15.3%
Other recognised bots0.5%

I publish a markdown mirror of this page too, at /visibility.md, so a later window can include the readers of the page you are reading now.

Robots.txt etiquette, per crawler

10 days, without 2026-08-22 · before the 2026-08-24 policy changeRobots.txt fetches per unit of content taken, both denominators shown for every crawler.

Both columns divide the same numerator, a crawler's robots.txt fetches in this window. The first divides by the HTML pages that crawler took. The second divides by its content fetches with the markdown mirror counted as content, which is the stricter reading. Showing only the larger ratio would be a cherry-pick, so every row carries both. This is a table rather than a chart on purpose, because an observed zero deserves words, never a zero-length bar.

Every name here is self-identified, and etiquette observed is not obedience proven. On 2026-08-24 I changed the site's robots policy, so this table is the before-baseline, and a later refresh adds the after columns beside these.

Robots.txt fetches per fetch of content, 10 days without 2026-08-22, before the 2026-08-24 policy change.
CrawlerPer HTML page takenrobots.txt fetches ÷ pagesPer content fetchmarkdown counted as content
Applebot4.00.8
ClaudeBot3.72.0
OAI-SearchBot1.51.4
Googlebot0.70.6
bingbot0.10.1
PerplexityBot0.10.1
GoogleOtherNo robots.txt fetch observed in 10 days.
GPTBotNo robots.txt fetch observed in 10 days.
ChatGPT-UserNo robots.txt fetch observed in 10 days.
AmazonbotNo robots.txt fetch observed in 10 days.
YandexBotNo robots.txt fetch observed in 10 days.
CCBotNo robots.txt fetch observed in 10 days.
BaiduspiderNo robots.txt fetch observed in 10 days.

What Google shows, and what it withholds

21 settled days · to 18 Aug 2026 · Google Search ConsoleImpressions whose query Google named, over all impressions, settled days only.

Across 21 settled days, Google named the search query behind 2 of 245 impressions, which is 0.8%. The rest are queries Google declined to name. An impression is one appearance of this site in someone's results, and a day only counts here once Search Console has stopped revising its figures, which is what I call settled.

The withholding is a known mechanism rather than something odd about this site. Ahrefs measured it at 46.77% across the sites it studied, so a site this small reads as the far end of the same mechanism, where almost every query disappears. That external figure stays in prose and never enters a chart or table beside this site's own. The interesting number is not the share but how it moves, so the table below gains a row at each refresh and only ever holds windows that already exist.

The series, as it grows. Query rows are a floor, never a total: the gap is queries Google declined to name.
Window endSettled daysNamed-query impressionsAll impressionsShare
18 Aug 20262122450.8%

First signals, and what I am holding back

First signals are observations too few to claim a pattern, published with their size stated so nobody mistakes them for one. At launch, exactly one figure on this page is held at that strength, the AI slice of the markdown mirror's readers above, and nothing else here is built on it.

The findings below are withheld on purpose, which is itself a statement about method rather than a promise of coming content. Each row names the condition its figure waits on, and a row leaves this list only through a review round.

  • llms.txt before and afterThree to four weeks of after-data; the before-baseline and go-live date are already recorded.
  • Crawl versus credit, Google editionAll joint Search Console days settled.
  • AI training / search / user-fetch splitAI page views in the hundreds; one crawler session currently moves the split by 24 points.

How this is measured

Two archives sit behind this page, and they answer different questions. The traffic archive holds sampled Cloudflare analytics for this site, and answers what asked for what. The verification archive holds daily counts from checking crawler-name claims against published address ranges, and answers whether a name's addresses matched what its operator publishes. Google Search Console is the third source, and it reports what Google's index did with the site afterwards.

Four rules shape every traffic figure. I count eyeball traffic only, which is Cloudflare's own term for requests arriving from outside its network. A page view is a page-shaped path answered 2xx or 404, because a 404 on this site returns a real HTML page, and counting successes only would miss most of what actually happens here. My own fetches and this project's tooling are excluded before anything is counted. And nothing here is a tally. Every Cloudflare-derived figure is an estimate: the dataset is sampled, with intervals up to 3.3 observed.

Everything the classification list does not match lands in one bucket, labelled "Unrecognised clients, not verifiable as people". That bucket is never presented as human traffic, in any chart, table or download, and the section on what the unrecognised traffic claims to be shows why.

Every client name on this page is the name the client gave itself. A name is not proof of who sent the request, and no operator is accused of anything here.

Since the verification collector first ran on 24 August 2026, every request claiming a known crawler's name is compared, as it is collected, against the address ranges that crawler's operator publishes. The comparison happens in memory, only counts are stored, and the addresses are discarded. The ranges are snapshotted the same day, so any verdict can be rechecked against exactly what the operator published then.

Five verdict words appear on this page, and they are not interchangeable. Verified means the requests came from inside the operator's published ranges. Not verified means the operator publishes ranges and the requests came from outside them. Unverifiable means the operator publishes no ranges at all, so no check is possible. That is a fact about the operator's documentation, never a verdict on the crawler, and it is why Amazon's Amazonbot and Anthropic's ClaudeBot can show no number. Not covered means the day sits outside verification's reach. Not a user agent means the name is documented as never being an HTTP user agent at all, so no address range could redeem it.

Verification cannot prove intent. It shows that an address sat inside or outside a published range, which is measured. Calling the rest a costume is a reading, and I keep the two apart everywhere on this page. The check also reaches back only seven days from its first run, so days before that reach are unverified permanently. No refresh may add a verdict to an earlier day, and no refresh may turn the reading into a proven claim, whatever later days show.

Some things are deliberately not collected or published anywhere on this site. IP addresses. Query strings. Referrers. Bot scores. Raw request paths. Raw user-agent strings. A count of distinct addresses is a count, and it is the only address-shaped thing this page will ever show.

Figures are shares and ratios rather than absolute visitor counts, because an absolute count invites judging the site's size instead of the finding. An event count appears only as method context inside one named episode, and always with its denominator and its window in the same block.

Search Console figures count settled days only. Google keeps revising its most recent days, so an unsettled day is excluded entirely rather than shown provisionally.

Days that look odd have public explanations. 2026-08-22 carries the wave hour. 2026-08-23 carried a second burst of requests for pages that never existed, which is why the 404 table names a window without both days. And on 2026-08-24 I changed the site's robots policy, added a tripwire behind it with a block for Bytespider, and took /llms.txt live, so the launch window ends the day before and the etiquette table is a before-baseline.

You can download every figure on this page as JSON or CSV and check any number against them. The build cost of this site is measured the same way on the build cost page, and the structured data of the wider web on the structured data reports.

11 days of detailed data, 13 August 2026 to 23 August 2026. 21 settled Search Console days. Built 25 August 2026.

The figures come from this site's own Cloudflare zone analytics and from Google Search Console, read through collectors I run myself. Chrome release dates come from the Chromium Dash schedule service.