Markdown

Report

How these figures are made

I build every figure on this site from three sources that measure the web in different ways, at different times, for different reasons. Google publishes a range for how many sites use each term, never an exact number. The Web Almanac counts exact sites, inside one sample it crawls once a year. Web Data Commons counts exact sites too, but publishes no licence for its data, so I can cite it and I can never copy it into a download. Two more sources document the web rather than measure it: Google's rich result requirements and Schema.org's own release history. This page explains how those five become the reports you read, where they can be compared and where they cannot, and three places I got a figure wrong and had to fix it.

Where two exact counts meet

Google and the Web Almanac are the only two sources here that both publish an exact-sounding number for the same term, so this is the one place I can check one measurement against another. Google sorts every term into one of six bands and never publishes the count inside a band. The Web Almanac publishes an exact count of the sites it found carrying a term, but only within one of four crawl slices, a mobile home page, a desktop home page, a mobile secondary page, or a desktop secondary page. When a type appears in both, I can ask a narrow question. Does the Web Almanac's exact count actually land inside the band Google says it should? Rarely.

I found 473 real types that appear in both sources for the newest archived month. I compare each one on exactly one crawl slice at a time. Summing the Web Almanac's four slices together would count the same domain more than once, if it carried the type on more than one kind of page. None of the four slices agrees with Google's band more than 15.05% of the time, and the worst of the four manages barely 9.66%.

How often the Web Almanac's exact site count for a type actually lands inside the band Google publishes for the same type, one crawl slice at a time, never summed across them.
Crawl sliceTypes comparedFell inside Google's bandShare inside
Mobile home page4326515.05%
Desktop home page4174811.51%
Mobile secondary page4395011.39%
Desktop secondary page435429.66%

When the two disagree, it is almost always the Web Almanac sitting below Google's band, not above it. Google counts across its whole index. The Web Almanac counts a sample of about 15.4 million mobile origins and 12.2 million desktop ones, one home page and one secondary page each, so most of a site's other pages, and most sites outside the sample entirely, are simply not there to be counted. That gap is sample size, not an error in either source, and no correction factor turns one number into the other. I compare the two here. I never combine them. A blended figure would look more confident than either source really is, and it would be a number you could not check against a download, because no download would agree with it either.

The five sources, and what I may publish

These figures come from five sources, and each one measures or documents a different piece of the web this project reports on. Each licenses its data on its own terms. The table below names each one, what it licenses, and whether I may redistribute its figures at all. I read every licence field straight from the same dataset the reports use, so it cannot say something different here than it says on the page where the figures actually appear.

Every source this site reads, its population, when it was last read, and its own licence.
SourcePopulationLast measuredLicenceMay I redistribute it
Google and the Schema.org community, Schema.org usage statistics datasetWebsites in Google's index. Each site is counted once however many pages carry the markup, so this measures adoption by site, not volume of markup.July 2026No licence statement of its own. Both defensible readings — Apache 2.0 as repository content, or CC BY-SA 3.0 as site data — permit archiving, republishing and commercial use with credit.Yes, with credit
HTTP Archive Web Almanac 2025, SEO chapter resultsAbout 15.4 million mobile and 12.2 million desktop origins from the Chrome UX Report list, home page and one secondary page each.1 July 2025Apache License 2.0.Yes, with credit
Google Search Central rich result requirementsNot a measurement of the web. What Google requires and recommends per rich result feature.11 August 2026CC BY 4.0.Yes, with credit
Schema.org dated release snapshotsNot a measurement of the web. These are the vocabulary itself, used only to date when a term was invented.11 August 2026CC BY-SA 3.0 for the vocabulary and documentation.Yes, with credit
Web Data Commons structured data class statisticsPay-level domains in a Common Crawl corpus. Ten yearly releases, November 2015 to December 2024; the series has stalled.December 2024None. Only the extraction software is licensed. The data and workbooks carry no licence at all.No, cite and link only

The same licence named above for Schema.org's dated releases also covers something more concrete, the description column you read in the explorer table. I reproduce Schema.org's own definitions there. A claim that a term is barely used is worth much less than Schema.org's own words for what the term means, with a link so you can check it yourself. Reusing that one column carries the same share-alike condition Schema.org attaches to it. Reusing anything else on this site does not, because everything else here is my own analysis, not a copy of theirs.

The rules the pipeline enforces

Before any figure reaches a page, a pipeline runs nine numbered checks and refuses to publish if one of them fails. Eight of them write their own sentence into the dataset every time they run, quoted below exactly as generated, ending with the ninth, in the last row. The missing number is the second. It said Web Data Commons' counts and Google's bands must never be combined into one figure, and once the Web Almanac arrived as a third source of exact counts, that rule was folded into the eighth, which now says the same thing for any two of the three.

Every enforced rule with a sentence in the published dataset, in number order.
RuleWhat it enforces
Rule 1Annotation rows removed before any figure was calculated.
Rule 3Terms that cannot be dated are excluded from the adoption-lag figures; 1610 of 2487 were excluded.
Rule 4A term that changes direction between months is reported as unstable, not as movement.
Rule 5Every figure carries the snapshot it came from.
Rule 6Vocabulary counts use Schema.org's own terms only: 937 classes and 1529 properties.
Rule 7Web Almanac type strings were filtered against the vocabulary before joining.
Rule 8The measurements sit in separate blocks and no figure draws on more than one.
Rule 9Every rich result requirement carries its page address and that page's last-updated date.

Traps worth knowing about

A few numbers in this dataset look like they answer the same question and do not. The table below lists the ones most likely to catch you out, with what is actually true beside each one.

The misreadings a number here is most likely to invite, and what actually applies.
TrapWhat is actually true
Using 5,545 as the denominator for any share of terms.5,545 is rows in Google's raw file. 3,058 of them are annotation markers, not terms. The real base is 2,487.
Reading a Web Almanac count as if it confirms or corrects a Google band.They are two different measurements of two different populations. Compare them, as the table above does, and never blend them into one figure.
Treating a missing Web Almanac slice as a count of zero.41 of the 473 real types carry no mobile home page figure at all. A missing measurement is not a measured zero.
Adding the Web Almanac's four crawl slices together for one total.A domain that carries a type on both its home page and a secondary page would be counted twice. Every figure here stays inside the one slice it was measured on.
Mixing up the 1,589 undatable terms with the 1,610 excluded from the adoption chart.1,589 predates the archive itself, classes and properties only. 1,610 also excludes enumeration values, so it is the one that adds back to 2,487 with the 877 datable terms.
Treating the 282 rows of other vocabularies as part of the 682 distinct type strings.682 counts distinct type strings. 282 counts spreadsheet rows across four crawl slices, a different question with a different denominator.
Writing a Google band such as 10M+ as if it were a number.It is a range. Google never publishes the count inside it, here or anywhere else.

What the archive keeps

I keep every monthly file exactly as Google published it, before I touch anything, so a figure on this site can always be checked against what Google actually said that month, rather than against a page that may since have changed. Three months are archived so far, and the count grows by one every time I run the collector again. Web Data Commons stopped publishing after its December 2024 release, so its own dateline simply ends there, rather than being extended or guessed at.

Where I have already been wrong

Three mistakes reached a live page during this project, and each one is worth naming rather than folding quietly into a changelog nobody reads.

My first release of the flagship report mislabelled a property count. Google asks for a property once per feature, so a property that nine features all ask for was counted nine separate times. I added those counts together, called the total 314, and published 109 of them as thinly used. Counting each property once instead gives a different picture, 206 properties present in the usage file and 103 of them thinly used, a share of 50.00%. The corrected figure is larger, not smaller.

The same report also carried a caveat saying some rows were enumeration values rather than properties, and had not been stripped out. I went back to check, expecting to find some. There was nothing to strip. The caveat was true about the risk and wrong about the fact, so I replaced it with the honest one, and I checked thousands of the report's own numbers before and after to make sure nothing else had shifted while I was in there.

A field counting live features included one Google is already phasing out, so it read 33 when 32 features are actually live. That single number fed a stat row on two separate report pages, so both were wrong together until I split it into two fields, one for every feature captured and one for only the live ones, each named for what it actually counts.

Every one of those three is now guarded by a check that stops the build if it happens again, and I proved each check by deliberately breaking it and watching the build fail. That is not the same claim as a clean record. It is the reason to trust the figures here at all, the checks and the corrections, not an absence of mistakes I have not found yet.

The downloads

Every figure on this site is also a file you can open yourself. The dataset behind every report, the term-by-term table behind the explorer, and a spreadsheet version of the whole dataset are all linked below, built from the same pipeline run as the pages themselves.

How this page itself is built

I built this page from the same two files every other report reads, so the sources table, the rule list, and the comparison above update on their own the next time I run the pipeline again. I retype nothing here by hand, which is also the whole reason the rest of this site is worth reading.

Snapshot July 2026. 3 months archived. Built 19 August 2026.

Analysis © Eduard Dziak, licensed under the Apache License 2.0. Please credit eduarddziak.com with a link. Source data: Google and the Schema.org community, Schema.org usage statistics dataset. Rich result requirements: Google Search Central, licensed CC BY 4.0. Term descriptions: Schema.org, licensed CC BY-SA 3.0, each term linking to its own page. Crawl counts: HTTP Archive Web Almanac, licensed Apache License 2.0.