Structured data
What the web actually marks up
Google publishes a file every month counting how many websites use each Schema.org term. Almost every summary of that file you will find repeats the same counting error, because more than half of it is not vocabulary at all. I archive every month, take those rows out before counting, and show you both figures so you can see the size of the difference for yourself.
- Real terms2,487
- Classes937
- Properties1,529
- SnapshotJuly 2026
The number most people publish is wrong
July 2026 · six ranges, no exact numbersGoogle and the Schema.org community, Schema.org usage statistics dataset. Websites in Google's index.
The file holds 5,545 rows, but 3,058 of them end in -input or -output. Those are markers from Schema.org's Actions syntax, exactly two for every real property, and you would never write one into a page. Almost all of them sit in the smallest band, so leaving them in drags the whole distribution downwards and makes genuinely rare terms look ordinary.
| What is counted | Rows | Share under 1,000 websites |
|---|---|---|
| Real terms only | 2,487 | 46.68% |
| Every row in the file | 5,545 | 76.05% |
| Annotation rows removed first | 3,058 | Not vocabulary |
How adoption is distributed
July 2026 · six ranges, no exact numbersGoogle and the Schema.org community, Schema.org usage statistics dataset. Websites in Google's index.
These are ranges rather than exact counts. Google sorts every term into one of six bands and never publishes the figure inside a band, so a term you can see sits between one and ten million websites has no public exact number anywhere, on this page or any other.
| Websites using the term | Terms | Share of real terms |
|---|---|---|
| < 1K | 1,161 | 46.68% |
| 1K - 10K | 575 | 23.12% |
| 10K - 100K | 423 | 17.01% |
| 100K - 1M | 173 | 6.96% |
| 1M - 10M | 104 | 4.18% |
| 10M+ | 51 | 2.05% |
| All real terms | 2,487 | 100.00% |
Three measurements, three different webs
Three organisations count structured data and none of them count the same thing. Google measures its whole index and publishes ranges. The Web Almanac measures about 15 million origins and publishes exact counts. Web Data Commons measured a Common Crawl corpus and stopped in December 2024. I never combine any two of them into one figure, because there is no honest way to do it. An exact count of one set of websites is not the number hiding inside a range measured across another.
| Source | What it counts | Grain | Last measured | May be republished |
|---|---|---|---|---|
| Google and the Schema.org community, Schema.org usage statistics dataset | Websites in Google's index. Each site is counted once however many pages carry the markup, so this measures adoption by site, not volume of markup. | six ranges, no exact numbers | July 2026 | Yes, with credit |
| HTTP Archive Web Almanac 2025, SEO chapter results | About 15.4 million mobile and 12.2 million desktop origins from the Chrome UX Report list, home page and one secondary page each. | exact site counts | 1 July 2025 | Yes, with credit |
| Web Data Commons structured data class statistics | Pay-level domains in a Common Crawl corpus. Ten yearly releases, November 2015 to December 2024; the series has stalled. | exact site counts | December 2024 | No. Cited and linked only |
Reports
Questions
- How many Schema.org terms are actually used on the web?
- Google's July 2026 file covers 2,487 real terms. That breaks down into 937 classes, which are the types you mark something up as, 1,529 properties, which are the fields inside them, 20 enumeration values, and one term Schema.org has never actually published. 46.68% of those terms appear on fewer than a thousand websites, and only 51 of them reach ten million or more.
- Why does this page show a different percentage from other sites?
- Because 3,058 rows in the file are not vocabulary terms at all. They end in -input or -output, there are exactly two for every real property, and they come from Schema.org's Actions syntax rather than from anything you would mark up. Leave them in and you get 76.05% of terms below a thousand websites, which is the figure you will find published elsewhere. Take them out first and you get 46.68%. You can check this yourself against the file, which is why I publish both numbers rather than only the corrected one.
- Are these exact numbers of websites?
- No, and nobody has them. Google publishes six bands rather than exact counts, so a term is known to sit somewhere between one thousand and ten thousand websites without the real figure ever being made public. If you find a page showing an exact count for Google's data, it has either been invented or taken from a different measurement of a different set of websites.
- Where does the data come from?
- Google publishes it monthly in the Schema.org repository. I keep a checksummed copy of every file, so the record survives even if the source stops. The rich result requirements are separate. I read those from the markup of Google Search Central's own pages on 11 August 2026, rather than typing them from memory, and each one carries the address and date it came from.
- Can I reuse these numbers?
- Yes. You can download the full dataset as JSON or CSV at a stable address, and the analysis is licensed under the Apache License 2.0. All I ask is a credit to eduarddziak.com with a link. One source works differently. Web Data Commons publishes no licence for its data at all, so I cite and link its figures but never include them in a download.
- How often does this update?
- Google publishes a new file each month, and I archive it, check it against the previous months, and rebuild every figure from the archive. You are looking at July 2026, rebuilt on 12 August 2026.
How this is measured
Every month Google publishes a file counting how many websites in its index use each Schema.org term, and every month I archive a copy of it. The figures you see here are rebuilt from that archive rather than from a live download, so any number on this page can be traced back to the month it came from. I am working from the July 2026 file, and the archive behind it holds 3 monthly snapshots so far.
Three organisations measure how much of the web uses Schema.org, and they do not measure the same thing. Google covers its whole index but publishes six bands instead of exact counts, so a term known to sit between one and ten million websites has no public figure inside that range. The Web Almanac publishes exact counts, but only for the origins it crawled, which makes it precise about a sample and silent about everything outside it. Where Google is complete is the vocabulary itself. All 1,529 defined properties appear in its file, so no term is missing from the census even though no term has an exact count.
I apply one correction before counting anything, and it is the reason this page disagrees with most published summaries. The file holds 5,545 rows, but 3,058 of them end in -input or -output. Those are markers from Schema.org's Actions syntax, exactly two for every real property, and nobody marks up a website with them. Almost all of them sit in the smallest band, so counting them drags the average down and makes genuinely rare terms look ordinary. Over all rows you get 76.05% of terms below a thousand websites. Over the 2,487 real terms you get 46.68%, which is the figure I publish.
The Web Almanac figures come from a crawl on 1 July 2025, a different population from Google's index and more than a year earlier. I never combine its counts with Google's bands into a single number, because an exact count of one set of websites is not the hidden figure inside a range measured across another.
You can download the same figures as JSON or CSV and check any number on this page against them.
Snapshot July 2026. 3 months archived. Built 12 August 2026.
Analysis © Eduard Dziak, licensed under the Apache License 2.0. Please credit eduarddziak.com with a link. Source data: Google and the Schema.org community, Schema.org usage statistics dataset. Rich result requirements: Google Search Central, licensed CC BY 4.0. Term descriptions: Schema.org, licensed CC BY-SA 3.0, each term linking to its own page.