Structured data
What the web actually marks up
Google publishes a monthly file counting how many websites use any given Schema.org term. Almost all the summaries of that file you will find repeat the same counting error, because more than half of it is not vocabulary at all. I archive all of them, take those rows out before counting, and show you both figures. The difference is not small.
- Real terms2,487of 5,545 rows in the file
- Classes937
- Properties1,529
- SnapshotJuly 2026
Key findings
- 46.68% of the 2,487 real Schema.org terms are used by fewer than 1,000 websites.
- The figure most sites republish, 76.05%, counts 3,058 annotation rows that are not vocabulary.
- Only 51 terms are on ten million websites or more.
- Google publishes six usage bands and no exact count inside any of them.
- The vocabulary Google measures holds 937 classes, 1,529 properties, and 20 enumeration values.
- Three sources measure structured data on the web, and no two of their figures can be combined.
46.68% of terms sit under 1,000 websites, not 76.05%
July 2026 · six ranges, no exact numbersGoogle and the Schema.org community, Schema.org usage statistics dataset. Websites in Google's index.
The two bars show the same measurement counted two ways. The dashed bar is the wrong way.
| What is counted | Rows | Share under 1,000 websites |
|---|---|---|
| Real terms only | 2,487 | 46.68% |
| Every row in the file | 5,545 | 76.05% |
| Annotation rows removed first | 3,058 | Not vocabulary |
The file holds 5,545 rows, but 3,058 of them end in -input or -output. Those are markers from Schema.org's Actions syntax, two per real property, and you would not write one into a page. Almost all of them sit in the smallest band, so leaving them in drags the whole distribution downwards and makes genuinely rare terms look ordinary. I expected the correction to move the headline by a point or two. It moved it by more than that. If the figure you have seen elsewhere is the higher one, the markers are the whole difference.
Only 51 terms are on ten million websites or more
July 2026 · six ranges, no exact numbersGoogle and the Schema.org community, Schema.org usage statistics dataset. Websites in Google's index.
The bars count terms in Google's six usage bands for July 2026.
| Websites using the term | Terms | Share of real terms |
|---|---|---|
| < 1K | 1,161 | 46.68% |
| 1K - 10K | 575 | 23.12% |
| 10K - 100K | 423 | 17.01% |
| 100K - 1M | 173 | 6.96% |
| 1M - 10M | 104 | 4.18% |
| 10M+ | 51 | 2.05% |
| All real terms | 2,487 | 100.00% |
Nearly half of the vocabulary sits in the smallest band. The counts fall away fast above it. Adoption concentrates in a small set of terms the whole web shares. That is the shape I expected. The steepness still surprised me. For your own markup it means the safe terms are few and well known. Anything outside them puts you in thin company.
The labels are ranges, because Google does not publish the exact figure inside a band.
Three measurements, three different webs
The dateline marks when the three sources measured. It shows no quantities at all, on purpose.
| Source | What it counts | Grain | Last measured | May be republished |
|---|---|---|---|---|
| Google and the Schema.org community, Schema.org usage statistics dataset | Websites in Google's index. Each site is counted once however many pages carry the markup, so this measures adoption by site, not volume of markup. | six ranges, no exact numbers | July 2026 | Yes, with credit |
| HTTP Archive Web Almanac 2025, SEO chapter results | About 15.4 million mobile and 12.2 million desktop origins from the Chrome UX Report list, home page and one secondary page each. | exact site counts | 1 July 2025 | Yes, with credit |
| Web Data Commons structured data class statistics | Pay-level domains in a Common Crawl corpus. Ten yearly releases, November 2015 to December 2024; the series has stalled. | exact site counts | December 2024 | No. Cited and linked only |
Three sources, three webs. Google measures its whole index and publishes ranges. The Web Almanac measures about 15 million origins once a year and publishes exact counts. Web Data Commons measured a Common Crawl corpus and stopped in December 2024. The one I lean on is Google, because it is the only one that covers the whole vocabulary. The one I would reach for to sanity-check a single term is the Web Almanac. If you see a page blend the two into one number, nobody measured that number.
I do not combine any two of these into one figure, because an exact count of one set of websites is not the number hiding inside a range measured across another.
Reports
- The properties Google recommends and almost nobody uses103 of 206 recommended properties are on fewer than 100,000 websites.
- Every Schema.org term, and how much the web uses itAll 2,487 terms in one table you can search and sort, with Schema.org's own definitions.
- What each industry can mark up, and what it actually does578 of the 937 classes sit in a real industry, and four of the thirteen industries have no rich result Google documents at all.
- What changed, and how long adoption takes136 terms ended three archived months in a different usage band, 87 have been superseded with a documented replacement, and 22 are already on 100,000 websites before Schema.org has ratified them.
- The properties Google requires, and what they rest on30 of 32 live features need a required property, and one of them rests on a property used by fewer than 1,000 websites.
- Nearly a third of the type names on the web are not real209 of the 682 distinct type strings the Web Almanac's 2025 crawl found are not real Schema.org types, 30.65% of everything written.
- Where the vocabulary lives, and what makes a term catch onHealth is the largest extension field Schema.org publishes and the least used, 75.00% of its terms under a thousand websites. Properties attached to more types get used more, without exception.
- How these figures are madeThe five sources and their licences, the rules the pipeline enforces, and three mistakes I found in my own published figures and fixed.
Questions
- How many Schema.org terms are actually used on the web?
- Google's July 2026 file covers 2,487 real terms. That breaks down into 937 classes, which are the types you mark something up as, 1,529 properties, which are the fields inside them, 20 enumeration values, and one term Schema.org has not actually published. 46.68% of those terms appear on fewer than a thousand websites. Only 51 of them reach ten million or more.
- Why does this page show a different percentage from other sites?
- Because 3,058 rows in the file are not vocabulary terms at all. They end in -input or -output, two per real property, and they come from Schema.org's Actions syntax, not from anything you would mark up. Leave them in and you get 76.05% of terms below a thousand websites, which is the figure you will find published elsewhere. Take them out first and you get 46.68%. You can check this yourself against the file. That is why I publish both numbers and not only the corrected one.
- Are these exact numbers of websites?
- No, and nobody has them. Google publishes six bands, not exact counts, so a term is known to sit somewhere between one thousand and ten thousand websites without the real figure ever being made public. If you find a page showing an exact count for Google's data, it has either been invented or taken from a different measurement of a different set of websites.
- Where does the data come from?
- Google publishes it monthly in the Schema.org repository. I keep a checksummed copy of the files, so the record survives even if the source stops. The rich result requirements are separate. I read those from the markup of Google Search Central's own pages on 11 August 2026, not from memory, and each one carries the address and date it came from.
- Can I reuse these numbers?
- Yes. You can download the full dataset as JSON or CSV at a stable address, and the analysis is licensed under the Apache License 2.0. All I ask is a credit to eduarddziak.com with a link. One source works differently. Web Data Commons publishes no licence for its data at all, so I cite and link its figures and keep them out of the downloads.
- How often does this update?
- Google publishes a new file monthly, and I archive it, check it against the previous months, and rebuild the figures from the archive. You are looking at July 2026, rebuilt on 19 August 2026.
How this is measured
Google publishes a monthly file counting how many websites in its index use any given Schema.org term. I archive a copy the same month. I rebuild the figures you see here from that archive, not from a live download, so you can trace any number on this page back to the month it came from. I am working from the July 2026 file. The archive behind it holds 3 monthly snapshots so far.
Three organisations measure how much of the web uses Schema.org. They do not measure the same thing. Google covers its whole index but publishes six bands instead of exact counts, so a term known to sit between one and ten million websites has no public figure inside that range. The Web Almanac publishes exact counts, but only for the origins it crawled, which makes it precise about a sample and silent about the rest. Where is Google complete? In the vocabulary itself. All 1,529 defined properties appear in its file, so no term is missing from the census even though no term has an exact count. When you check a figure of mine, the first thing to ask is which of the three it came from.
I apply one correction before counting anything. It is the reason my figures disagree with most published summaries. The file holds 5,545 rows, but 3,058 of them end in -input or -output. Those are markers from Schema.org's Actions syntax, two per real property, and nobody marks up a website with them. Almost all of them sit in the smallest band, so counting them drags the average down and makes genuinely rare terms look ordinary. Over all rows you get 76.05% of terms below a thousand websites. Over the 2,487 real terms you get 46.68%, which is the figure I publish. If you have quoted the higher one, you have quoted the markers.
The Web Almanac figures come from a crawl on 1 July 2025, a different population from Google's index and more than a year earlier. I do not combine its counts with Google's bands into a single number, because an exact count of one set of websites is not the hidden figure inside a range measured across another.
You can download the same figures as JSON or CSV and check any number on this page against them.
Snapshot July 2026. 3 months archived. Built 19 August 2026.
Analysis © Eduard Dziak, licensed under the Apache License 2.0. Please credit eduarddziak.com with a link. Source data: Google and the Schema.org community, Schema.org usage statistics dataset. Rich result requirements: Google Search Central, licensed CC BY 4.0. Term descriptions: Schema.org, licensed CC BY-SA 3.0, each term linking to its own page. Crawl counts: HTTP Archive Web Almanac, licensed Apache License 2.0.