How the Census is run
This page is about the survey. The Index methodology is a different document: it says what a score means. This one says how these particular scores were obtained, and what they can and cannot support.
One standard, no special cases
A census site is scanned by the same scanner, against the same 11.0 tests, and scored by the same engine as a site whose owner asked us to look at it. There is no census mode. If the two could diverge, comparing a census figure to your own score would be a lie, and the comparison is most of what makes the census useful.
Every result records which version of the Index it was scored against and the date that version took effect. A standard that moves cannot quietly rewrite a finding published under an older one.
Why we scanned you without asking
public-interest census: a declared bot with a published policy at https://readinessindex.io/bot, honoring robots.txt addressed to AIRBot, reading public pages only. No per-site permission was sought or given.
That sentence is recorded beside every single result, not asserted once here — because "we had permission" and "this is a public measurement" are different claims and only the second one is true.
The crawler is AIRBot. It identifies
itself, it never impersonates another company's user agent, and it obeys any
robots.txt rule addressed to AIRBot — full stop, no
exception for the census. A site that turns us away is recorded as having
turned us away and stays in the population: excluding sites that block
crawlers from a survey about whether sites block crawlers would point the
selection bias straight at the thing being measured.
If you would rather not be in the next one, the bot page says exactly what to add. We would rather you didn't, and we will still count you as having said no.
How hard we crawl
- One request every 3 seconds to any single organization. Never more, whatever else is running.
- 16 organizations at once, which are 16 different hosts and so costs none of them anything.
- 32 URLs per site: the homepage, one page the organization's own navigation says matters most, and 30 sampled from its sitemap and links.
- 429 and 503 are obeyed, with backoff, and a site that keeps saying no is abandoned rather than pressed.
- Every raw response is cached, so a re-run or a re-score costs the site nothing.
Thirty is the per-site page sample. It is not what carries a finding — the n behind a cohort claim is the number of organizations scanned, and it is printed beside every share.
Three ways a sector average can lie
Each of these travels with every published aggregate, because each is a way a headline number can be true and misleading at once.
- Coverage
- The denominator is the number of sites actually scanned, never the cohort size. Rows with no published URL, refusals and failures are counted and named rather than dropped.
- Sample depth
- The 32-URL cap is a ceiling, not a promise. Some sites yield one page. The distribution is published so the cap cannot be read as a claim about what was actually read.
- Renormalization
- Two sites can differ in which tests applied to them at all. Scores are renormalized over the applicable points to stay comparable, and that denominator's distribution is published too — it is the number somebody would need to argue a comparison is unfair.
A census run declares no engagement profile, which is the correct state rather than a gap: the tests that depend on a business fact nobody told us come out N/A and leave the denominator. Nothing is scored zero for a question we never asked.
Where the lists come from
Every cohort is drawn from a published source we are legally free to build on and republish a derivative of. The filtering on top is ours and is disclosed as ours on each cohort page.
| Source | Publisher | Edition | License |
|---|---|---|---|
| Public .gov domain registry (current-full.csv) | Cybersecurity and Infrastructure Security Agency (CISA) | 2026-09-07 | A work of the United States federal government. Not subject to domestic copyright (17 U.S.C. § 105); the repository additionally dedicates it to the public domain. |
| IPEDS HD2023 (institutional characteristics) and DRVEF2023 (enrollment) | National Center for Education Statistics (NCES) | 2023-fall | A work of the United States federal government. Not subject to domestic copyright (17 U.S.C. § 105). |
| IRS Exempt Organizations Business Master File, filtered to health-care organizations by Wikidata classification | Internal Revenue Service (IRS), with Wikidata | 2026-09-07 | A work of the United States federal government. Not subject to domestic copyright (17 U.S.C. § 105). The classification comes from Wikidata under CC0. |
| Exempt Organizations Business Master File Extract | Internal Revenue Service (IRS) | 2026-09-07 | A work of the United States federal government. Not subject to domestic copyright (17 U.S.C. § 105). |
| Tranco top sites list | Tranco (Le Pochat et al., NDSS 2019) | 2026-09-06 | Published for research use and redistributable. Each daily list has its own permanent identifier, which is what a citation has to name — 'the Tranco list' on its own is not reproducible. |
| List of S&P 500 companies | Wikipedia contributors | 2026-09-08 | CC BY-SA 4.0. The constituent list is Wikipedia's compilation, which is share-alike — the same license this cohort goes out under. “S&P 500” is a registered trademark of S&P Dow Jones Indices LLC; naming the constituents is descriptive use, and nothing here is endorsed by or affiliated with S&P. |
| Tranco top sites list | Tranco (Le Pochat et al., NDSS 2019) | 2026-09-06 | Published for research use and redistributable. Each daily list has its own permanent identifier, which is what a citation has to name — 'the Tranco list' on its own is not reproducible. |
| Wikidata: employer identification number (P1297) and official website (P856) | Wikimedia Foundation and Wikidata contributors | 2026-09-07 | CC0 1.0 Universal. Wikidata's data is dedicated to the public domain and carries no attribution requirement, though the query and its date are recorded here because a result set that cannot be reproduced is not a citation. |
And what we turned down
Better lists exist. We cannot republish a cohort derived from them, and a source we cannot cite is a source that makes the census uncheckable.
- US News Best Colleges — Copyrighted compilation; ranks prestige, which the Census does not measure.
- Forbes / NPT largest charities — Copyrighted compilations.
- American Hospital Association directory — Licensed; redistribution prohibited.
- BuiltWith / Wappalyzer technology datasets — Licensed; the terms do not permit republishing a derived list.
What we publish, and what we never will
Published: everything measured, named, with the SHA-256 of each result recorded at the moment the scan finished. Nothing is anonymized. A finding nobody can check is an opinion, and the chain from a number on a page to a file to a run is the whole argument for believing any of this.
Never published: a contact route for anyone. We keep a way to reach the organizations we scan, so that a body surveyed without being asked can at least argue with the result — role addresses over people, always sourced, never guessed from a staff directory. That store sits outside the published dataset by construction, not by a filter somebody has to remember.
The cohorts
| Cohort | What it is | Size |
|---|---|---|
| state-portals | The primary web portal of each US state government. | 50 |
| universities | The 100 largest US universities by total enrollment. | 100 |
| health-systems | The 50 largest US nonprofit health systems by total revenue. | 50 |
| nonprofits | The 100 largest US nonprofits by total revenue. | 100 |
| sp500 | The 500 companies in the S&P 500. | 500 |
| wordpress-100 | The 100 highest-ranked WordPress sites in the Tranco list. | 100 |
| drupal-100 | The 100 highest-ranked Drupal sites in the Tranco list. | 100 |
| Total | 941 distinct websites | 1000 |
Why those two numbers differ
The cohorts add up to 1000 rows across 941 distinct websites, because 51 sites belong to more than one of them. A children's hospital is one of the largest nonprofits and one of the largest health systems. A university can run Drupal. Kaiser Permanente files as six separate nonprofits and answers on one address.
Those rows stay in both cohorts, and the site is read once. The cohort is the unit of analysis here — the question this census answers is how state portals compare with universities, or nonprofits with listed companies, and a member sitting in two populations does not damage either comparison. What it would damage is a claim about the whole, so we do not make one: this is a census of 1000 rows over 941 websites, and never 1000 distinct organizations.
Is your site ready for AI?
Scan a page, get a number you can verify. Free.