# The AI Readiness Index

**Version 11.0 — specification**
Effective 2026-09-24.

**License:** This document and the tests it defines are published under
[CC BY-SA 4.0](https://creativecommons.org/licenses/by-sa/4.0/). Copy it, adapt
it, build your own scanner from it — for any purpose, free, forever — as long
as you credit TEN7 as the creator of the AI Readiness Index and license
what you build the same way. The scanner that computes a score from this
document is separate, proprietary software. Full terms at
[readinessindex.io/license](https://readinessindex.io/license).

## What this is

This document defines what a website has to do to score 100 on the AI Readiness Index. It is written for two readers at once: the engineer building the scanner in this repository, and the consultant explaining a score to a client.

Every check has an ID, a weight, a detection method, and a scoring rubric. If a check is not described here, the scanner should not score it.

We built this because "is your site ready for AI" is currently an unanswerable question. Vendors assert things. Nobody measures. Our answer is a number you can reproduce, defend line by line, and re-run in six months to show what changed.

The Index measures one thing: **can machines find, read, trust, and use this website?** It does not measure design quality, marketing effectiveness, or whether the content is any good. A beautiful site can score 12. A plain one can score 94.

## How to read a score

A result has three parts, and all three matter.

**The score (0–100).** A single weighted number across eight dimensions.

**The gate status.** Two checks are gates. If a site blocks AI crawlers in `robots.txt` or rejects them at the CDN, `gated` is `true` — a separate fact reported beside the score, not a ceiling on it. The score still says how well the site is built; the gate says whether any of it is reachable today. A perfectly marked-up site that returns 403 to `ClaudeBot` is not badly built, it is invisible, and those are different problems with different fixes. Blending them into one capped number made a well-built gated site indistinguishable from a badly-built one — see the changelog for 4.0.

**The automated coverage.** The percentage of applicable weight the scanner verified on its own. Some tests — log retention, a monthly prompt panel, whether schema is generated from fields or hardcoded in a template — cannot be seen from outside. Those are scored from attestation and flagged. A score of 82 with 91% automated coverage is a much stronger claim than 82 with 64%.

Report all three together. A bare number invites arguments we don't need to have.

### Grade bands

| Score | Grade | What it means |
|---|---|---|
| 90–100 | Exemplary | Machine-legible by design. Remaining work is optional or emerging. |
| 75–89 | Strong | Solid fundamentals with specific, addressable gaps. |
| 60–74 | Adequate | Findable and readable, but leaving significant value unclaimed. |
| 40–59 | Weak | Partially legible. Real content is not reaching answer engines. |
| 0–39 | Not ready | Structural problems. Start with retrievability before anything else. |
| Any | **Gated** | A separate fact, not a grade. AIR-1.1 or AIR-1.2 failed: AI crawlers cannot reach the site, so none of the score above is reaching one today. Before 4.0 this capped the total at 40, which blended two facts into a number comparable to neither. |

`grade` is always the band the real, uncapped `total` falls in — a gate failure never changes it. `gated` is the separate boolean the report checks to decide whether to lead with the access problem, so it can say "Strong at 81.2, and gated — none of that is reaching an assistant today" rather than either hiding the score behind the word Gated or letting it stand alone with no warning attached.

## The scoring model

### Bands

Every check scores on a 0–4 band. The check's contribution is `weight × (band ÷ 4)`.

Unless a test says otherwise, use the **default coverage rubric** against the sampled page set:

| Band | Condition |
|---|---|
| 4 | Present and correct on ≥95% of applicable sampled pages |
| 3 | Present and correct on 75–94% |
| 2 | Present on 40–74%, or present everywhere but with correctness errors |
| 1 | Present on 1–39% |
| 0 | Absent, or present but broken everywhere |

Coverage percentages are weighted, not raw page counts. Each money page (see Sampling, step 4) counts as **2** in both the numerator and the denominator; every other sampled page counts as 1. Ten money pages in a 250-URL sample therefore score as a 260-page sample. A failure on `/apply` costs twice what the same failure on a 2019 press release costs, which is the reason the client names those URLs in the first place.

Tests that are inherently site-wide and binary (a file exists, a header is set) define their own bands in the test entry.

### Points

| # | Dimension | Points | Tests |
|---|---|---|---|
| 1 | Primary Indicators | 30 | 13 |
| 2 | Rendering and extraction | 15 | 6 |
| 3 | Structured data and the entity graph | 15 | 8 |
| 4 | Crawl, index and licensing hygiene | 12 | 10 |
| 5 | Authorship, provenance, freshness | 12 | 6 |
| 6 | Agent interfaces | 8 | 4 |
| 7 | Alternate representations | 8 | 4 |
| | **Total** | **100** | **51** |

Beside the hundred, and counted separately because they are not part of it: 7 leading indicators and 9 best practices.
Dimensions 1, 2 and 3 (First contact, Rendering and extraction, and Structured data and the entity graph) carry 57 of the 100 points — just over half. First contact is heaviest by construction: it is the highest-signal test pulled from six other dimensions, not a new editorial judgment about what matters most.

Dimension 8 carries only 7 of 100 points despite being the most commercially valuable work we do. Weight here reflects what a scanner can observe from outside a site, not what the work is worth. Track measurement maturity in the engagement, not in this number.

Dimension 9 carries none of the 100 — by design, not omission. It holds every `leading`-tier test (see *Versioning and change control*, 4.0), grouped on its own rather than scattered through the dimensions that would otherwise host them, so a site's score can never be quietly diluted by a practice almost nobody has built yet. They are still run and still reported.

### Applicability and N/A

Each test declares an applicability rule:

- **`always`** — scored on every site.
- **`conditional`** — scored only if a precondition holds. State the precondition in the check.
- **`vertical`** — scored only for a declared client vertical. Fifteen are published; `other` is one of them and carries no expected type set, so a client who tells us their industry is not on the list is not marked against a standard that does not exist for them.

A conditional test states its precondition in one of two forms, and never in both:

- **A named site fact** — a single named predicate over collected evidence, such as `sample_contains_tables` or `site_has_search`. Predicates are code, not expressions in a config file, so each one is unit-tested and each one names itself in the N/A reason an auditor reads.
- **A dependency on another check** — written as `requires: AIR-4.1 ≥ 2`. A dependency is satisfied when the named check scores **band 2 or higher**. Below band 2 the foundation is broken, and measuring what sits on top of it produces a finding nobody can act on: telling a client their malformed `llms.txt` is well cached is noise. A dependency that is itself N/A cascades N/A to everything that depends on it.

Four checks carry a dependency: AIR-4.2 → AIR-4.1, AIR-7.2 → AIR-7.1, AIR-7.4 → AIR-1.10 **or** AIR-7.1, BP-6 → BP-5.

#### Facts the scanner cannot observe

Eight checks are conditional on a business fact no crawler can determine: AIR-4.1 and AIR-4.3 (licensing and access posture), AIR-5.3 (YMYL content), AIR-5.5 (research and policy publishing), AIR-5.6 (original asset production), LI-7 (permission for full-text inclusion), and BP-1 with BP-2 (named critical flows). AIR-3.3 needs a declared vertical.

These come from an engagement profile supplied per client — never from inference. The scanner does not sniff a site for medical keywords and decide it is YMYL. A guess inside a number we ask a client to defend line by line is worse than saying we have not checked.

But not checking is not the same as not scoring, and from 1.0 to 9.x the Index treated them as the same thing.

#### Every test is scored (10.0)

**On any run a person is shown, every test outside dimension 9 comes back with a band.** The denominator is the full 100 points, on every site, every time.

Applicability resolution still runs, and still decides exactly what it always decided: whether this test has anything to find on this site. What changed in 10.0 is only what happens next. A verdict that used to end the test's participation in the run now resolves to a band, in one of three ways:

| What resolution found | Band | What the report says |
|---|---|---|
| The site has none of the thing the test is about — a precondition that evaluated false, a vertical the site is not in, a dependency whose parent had nothing to find | **4** | *There is none of this on the site, so there is nothing to fix.* |
| The answer is real but invisible from outside — an undeclared profile fact, a manual test with no attestation, or evidence we read and could not settle | **2**, held | *We cannot see this from outside your site. Confirm it and this scores in full.* |
| The test this one depends on is not passing | **0** | *The test this one builds on is not in place yet.* |

The middle row is a **ceiling**, not an earned band, and uses exactly the mechanism `unattested_ceiling` already used: the band is held at 2, `band_capped_at` and `cap_reason` record why, and `earnable_points` falls accordingly. A held test can never read as a pass, and confirming it can always move the number. Points above a ceiling are also excluded from the backlog: a test sitting at its ceiling has nothing on the site to fix, and a list headed *what to fix* must not open with work nobody can collect.

Every one of these carries its `unscorable_reason` in the result file and an evidence row naming what was looked for, what came back, and what it scored — the same audit trail as any other test. An auditor can still see exactly what we chose not to measure; it is now a fact about a scored row rather than a missing one.

**Why.** A test with no band was invisible to the arithmetic *and* to the reader. A self-serve visitor declares nothing, so more than half the Index left the run — a third of it purely for want of an engagement profile — and the score that came back was a percentage of a denominator that silently changed from site to site. Two sites could not be compared, and a client reading their own report found rows that gave them neither a result nor anything to do. Scoring every test costs the Index nothing it was actually measuring and gives back the property its name claims: a hundred points, all of them accounted for.

**What this is not.** It is not scoring an N/A as zero — penalizing a museum for having no `Physician` markup is the mistake this Index was built to avoid, and the first row above is what prevents it. Nor is it paying a site for our own blind spots: anything we could not see is held at 2, never granted, which is the same dishonesty pointing the other way.

**Two runs are exempt.** A **quick scan** looks at thirteen tests and says so; scoring the other fifty-five would hand it points for work it never did and move its ceiling from site to site. The **census** exists to report what share of sites do a given thing, so a test it could not judge has to stay out of that tally rather than be counted as a site with nothing to fix. Both run with `score_everything` off, and neither is a report a client is handed.

```
score = (Σ earned_points ÷ Σ applicable_weight) × 100
```

### Resolving external identifiers

AIR-1.12 asks whether a site's `sameAs` links point at real authority records. Classifying a URL as a tier-1 identifier needs no network call; confirming it still resolves does, and those hosts belong to somebody else — ORCID, Wikidata, a registry, a university directory.

Requesting them tells a third party which site is being audited. That is a real disclosure, and it is why resolution is a deliberate choice rather than a default in the reference scanner's CLI. A scan that does not resolve them holds AIR-1.12 at band 3: the links are classified, not confirmed.

Two rules apply wherever it is turned on:

- **A third party's availability is never a fact about the client's site.** A timeout, a DNS failure or a rate limit against ORCID records that one link as unconfirmed. It must not fail the scan, and it must not score the client down.
- **Say that it happened.** The evidence records which external hosts were contacted, so a client can see what their audit disclosed and to whom.

### Rounding, and the full accounting

**Every points figure shown to a reader is a whole number. Every points figure used in the arithmetic is not.**

A test contributes `weight × band ÷ 4`. No weight in the Index divides by four, so a weight-2 test at band 3 contributes exactly 1.5 points and a weight-5 test at band 2 contributes exactly 2.5. Reports print those as 2 and 3.

This is deliberate. The Index exists to be defended out loud, line by line, in a room with the people who have to do the work. Nobody argues about three quarters of a point, and a column of figures like `2.25` reads as a spreadsheet artifact rather than a judgment somebody made. Whole numbers are the readable choice.

They are also, in aggregate, wrong — and saying so is not optional:

- **A run's printed rows do not sum to its printed score.** On a typical run roughly a third of the rows round up or down. Adding the printed column can land ten points or more away from the printed total.
- **The score is never computed from the rounded rows.** It is computed from the exact contributions, summed at full precision, divided by the total weight, and rounded exactly once, at the end.

Both numbers are correct. They answer different questions, and only one of them is the score.

**Every result therefore carries a full accounting**, reachable from the foot of the report and from the stored result file: every test, its weight, its band, the arithmetic `weight × band ÷ 4` written out, its exact contribution, the figure the report printed instead, and the final division. Nothing on that page is rounded except where it says so. It recomputes the score from weight and band rather than reprinting the stored figure, so a disagreement between the accounting and the report would be visible rather than silent.

An implementer of this specification must do the same. Print whole numbers if you like, but publish the arithmetic, and never compute a score from figures you have already rounded.

*The one real fix — re-weighting every test to a multiple of four, so a 400-point Index lands every band on a whole number — is not made here. It would be a second non-comparable version boundary bought for a cosmetic gain, and the accounting already makes the arithmetic checkable.*

### Gates

Two checks are gates: **AIR-1.1** and **AIR-1.2**.

If either scores 0 or 1, set `gated: true` and record which one in `gate_failures`. Nothing else changes: `total` is computed exactly as in *Scoring* above, from every applicable check including the gate checks themselves, and `grade` is the band that real `total` falls in.

Before 4.0, a gate failure capped the final score at 40 (`total = min(ungated_score, 40)`) and forced `grade` to `Gated`. That produced one number standing for two different facts — how well the site is built, and whether any of it is reachable today — and made them impossible to tell apart: a gated site with excellent structure scored identically to a gated site with none. Reporting `gated` and `total` as two separate facts says both; capping said neither.

### Verification mode

Every test declares how its band was determined:

- **`auto`** — the scanner determined the band from evidence it collected.
- **`auto-partial`** — the scanner produced a strong signal, but a human must confirm. Heuristics that can be fooled live here.
- **`manual`** — not observable from outside. **No scored test uses this mode as of 11.0.** Every point in the hundred is something the scanner can see for itself; what cannot be observed is reported as a best practice instead. The mode remains defined because the best practices use it.

Automated coverage counts `auto-partial` at half weight, because a strong signal awaiting human confirmation is genuinely half-verified:

```
automated_coverage = (Σ auto_weight + 0.5 × Σ auto_partial_weight) ÷ Σ applicable_weight
```

The half is not arbitrary. Counting `auto-partial` at zero puts a ceiling of 0.83 on every possible run, because the always-applicable `auto-partial` checks are 15.75 points that never leave the denominator. Counting it in full reports 0.96 for a site nobody has confirmed anything about, and erases the distinction the mode exists to draw.

Never let a `manual` check silently default to 4. An unanswered attestation scores `null` and drops out of the denominator, and the result records it as unverified.

#### Attestation against an `auto-partial` check

Four `auto-partial` checks have a top band that asks something no scanner can see. This document says so in each case, and each has a different observable ceiling because their band tables differ:

| Test | Unattested ceiling | What the attestation confirms |
|---|---|---|
| LI-3 | 3 | That IndexNow submissions fire on publish, not on a schedule |
| AIR-1.10 | 3 | That `llms.txt` is generated from the content model rather than hand-maintained |
| AIR-3.1 | 3 | That JSON-LD comes from mapped fields, confirmed with the build team |
| BP-9 | 3 | That sitemaps are submitted and alerts reach a named person |

Each of these four now has a band 3 that describes exactly what a scanner can see, so the attestation lifts 3 to 4 rather than jumping a gap. A test whose band table skips a value cannot express *mostly right*, and *mostly right* is the state most real sites are in.

An attestation may **raise the ceiling** on these four. It never sets a band and never lowers one: the scanner scores what it observed, and a supplied attestation permits that observed band to rise as far as the evidence supports. The result records which points came from a claim rather than a probe, and the report prints them as attested.

Without an attestation the test scores at its observable ceiling and the report says what would be needed to go higher. A client who did the work can prove it; a client who says nothing earns nothing.

Every other `auto-partial` check is fully observable. It carries that mode because the heuristic can be fooled — AIR-4.3 compares a signed request against an unsigned one, AIR-6.3 calls a live tool — not because it needs a client to vouch for it.

#### Tests that need a previous run

Three checks are defined against change over time: AIR-4.8 (is `lastmod` a real signal or deploy noise), AIR-2.4 (do anchor IDs survive a deploy), and AIR-5.4 (do dates move only when content moves). A fourth, selector churn on critical-flow elements, moved to BP-2 in 11.0 — measured every run as a Best Practice rather than banded and capped, since it no longer carries points to cap.

On a first run there is nothing to compare against. Score the observable half, **cap the band at 3**, set `confidence: medium`, and record `stability_unconfirmed` in the evidence. AIR-5.4 already words its band 3 this way; the other three follow the same rule.

Supplying a previous run directory as a baseline lifts the cap and permits band 4.

Every ceiling — this one, the attestation ceiling above, and a test's own declaration
that it could not see further — records the reason it applied. A client told a number
was held below the evidence is owed the sentence explaining what would lift it. This makes the re-audit a first-class input to the score rather than a report-writing exercise, and it gives the client a visible, earnable reason the number rises next quarter.

## How the audit runs

### Sampling

Most tests score against a sampled page set, not the whole site. Build the sample like this:

1. **Always include:** the homepage, `/robots.txt`, `/sitemap.xml` (and every child sitemap), `/llms.txt`, `/.well-known/` probes, and every URL in the site's primary navigation.
2. **Stratify by content type.** Use the sitemap index, URL patterns, or a declared content-type map. Take up to 10 URLs per type.
3. **Add depth.** Include at least 5 URLs three or more levels deep, to catch templates the homepage never exercises.
4. **Add the money pages.** The client names up to 10 URLs that matter most — programs, services, donate, apply, contact. These are weighted double in coverage math.
5. **Cap the crawl.** Default 250 URLs. Configurable. Record the actual count.

Minimum viable sample is 25 URLs. Mark the run `low_confidence` only when the sample is under 25 **and** more URLs were discovered than were sampled. A twelve-page nonprofit site that was crawled in full is a complete measurement, not a weak one.

#### The representation sub-sample

Four checks want an extra request against every sampled URL: AIR-7.1 (`{path}.md`), AIR-7.3 (`Accept: text/markdown`), AIR-7.4 (cache behavior), and AIR-4.4 (`X-Robots-Tag` on non-HTML). At the default 250-URL cap and one request per second, that is over 800 requests and roughly fifteen minutes of politeness delay stacked on top of the base crawl, against infrastructure that often belongs to a hospital.

Those four tests run against a **representation sub-sample**: 25 URLs by default, drawn from the main sample, always including every money page. The size is configurable and the selected URLs are recorded in the run metadata, so the sub-sample is reproducible across audits. Twenty-five pages resolves a rubric whose finest distinction is 75% against 95%.

### Fetching

Fetch every sampled URL twice:

- **Raw** — plain HTTP GET, no JavaScript, with a declared user-agent.
- **Rendered** — a headless browser with JavaScript enabled.

The difference between those two responses is the entire evidence base for AIR-2.1 and AIR-2.2, which together carry 9 points. Do not skip the second fetch.

Separately, probe a small set of URLs with each AI crawler user-agent to feed AIR-1.2. There is no other way to measure what the edge does: `robots.txt` states an intention, a CDN rule states a fact, and the two disagree often enough that a readiness score which skipped the question would be worth less than the awkwardness of asking it.

Say plainly what that means, because it is the part an implementer will be asked about. **The wire-level `User-Agent` on those requests carries another company's product token** — `GPTBot`, `ClaudeBot`, and the six others in the table below. Three constraints make that measurement rather than impersonation, and an implementation that drops any of them is doing something else:

1. **The requests still identify you.** Send your own `From` header and an `X-Audit-Operator` header naming your crawler on every probe. An operator inspecting the whole request, rather than the user-agent alone, can see who it was.
2. **Probe only URLs you were already allowed to fetch.** The probe runs over the sample, and the sample is filtered by `robots.txt` first. A probe must never reach, under somebody else's name, a page you declined to reach under your own.
3. **Publish the behavior before you do it**, at the URL in your own user-agent string, alongside how to opt out. Disclosure volunteered reads differently from disclosure discovered.

#### The agent classes

The Index groups the ten agents three ways. AIR-1.1 evaluates all ten against `robots.txt`. AIR-1.2 probes only the eight that actually issue requests.

| Class | Agents | AIR-1.1 | AIR-1.2 probe |
|---|---|---|---|
| Answer-serving | `OAI-SearchBot`, `ChatGPT-User`, `PerplexityBot`, `Claude-User` | yes | yes |
| Training and corpus | `GPTBot`, `ClaudeBot`, `CCBot`, `Bytespider` | yes | yes |
| Directive-only | `Google-Extended`, `Applebot-Extended` | yes | **no** |

`Google-Extended` and `Applebot-Extended` are not crawlers. They are opt-out tokens that Google and Apple honor in `robots.txt`; no request ever arrives carrying either string. Blocking them is a rights decision with no fetch behavior attached, which is why they count in AIR-1.1 and cannot be probed in AIR-1.2. A user-agent matrix that lists them is describing a crawler that does not exist.

Where a band anchor says **major agents**, it means the four answer-serving agents plus `GPTBot` and `ClaudeBot`.

### Politeness

Respect `Crawl-delay`. Default to 1 request per second with 4 concurrent workers. Back off on 429 and 503, honoring `Retry-After` where it is given, and abandon a host that keeps refusing rather than retrying into a wall. Cache every response to disk so re-scoring never re-crawls.

### Consent

Two different questions, and conflating them is how a scanner ends up doing something indefensible.

**Whose site may you scan?** A client audit runs with the client's agreement, recorded before the crawl. A survey of a public population — see *Populations*, below — runs without asking each site first, on the same footing as any other published measurement of the public web. What is not acceptable in either case is scanning under a consent you did not obtain. Record which basis a run used, in the run itself.

**What may you fetch once you are there?** `robots.txt`, and specifically the group addressed to your own crawler. Name your crawler in your user-agent string, publish a page at the URL that string points to, and stop when a site names you and disallows you.

A scanner may reasonably decide that a blanket `User-agent: *` disallow does not address it — a `*` rule is aimed at the anonymous scrape, and a named crawler with a published policy, a rate limit and a working opt-out is not that. If you make that choice, **publish it on your crawler's own page** and honor `Crawl-delay` under `*` regardless, because a rate limit is not an exclusion. What is not defensible is making the choice quietly.

### Populations

A survey scores a defined population and publishes the aggregate. Two rules keep it a measurement rather than an accusation.

**Cite the population.** Record the source list, its publication date and its URL. "How did you pick these" is the first question anyone asks, and the answer has to be a citation rather than a judgment call. Use sources you may lawfully redistribute, and claim no ownership over somebody else's list.

**Publish the version and the date.** Every result records which version of this Index it was scored against and the date that version took effect. Without both, an old number silently starts meaning something new the next time the Index changes.

### Reproducibility

Store the raw evidence — headers, HTML, extracted JSON-LD, screenshots — in a run directory keyed by timestamp. A score you cannot reproduce next quarter is an opinion, not a measurement.

## Result schema

The scanner writes one `result.json` per run. Everything else (Markdown report, HTML dashboard, remediation backlog) renders from this file.

```json
{
  "schema_version": "1.0",
  "index_version": "1.0",
  "site": {
    "name": "Example University",
    "base_url": "https://example.edu",
    "vertical": "higher-ed",
    "cms": "drupal-10"
  },
  "run": {
    "id": "2026-09-01T20-14-33Z",
    "started_at": "2026-09-01T20:14:33Z",
    "finished_at": "2026-09-01T20:41:02Z",
    "scanner_version": "0.1.0",
    "urls_sampled": 187,
    "representation_sample": 25,
    "low_confidence": false,
    "baseline_run_id": null,
    "extraction_method": "readability-density",
    "extraction_version": "1"
  },
  "score": {
    "total": 61.4,
    "gated": false,
    "gate_failures": [],
    "grade": "Adequate",
    "applicable_weight": 94.5,
    "earned_points": 58.0,
    "leading_adopted": 2,
    "leading_total": 6,
    "earnable_points": 88.0,
    "automated_coverage": 0.87
  },
  "dimensions": [
    {
      "id": 3,
      "name": "Rendering and extraction",
      "weight": 20.0,
      "applicable_weight": 20.0,
      "earned_points": 9.5,
      "score_pct": 47.5
    }
  ],
  "checks": [
    {
      "id": "AIR-2.1",
      "title": "Primary content is present without JavaScript",
      "dimension": 3,
      "weight": 6.0,
      "applicability": "always",
      "applicable": true,
      "na_reason": null,
      "verification": "auto",
      "band": 1,
      "earned_points": 1.5,
      "confidence": "high",
      "summary": "Raw HTML contains 22% of rendered text on average across 187 URLs.",
      "evidence": [
        {
          "url": "https://example.edu/admissions",
          "raw_text_chars": 812,
          "rendered_text_chars": 9440,
          "coverage": 0.086
        }
      ],
      "remediation": {
        "severity": "critical",
        "effort": "large",
        "summary": "Server-render the decoupled front end.",
        "detail": "Program and admissions templates ship an empty shell..."
      }
    }
  ],
  "remediation_backlog": [
    {
      "rank": 1,
      "check": "AIR-2.1",
      "points_available": 4.5,
      "severity": "critical",
      "effort": "large"
    }
  ]
}
```

#### Schema notes

- `site.cms` is best-effort, from the `generator` meta tag and response headers. It is nullable and nothing scores against it.
- A dimension with no applicable checks reports `score_pct: null`, not `0`. Zero means measured and failed.
- **Rounding happens once, at the end.** All arithmetic runs unrounded; the values written to this file are rounded to one decimal for display. The per-check `earned_points` therefore will not sum exactly to `score.earned_points`, and neither is wrong. Do not "fix" this by rounding intermediates — that reintroduces the drift the single rounding step exists to prevent.
- `extraction_method` and `extraction_version` record how main-content text was isolated. When extraction improves, old scores stay interpretable because the method that produced them is on the record.

### Remediation ranking

Sort the backlog by **points available** (`weight − earned_points`), then by effort ascending, then by check ID ascending. The last key is not cosmetic: two runs of the same evidence must produce the same backlog in the same order, or the re-audit comparison is unreadable.

A check appears in the backlog only if it is applicable, has a band, and has points available above zero. N/A checks, unattested `manual` checks, and checks already at band 4 are not remediation items.

Use three severities — `critical`, `major`, `minor` — and three effort sizes — `small`, `medium`, `large`. Effort is a property of the check, defined in this document. Severity is derived: gate failures and any check scoring **0 or 1** with weight ≥3.0 are `critical`; band ≤2 with weight ≥1.5 is `major`; everything else is `minor`.

Band 1 on a heavy check is a critical finding. A site whose program pages ship 22% of their content scores AIR-2.1 at band 1, not 0, and calling that `major` alongside a missing caption on a table understates it to the person deciding what to fund.

The client-facing output is not the score. It is the ordered list of what to fix and what each fix is worth.

---

# The tests

Each entry gives: points, applicability, verification mode, effort, what the test measures, how to detect it, what evidence to keep, and the band rule when it differs from the default coverage rubric.

## #1 · Primary Indicators · 13 tests · 30 pts

**Points 30**

### AIR-1.1 — No blanket robots.txt blocks on AI crawler user-agents


- **Points:** 5
- **Gate**
- **Applicability:** always
- **Verification:** auto
- **Effort:** small
- **Measures:** Whether `robots.txt` disallows the crawlers that feed answer engines.
- **Detect:** Parse `robots.txt`. Evaluate each of `GPTBot`, `OAI-SearchBot`, `ChatGPT-User`, `ClaudeBot`, `Claude-User`, `PerplexityBot`, `Google-Extended`, `Applebot-Extended`, `CCBot`, `Bytespider` against the rule set. Resolve `User-agent: *` fallbacks correctly, including the longest-match group rule.
- **Evidence:** Raw `robots.txt`, the parsed rule tree, and a per-agent allow/deny verdict for `/` and for three sampled content URLs.
- **Bands:** anchored on *which* agents are blocked, not how many.
  - **4** — every listed agent may fetch content paths.
  - **3** — one training or directive-only agent blocked; every answer-serving agent allowed.
  - **2** — two or more training or directive-only agents blocked, up to and including all of them; every answer-serving agent allowed.
  - **1** — exactly one answer-serving agent blocked.
  - **0** — `Disallow: /` for `*`, or two or more answer-serving agents blocked.
- **Why the gate turns on the answer-serving class:** the Index measures whether machines can find, read and *use* a site. An answer engine blocked at `robots.txt` cannot answer a question about the client, which is the failure the gate exists to catch. A training block is a rights posture, and this document declines to judge rights postures — AIR-1.4 scores any combination of `yes` and `no` at 4 for exactly that reason. A university that blocks `GPTBot` and welcomes `ChatGPT-User` is reachable, and gating it would be measuring our approval rather than machine reach.
- **Consequence:** blocking every training crawler costs 2 of the 5 points and does not gate. Blocking a single answer-serving agent gates.
- **Absent robots.txt:** a 404 scores 4. Nothing is disallowed, so nothing is blocked. This is correct and frequently queried; a missing file is not a failure of this test.
- **Note:** A deliberate, documented block is a business decision, not a defect. Record it, score it honestly, and say in the report that the client chose it. The number should reflect machine reach, not our approval.
- **Fix:** Remove inherited blocks. Keep only directives the client can explain.

### AIR-1.2 — CDN and WAF do not silently reject AI crawlers


- **Points:** 5
- **Gate**
- **Applicability:** always
- **Verification:** auto
- **Effort:** medium
- **Measures:** Whether the edge returns real content to AI user-agents, regardless of what `robots.txt` permits.
- **Detect:** Request 10 sampled URLs with each of the eight fetching AI user-agents — the answer-serving and training classes defined under Fetching. `Google-Extended` and `Applebot-Extended` are excluded; they issue no requests. Record status code, response size, and whether the body matches the baseline fetch. Flag 403, 429, 503, interstitial challenge markup, and bodies more than 40% smaller than baseline.
- **Evidence:** A user-agent × URL matrix of status codes, sizes, and challenge detection.
- **Bands:** 4 — all agents get 200 with full content. 3 — one agent rate-limited but eventually served. 2 — one agent blocked or challenged. 1 — several blocked. 0 — most blocked.
- **Fix:** Audit Cloudflare bot settings, Pantheon AGCDN rules, Fastly VCL, and origin WAF. Verify by re-running this probe, not by reading configuration.

### AIR-1.3 — No interactive challenges on read-only paths


- **Points:** 1
- **Applicability:** always
- **Verification:** auto
- **Effort:** small
- **Measures:** Whether anonymous GETs on browse, search, listing, and detail pages return content without a challenge.
- **Detect:** Look for Turnstile, reCAPTCHA, hCaptcha, and JS-challenge signatures in raw responses on non-form URLs. Detect meta-refresh interstitials and challenge cookies.
- **Evidence:** Per-URL challenge detection with the matching signature.
- **Fix:** Scope challenges to POST endpoints and authenticated routes.

### AIR-1.4 — Content Signals declared in robots.txt


- **Points:** 3
- **Applicability:** always
- **Verification:** auto
- **Effort:** small
- **Measures:** Whether the site states a machine-readable position on `search`, `ai-input`, and `ai-train` using Cloudflare's Content Signals Policy.
- **Detect:** Parse `Content-Signal:` lines. Validate syntax (`key=yes|no`, comma-delimited) and check for contradiction with `Disallow` rules.
- **Bands:** 4 — all three signals declared, syntactically valid, consistent with `robots.txt`. 3 — declared but one signal omitted. 2 — declared with a syntax error or an internal contradiction. 1 — a partial or malformed attempt. 0 — absent.
- **Note:** Any combination of `yes`/`no` scores 4. We are measuring that a deliberate position exists, not which position it is.
- **Fix:** Run the policy conversation with the client, then add the line.

### AIR-1.5 — Human-readable AI policy comment in robots.txt


- **Points:** 1
- **Applicability:** always
- **Verification:** auto-partial
- **Effort:** small
- **Measures:** Whether a person reading `robots.txt` finds a plain-language statement of the organization's position.
- **Detect:** Extract `#` comment blocks. Require ≥120 characters of prose and at least two policy terms (`AI`, `training`, `crawl`, `license`, `permission`).
- **Bands:** 4 — substantive block present and consistent with the machine signals. 2 — a comment exists but is thin or boilerplate. 0 — none.
- **Fix:** Write four sentences above the signals. Journalists and counsel read this file.

### AIR-1.6 — No stray noarchive or nosnippet directives


- **Points:** 1
- **Applicability:** always
- **Verification:** auto
- **Effort:** small
- **Measures:** Whether the site suppresses the excerpting that answer engines depend on.
- **Detect:** Parse meta robots tags and `X-Robots-Tag` headers across the sample for `noarchive`, `nosnippet`, and `max-snippet:0`.
- **Bands:** 4 — none present. 2 — present on a minority of pages. 0 — present site-wide.
- **Fix:** Remove them. They are almost always left over from a legacy SEO module.

### AIR-1.7 — Every page emits a valid, self-consistent canonical


- **Points:** 2
- **Applicability:** always
- **Verification:** auto
- **Effort:** small
- **Measures:** Presence and correctness of `rel=canonical`.
- **Detect:** For each sampled page, extract the canonical and resolve it. Flag: missing, multiple, cross-host, pointing at a redirect, pointing at a non-200, or pointing at the homepage from a deep page.
- **Fix:** Emit a self-referencing canonical by default and correct the exceptions.

### AIR-1.8 — Semantic HTML structure


- **Points:** 2
- **Applicability:** always
- **Verification:** auto
- **Effort:** medium
- **Measures:** Whether document structure carries meaning or only styling.
- **Detect:** Per page, require exactly one `<main>`, at least one `<article>` or `<section>` wrapping the primary content, and a `<nav>` for primary navigation. Compute a div-to-semantic-element ratio in the content region and flag pages above 12:1.
- **Fix:** Replace structural divs with `article`, `section`, `nav`, `aside`, `figure`, and `dl`.

### AIR-1.9 — One h1 and a non-skipping heading hierarchy


- **Points:** 2
- **Applicability:** always
- **Verification:** auto
- **Effort:** small
- **Measures:** Whether headings segment the page into topics a model can follow.
- **Detect:** Count `h1` elements. Walk the heading tree and flag skipped levels, empty headings, and headings used purely for styling (no following content).
- **Bands:** Default coverage rubric, where a page passes only if it has exactly one non-empty `h1` and no skipped levels.
- **Fix:** Move the site name out of `h1`. Fix skips in the templates, not the content.

### AIR-1.10 — /llms.txt exists and is generated from the content model


- **Points:** 2
- **Applicability:** always
- **Verification:** auto-partial
- **Effort:** medium
- **Measures:** Presence, validity, and freshness of `/llms.txt`.
- **Detect:** Fetch `/llms.txt`. Validate against the llmstxt.org structure: an `# H1` title, an optional blockquote summary, and `##` sections of Markdown links. Resolve a sample of the linked URLs. Compare listed URLs against the sitemap to estimate coverage and staleness. Detect the Drupal `llms_txt` module by response headers or known output patterns.
- **Bands:** 4 — valid, links resolve, coverage tracks the sitemap, generation confirmed. 3 — valid and current but hand-maintained. 2 — present with broken links or clear staleness. 1 — present but malformed. 0 — absent.
- **Fix:** On Drupal, use the `llms_txt` module with token-driven menus. A hand-written file is stale within a quarter.

### AIR-1.11 — A single linked entity graph with @id references


- **Points:** 2
- **Applicability:** always
- **Verification:** auto
- **Effort:** medium
- **Measures:** Whether schema forms one coherent graph or dozens of disconnected assertions.
- **Detect:** Extract all JSON-LD across the sample. Build a graph of nodes and `@id` references. Measure: what fraction of **entity** nodes carry an `@id` — value objects such as `ListItem`, `PostalAddress`, `ContactPoint`, `EntryPoint` and `BreadcrumbList` are excluded, because they have no identity and nothing should ever reference them; whether `Organization` and `WebSite` have stable canonical `@id` URIs reused site-wide; whether `WebPage` nodes reference `WebSite` and `WebSite` references `Organization`; and how many duplicate definitions of the same entity exist.
- **Evidence:** The resolved graph, orphan node count, duplicate entity count.
- **Bands:** 4 — canonical `@id`s reused, chain intact, no duplicates. 3 — chain intact with minor duplication. 2 — `@id`s present but inconsistently reused. 1 — `@id`s rare. 0 — isolated blobs with no references.
- **Fix:** Define canonical `@id` URIs once and reference them everywhere.

### AIR-1.12 — sameAs links to external authority records


- **Points:** 2
- **Applicability:** always
- **Verification:** auto
- **Effort:** small
- **Measures:** Entity disambiguation — whether a model can be certain which organization this is.
- **Detect:** Extract `sameAs` from `Organization` and `Person` nodes. Classify each URL by authority: tier 1 (Wikidata, ROR, NPI, IRS EIN, ORCID), tier 2 (Wikipedia, Crunchbase, LinkedIn, Candid), tier 3 (social profiles). Resolve each URL and confirm it returns 200 and references the organization back where possible.
- **Resolution is opt-in.** Those hosts are not the client's, they are not on the allowlist, and requesting them tells a third party which site is being audited. Classification from the markup alone separates a controlled identifier from a Facebook page; without resolution the test stops at band 3 and says why.
- **Bands:** 4 — at least one tier-1 identifier plus two others, all resolving. 3 — tier-1 present, some links stale. 2 — tier-2 and tier-3 only. 1 — social profiles only. 0 — no `sameAs`.
- **Fix:** Claim the Wikidata item. For universities, ROR is free and takes a day.

### AIR-1.13 — About, Contact, and organizational detail are complete


- **Points:** 2
- **Applicability:** always
- **Verification:** auto
- **Effort:** medium
- **Measures:** Whether a model can resolve who this organization actually is.
- **Detect:** Locate About and Contact pages. Score presence of: founding date, leadership names, postal address, service area, `contactPoint` entries differentiated by purpose, and public legal identifiers (EIN, ROR, DUNS). Check both the visible page and the `Organization` node.
- **Bands:** Fraction of those seven facts present in both HTML and schema: 4 — ≥6. 3 — 5. 2 — 3–4. 1 — 1–2. 0 — none.
- **Fix:** Thin About pages are a persistent weakness for nonprofits and institutions. Do not leave the key facts in an annual-report PDF.

## #2 · Rendering and Extraction · 6 tests · 15 pts

**Points 15**

### AIR-2.1 — Primary content is present without JavaScript


- **Points:** 6
- **Applicability:** always
- **Verification:** auto
- **Effort:** large
- **Measures:** The gap between the raw HTTP response and the hydrated DOM. This is the single heaviest test in the Index.
- **Detect:** For every sampled URL, extract main-content text from the raw fetch and from the rendered fetch. Compute `coverage = raw_chars ÷ rendered_chars` after normalizing whitespace and stripping nav, header, and footer regions. Average across the sample and report the per-template distribution.
- **Evidence:** Per-URL character counts, coverage ratio, and the first 500 characters of each version for spot-checking.
- **Bands:** 4 — mean coverage ≥0.95. 3 — 0.80–0.94. 2 — 0.50–0.79. 1 — 0.20–0.49. 0 — below 0.20.
- **Edge cases:** Coverage above 1.0 happens when JavaScript *removes* content that the raw response contained. Clamp the per-URL value to 1.0 for banding and keep the unclamped ratio in the evidence. A URL whose rendered text is empty is excluded from the mean and the exclusion is recorded — it is a fetch problem, not a coverage measurement.
- **Note:** Report the worst-performing template by name. "Your program pages ship 8% of their content" lands harder than a mean.
- **Fix:** Server-render or statically render. For decoupled Drupal this is the largest line item in most remediations.

### AIR-2.2 — Progressive disclosure content ships in the initial HTML


- **Points:** 3
- **Applicability:** conditional — the site uses tabs, accordions, load-more, or infinite scroll
- **Verification:** auto
- **Effort:** medium
- **Measures:** Whether tabbed, collapsed, or lazily paged content exists before interaction.
- **Detect:** Identify disclosure widgets by ARIA roles (`tab`, `tabpanel`, `region` with `aria-expanded`), `<details>`, and common class patterns. For each, check whether the panel body is populated in the raw HTML. Detect infinite scroll by watching for XHR content growth on scroll in the rendered fetch, and check for a paginated fallback.
- **Bands:** 4 — every panel populated in raw HTML; infinite scroll has crawlable pagination. 3 — one pattern incomplete. 2 — panels populated on some templates only. 1 — most content loads on interaction. 0 — all disclosure content is fetched on demand.
- **Fix:** Render all panels and hide with CSS. Give infinite scroll a paginated fallback.

### AIR-2.3 — Real data tables with header cells


- **Points:** 2
- **Applicability:** conditional — the sample contains tables
- **Verification:** auto
- **Effort:** medium
- **Measures:** Whether tabular facts — tuition, hours, comparisons, fees — are machine-parseable.
- **Detect:** For each `<table>`, require `<th>` with `scope`, a `<caption>`, and `<thead>`/`<tbody>`. Separately flag layout tables (no `th`, presentational attributes, single row or column).
- **Bands:** 4 — every data table complete, no layout tables. 3 — minor omissions such as missing captions. 2 — header cells present without scope. 1 — tables present with no header cells. 0 — layout tables carrying data.
- **Fix:** Fix the tables that carry facts first. Convert layout tables to CSS.

### AIR-2.4 — Stable, human-readable anchor IDs on section headings


- **Points:** 2
- **Applicability:** always
- **Verification:** auto-partial
- **Effort:** medium
- **Measures:** Fragment-level citability — whether a model can deep-link a passage rather than a page.
- **Detect:** For headings below `h1`, check for an `id`. Classify each as slug-like (derived from heading text) or unstable (`section-3`, `block-a7f2c1`, framework hashes). Compare IDs across two runs to detect churn.
- **Bands:** 4 — ≥95% of subheadings carry slug-like IDs that are stable across runs. 3 — 75–94%. 2 — IDs present but a majority are generated or unstable. 1 — sparse. 0 — none.
- **First run:** cap at 3 without a baseline run. Stability cannot be asserted from one observation. See *Tests that need a previous run*.
- **Fix:** Derive IDs from heading text, not render order.

### AIR-2.5 — Alt text on content images, empty alt on decorative


- **Points:** 1
- **Applicability:** always
- **Verification:** auto
- **Effort:** medium
- **Measures:** Whether images have a textual representation.
- **Detect:** For every `<img>`, record presence of `alt`, its length, and whether the image sits inside a content region or a decorative one. Flag missing `alt`, filename-as-alt, and long alt on obviously decorative images.
- **Fix:** Audit existing content, then make the field required in the media entity so the problem stops recurring.

### AIR-2.6 — Key facts exist as HTML text, not only in images or PDFs


- **Points:** 1
- **Applicability:** always
- **Verification:** auto-partial
- **Effort:** large
- **Measures:** Whether high-value facts are trapped in a scan, an infographic, or a PDF.
- **Detect:** Identify pages whose visible text is thin relative to their images or that link to a PDF as the primary content. Extract PDF text and look for numbers, dates, and prices that appear nowhere in the site's HTML. Requires review before scoring.
- **Bands:** 4 — no high-value facts found only in binaries. 2 — some found only in PDFs. 0 — core facts (tuition, hours, deadlines, contact) exist only in binaries.
- **Fix:** Publish an HTML equivalent. Keep the PDF as a download, not as the source of truth.

## #3 · Structured Data and the Entity Graph · 8 tests · 15 pts

**Points 15**

### AIR-3.1 — JSON-LD is generated from mapped fields, not hardcoded


- **Points:** 2
- **Applicability:** always
- **Verification:** auto-partial
- **Effort:** medium
- **Measures:** Whether schema follows the content or rots on the next content type change.
- **Detect:** Heuristics only. Compare JSON-LD values against visible field values on the same page — a mapped implementation matches. Flag identical literal values repeated across pages of the same type where the visible content differs. Detect Schema.org Metatag output patterns on Drupal. Confirm with the build team.
- **Bands:** 4 — values track content across every sampled page of a type, and the build team confirms the mapping. 3 — values track content on every sampled page, generation unconfirmed. 2 — mixed, with some literals. 0 — clear hardcoding, or values contradict the visible page.
- **Fix:** Map Drupal fields to schema properties through configuration so editors maintain it.

### AIR-3.2 — BreadcrumbList on every page below the homepage


- **Points:** 2
- **Applicability:** always
- **Verification:** auto
- **Effort:** small
- **Measures:** Whether a page declares where it sits in the information architecture.
- **Detect:** Validate `BreadcrumbList` structure and confirm the trail matches the URL path or menu position. Flag trails that are a single item or that contradict the path.
- **Fix:** Generate from the menu, not by hand.

### AIR-3.3 — Vertical-specific schema types


- **Points:** 4
- **Applicability:** vertical — scored against the client's declared vertical only
- **Verification:** auto
- **Effort:** large
- **Measures:** Whether the site models the things it actually contains, using the right types.
- **Detect:** From the declared vertical, load the expected type set and check coverage against matching content types:
  - **Ecommerce:** `Product`, `Offer`, `AggregateRating`, `BreadcrumbList`
  - **Financial Services:** `FinancialService`, `BankOrCreditUnion`, `LoanOrCredit`, `InvestmentOrDeposit`
  - **Government:** `GovernmentOrganization`, `GovernmentService`, `Legislation`
  - **Healthcare:** `MedicalOrganization`, `Physician`, `MedicalCondition`, `MedicalWebPage`
  - **Higher Education:** `Course`, `EducationalOccupationalProgram`, `CollegeOrUniversity`, `EducationEvent`
  - **Hospitality:** `LodgingBusiness`, `Restaurant`, `Reservation`, `Menu`
  - **Local Services:** `LocalBusiness`, `Service`, `OpeningHoursSpecification`, `GeoCoordinates`
  - **Manufacturing:** `Organization`, `Product`, `ProductModel`, `Brand`
  - **Media:** `NewsArticle`, `VideoObject`, `Person`, `ImageObject`
  - **Nonprofit:** `NGO`, `DonateAction`, `Grant`, `FundingScheme`
  - **Professional Services:** `ProfessionalService`, `Service`, `Person`, `Review`
  - **Real Estate:** `RealEstateListing`, `Residence`, `Place`, `GeoCoordinates`
  - **Research:** `Dataset`, `DataCatalog`, `ScholarlyArticle`
  - **SaaS:** `SoftwareApplication`, `Offer`, `FAQPage`, `Organization`
  - **Other:** no expected set. The test is **not scored** and says so on the report.
- **Bands:** Fraction of expected types present and valid on the content that warrants them: 4 — ≥90%. 3 — 70–89%. 2 — 40–69%. 1 — 1–39%. 0 — none.
- **Note:** `Dataset` and `DataCatalog` are badly underused. Research organizations routinely hold catalogd data with no machine-readable description at all.
- **Fix:** Pick the applicable set deliberately. Do not implement types the content does not support.

### AIR-3.4 — FAQPage and QAPage on genuine question-and-answer content


- **Points:** 2
- **Applicability:** conditional — the site has real Q&A content
- **Verification:** auto
- **Effort:** small
- **Measures:** Correct markup on genuine Q&A — and the absence of fabricated FAQ blocks.
- **Detect:** Validate `FAQPage` and `QAPage` structure. Cross-check that every marked-up question appears in the visible page text. Flag pages where FAQ markup exists with no corresponding visible content.
- **Bands:** 4 — real Q&A marked up, nothing fabricated. 2 — partial coverage. 0 — absent where warranted. **Cap at 1 if fabricated FAQ blocks are detected**, regardless of coverage.
- **Fix:** Mark up what exists. Do not invent questions to earn the markup — it is a spam pattern and it degrades the page for people.

### AIR-3.5 — isAccessibleForFree, about, and mentions with entity references


- **Points:** 1
- **Applicability:** always
- **Verification:** auto
- **Effort:** medium
- **Measures:** Access status and topical entity linkage.
- **Detect:** Check for `isAccessibleForFree` on substantive content. Check `about` and `mentions`, and classify their values as `@id` references or bare strings.
- **Bands:** 4 — access declared and subject entities referenced by identifier. 2 — properties present with string values only. 0 — absent.
- **Fix:** Entity references are how a model decides a page is *about* a subject rather than merely containing the word.

### AIR-3.6 — Product, Offer, and Service with real prices


- **Points:** 1
- **Applicability:** conditional — the site sells, charges, or offers something
- **Verification:** auto
- **Effort:** medium
- **Measures:** Whether commercial facts are marked up and current.
- **Detect:** Validate `Offer` structure. Compare `price`, `priceCurrency`, and `availability` against the visible page. Flag any mismatch and any `priceValidUntil` in the past.
- **Bands:** 4 — accurate and matching the page. 2 — present but drifting from visible values. 0 — absent, or demonstrably stale.
- **Note:** Stale pricing in schema is worse than none. A model will quote it with confidence.
- **Fix:** Source the schema value from the same field the page renders.

### AIR-3.7 — SearchAction on the WebSite node


- **Points:** 1
- **Applicability:** conditional — the site has search
- **Verification:** auto
- **Effort:** small
- **Measures:** Whether site search is exposed as a callable pattern.
- **Detect:** Extract `potentialAction` / `SearchAction` from the `WebSite` node. Execute the URL template with a test term and confirm real results come back.
- **Bands:** 4 — declared and the template returns results. 2 — declared but the template fails. 0 — absent.
- **Fix:** Small piece of the agent-interface story, and it costs almost nothing.

### AIR-3.8 — Schema validation of the structured data on the page

- **Points:** 2
- **Applicability:** always
- **Verification:** auto
- **Effort:** small
- **Measures:** Whether the structured data on the page is valid: that every block parses, that each node carries the properties its declared type needs to be usable, and that values are the shape their property requires.
- **Detect:** Parse every JSON-LD, microdata and RDFa block on every sampled page. A block that does not parse is a failure of the page it is on. For each node, check the properties a consumer needs for the type it declares, and check that dates are ISO 8601, that URLs are absolute, and that a price carries a currency.
- **Bands:** 4 — every block parses, every node carries what its type requires, every value is well formed. 3 — everything parses and one property or value is wrong somewhere. 2 — everything parses; required properties are missing on some nodes. 0 — a block on a sampled page does not parse.
- **Fix:** Schema rots silently. Nothing on the page changes when a date stops being a date or a required property is dropped in a template edit, and the markup keeps being served to machines that cannot use it. Validate what you publish rather than trusting it was right when it was written.
- **Changed in 11.0:** This used to ask whether schema validation ran in CI, which is not observable from outside and so awarded the same mark to every site. It measures the markup itself now.

## #4 · Crawl, Index and Licensing Hygiene · 10 tests · 12 pts

**Points 12**

### AIR-4.1 — RSL licensing document published and referenced


- **Points:** 1
- **Applicability:** conditional — the client holds licensable archives (journals, museums, associations, research bodies) or has opted into a licensing posture
- **Verification:** auto
- **Effort:** medium
- **Measures:** Whether machine-readable licensing terms exist under RSL 1.0 and are discoverable.
- **Detect:** Fetch `/license.xml` and any path referenced from `robots.txt`. Validate against the RSL 1.0 schema. Check for an HTTP `Link:` header on content responses.
- **Bands:** 4 — valid document, referenced from both `robots.txt` and headers. 3 — valid, referenced from one. 2 — present but invalid. 1 — referenced but missing. 0 — absent.
- **Fix:** Decide terms with the client first. The XML is the easy half.

### AIR-4.2 — RSL terms propagated to feeds and schema


- **Points:** 1
- **Applicability:** conditional — requires AIR-4.1 to pass
- **Verification:** auto
- **Effort:** small
- **Measures:** Whether licensing travels with the content through every channel, not just the well-known file.
- **Detect:** Check RSS and Atom feeds for licensing elements. Check sampled JSON-LD for a `license` property on licensable types.
- **Fix:** Add the license reference to feed generation and to the schema field mapping.

### AIR-4.3 — Web Bot Auth verification configured


- **Points:** 1
- **Applicability:** conditional — the client wants selective agent access rather than open or closed
- **Verification:** auto-partial
- **Effort:** large
- **Measures:** Whether the edge distinguishes cryptographically verified agents from unsigned scrapers.
- **Detect:** Probe `/.well-known/http-message-signatures-directory` if the site operates its own agents. Otherwise send a signed and an unsigned request and compare treatment.
- **Bands:** 4 — verified agents pass, unsigned agents claiming an AI UA are challenged. 2 — signature headers accepted but not acted on. 0 — no differentiation.
- **Fix:** Configure edge bot rules on `Signature-Agent` rather than user-agent strings.

### AIR-4.4 — Correct X-Robots-Tag on non-HTML resources


- **Points:** 1
- **Applicability:** conditional — the site serves PDFs, documents, media, or JSON endpoints
- **Verification:** auto
- **Effort:** small
- **Measures:** Whether indexing directives reach resources that cannot carry a meta robots tag.
- **Detect:** HEAD every sampled PDF, document, media file, and API endpoint. Record `X-Robots-Tag`. Flag `noindex` on substantive content and absent directives on endpoints that should be excluded.
- **Bands:** 4 — every non-HTML type carries a deliberate, correct value. 2 — mixed or partial. 0 — absent everywhere, or substantive PDFs marked `noindex`.
- **Fix:** Set headers by path pattern at the CDN or web server.

### AIR-4.5 — Missing pages return 404 or 410


- **Points:** 1
- **Applicability:** always
- **Verification:** auto
- **Effort:** small
- **Measures:** Soft 404s — "not found" content served with a 200 status.
- **Detect:** Request 10 deliberately invalid URLs under real path prefixes. Flag 200 responses. Separately, flag sampled pages whose body matches a not-found template or falls below a minimum content threshold while returning 200.
- **Evidence:** Probe URLs with status codes and body fingerprints.
- **Fix:** Return real status codes. Use 410 for content that is permanently gone.

### AIR-4.6 — Redirects resolve in a single hop


- **Points:** 1
- **Applicability:** always
- **Verification:** auto
- **Effort:** small
- **Measures:** Redirect chain depth and loops.
- **Detect:** Follow redirects without auto-resolution and count hops per sampled URL and per internal link target.
- **Bands:** 4 — no chain exceeds one hop. 3 — a few two-hop chains. 2 — chains of three or more. 1 — widespread chains, meaning both more than 10% of the sample and at least three URLs. 0 — loops present.
- **Fix:** Flatten the redirect table so every source points at its final destination.

### AIR-4.7 — Faceted, calendar, and parameterized URLs are controlled


- **Points:** 2
- **Applicability:** conditional — the site has search, filtering, or a calendar
- **Verification:** auto-partial
- **Effort:** medium
- **Measures:** Whether the crawlable URL space is finite and predictable.
- **Detect:** Crawl listing pages two levels deep with parameters followed. Measure URL growth rate and count distinct parameter combinations. Check whether those URLs are disallowed, `noindex`, or canonicalized to the unfiltered listing.
- **Bands:** 4 — parameter space bounded by robots rules, `noindex`, or canonicals. 3 — mostly controlled with gaps. 2 — partially controlled. 1 — minimal control. 0 — unbounded expansion observed.
- **Fix:** Combine disallow patterns, `noindex` on parameterized views, and canonicals back to the base listing.

### AIR-4.8 — Sitemap lastmod reflects real content changes


- **Points:** 2
- **Applicability:** always
- **Verification:** auto-partial
- **Effort:** medium
- **Measures:** Whether `lastmod` is a freshness signal or deploy noise.
- **Detect:** Parse all `lastmod` values. Flag clustering — more than 60% of URLs sharing a single date, or timestamps identical to the minute across content types. Cross-check against on-page `dateModified` where present. Re-run across two audits to confirm.
- **Bands:** 4 — values are distributed and match on-page dates. 3 — mostly accurate with some clustering. 2 — heavy clustering. 1 — all identical. 0 — absent.
- **First run:** cap at 3 without a baseline run. See *Tests that need a previous run*.
- **Fix:** Wire `lastmod` to the node's changed timestamp, not to cron.

### AIR-4.9 — Sitemap index split by type, with media sitemaps


- **Points:** 1
- **Applicability:** always
- **Verification:** auto
- **Effort:** small
- **Measures:** Whether sitemap structure makes crawl coverage diagnosable.
- **Detect:** Fetch `/sitemap.xml`. Confirm it is an index with child sitemaps. Check for image and video sitemaps where the site has substantial media. Validate against the sitemap schema and the 50k URL / 50MB limits.
- **Fix:** Generate per-content-type children and submit the index to both search consoles.

### AIR-4.10 — Paginated series are coherently signaled


- **Points:** 1
- **Applicability:** conditional — the site has paginated listings or multi-page articles
- **Verification:** auto
- **Effort:** small
- **Measures:** Whether page 2 and beyond are reachable and correctly related to page 1.
- **Detect:** Identify paginated series by URL pattern. Check for `rel=next`/`rel=prev`, a view-all page carrying the canonical, and whether deep pages are reachable by crawl.
- **Fix:** Prefer a view-all canonical where page weight allows.

## #5 · Authorship, Provenance, Freshness · 6 tests · 12 pts

**Points 12**

### AIR-5.1 — Author entity pages with credentials


- **Points:** 2
- **Applicability:** always
- **Verification:** auto
- **Effort:** medium
- **Measures:** Whether authors resolve to real, credentialed entities.
- **Detect:** Follow `author` references to their target pages. Require `Person` markup with a canonical `@id`, plus at least two of `jobTitle`, `affiliation`, `knowsAbout`, `sameAs`. Resolve `sameAs` (ORCID, LinkedIn, institutional profile).
- **Bands:** 4 — every recurring author has a resolvable entity page with verified identifiers. 3 — most do. 2 — author pages exist without markup or credentials. 1 — author names only. 0 — no author entities.
- **Fix:** A byline string asserts nothing. A resolvable entity with credentials does.

### AIR-5.2 — Bylines linked to author entities


- **Points:** 2
- **Applicability:** always
- **Verification:** auto
- **Effort:** small
- **Measures:** Whether substantive pages are signed, and whether the schema `author` is an `@id` reference rather than a repeated string.
- **Detect:** For each substantive page, check for a visible byline and an `author` property. Classify the value as `@id` reference or literal.
- **Bands:** 4 — visible byline plus `@id` reference on ≥95%. 3 — 75–94%. 2 — `author` present as a bare string. 1 — sparse. 0 — unsigned.
- **Fix:** Where the author is a department, name the department and model it as an entity.

### AIR-5.3 — reviewedBy on YMYL content


- **Points:** 2
- **Applicability:** conditional — the site publishes medical, legal, or financial content
- **Verification:** auto
- **Effort:** medium
- **Measures:** Whether expert review reaches the markup.
- **Detect:** On YMYL pages, check for `reviewedBy` referencing a credentialed `Person`, plus `lastReviewed` where the type supports it, plus a visible reviewer statement.
- **Fix:** Most healthcare and legal clients already run review workflows. The reviewer's identity simply never reaches the page.

### AIR-5.4 — datePublished and dateModified are accurate


- **Points:** 2
- **Applicability:** always
- **Verification:** auto-partial
- **Effort:** medium
- **Measures:** Whether the freshness signal is honest.
- **Detect:** Extract both dates from schema and visible markup. Flag: missing dates, `dateModified` before `datePublished`, future dates, and clustering that suggests a deploy stamp. Compare against sitemap `lastmod` (AIR-4.8) and against the previous audit run to confirm dates only move when content moves.
- **Bands:** 4 — dates present, internally consistent, and stable across a deploy with no content change. 3 — present and consistent, stability unconfirmed. 2 — present but clustered or contradicting the sitemap. 1 — one date only. 0 — absent.
- **Fix:** Wire both to the node's own timestamps. A deploy must not move them.

### AIR-5.5 — citation markup on primary-source references


- **Points:** 2
- **Applicability:** conditional — the site publishes research or policy content
- **Verification:** auto
- **Effort:** medium
- **Measures:** Whether sourcing is machine-readable.
- **Detect:** Check for `citation` on research and policy pages. Resolve DOIs where present. Compare against the count of outbound links to primary sources in the body.
- **Fix:** Credibility that a parser cannot see does not count.

### AIR-5.6 — C2PA Content Credentials on original assets


- **Points:** 2
- **Applicability:** conditional — the client produces original photography, primary research, or documents where authenticity matters
- **Verification:** auto
- **Effort:** large
- **Measures:** Whether provenance survives into the delivered file.
- **Detect:** Download sampled original images and PDFs and read C2PA manifests. Confirm the manifest survives the image derivative pipeline — this is where it usually breaks.
- **Bands:** 4 — valid manifests on originals and derivatives. 2 — present on originals, stripped by the pipeline. 0 — absent.
- **Fix:** Scope this deliberately. It needs signing infrastructure.

## #6 · Agent Interfaces · 4 tests · 8 pts

**Points 8**

### AIR-6.1 — OpenAPI specification for public APIs


- **Points:** 2
- **Applicability:** conditional — the site has a public or semi-public API
- **Verification:** auto
- **Effort:** small
- **Detect:** Look for the spec at conventional paths and from `/.well-known/`. Validate it, and confirm it is versioned and linked from the site.

### AIR-6.2 — /.well-known/ discovery entries for agent capabilities


- **Points:** 2
- **Applicability:** conditional — the site exposes any agent capability
- **Verification:** auto
- **Effort:** small
- **Detect:** Enumerate `/.well-known/` for MCP, NLWeb, OpenAPI, and `http-message-signatures-directory`. Validate each returns parseable metadata.

### AIR-6.3 — Form fields carry label, name, and autocomplete


- **Points:** 2
- **Applicability:** conditional — the site has public forms
- **Verification:** auto
- **Effort:** medium
- **Measures:** Whether an agent can tell what each field is for.
- **Detect:** For every public form field: a `<label for>` or equivalent accessible name, a meaningful `name` attribute (reject `field_1`, `input3`), and a valid `autocomplete` token where one applies (`given-name`, `email`, `postal-code`, `tel`). A token *applies* where the field type implies a purpose — `email`, `tel`, `url`, `password`. A search box or a free-text field has no sensible token and is not marked down for lacking one.
- **Bands:** Fraction of fields passing all three: 4 — ≥95%. 3 — 75–94%. 2 — 40–74%. 1 — 1–39%. 0 — none.
- **Fix:** Same change satisfies WCAG. Easy joint justification.

### AIR-6.4 — No captchas on browse, search, or filter interactions


- **Points:** 2
- **Applicability:** always
- **Verification:** auto
- **Effort:** small
- **Detect:** Submit search and filter forms and check for challenge responses. Complements AIR-1.3, from the application side rather than the edge.

## #7 · Alternate Representations · 4 tests · 8 pts

**Points 8**

### AIR-7.1 — Markdown companion for every canonical page


- **Points:** 2
- **Applicability:** always
- **Verification:** auto
- **Effort:** medium
- **Measures:** Whether a clean `.md` representation exists alongside each page.
- **Detect:** For each sampled content URL, request `{path}.md`. Require a 200 with `text/markdown` or `text/plain`, content that materially matches the page body, and no navigation chrome.
- **Fix:** Generate from the same source as the HTML. Strip nav, promos, and cookie banners.

### AIR-7.2 — link rel=alternate advertises the Markdown version


- **Points:** 2
- **Applicability:** conditional — requires LI-7 to pass
- **Verification:** auto
- **Effort:** small
- **Measures:** Discoverability of the alternate representation.
- **Detect:** Parse `<link rel="alternate" type="text/markdown">` from the head and resolve the href.
- **Fix:** Add it to the head template. A file nobody can find is not much use.

### AIR-7.3 — Accept: text/markdown content negotiation


- **Points:** 2
- **Applicability:** always
- **Verification:** auto
- **Effort:** medium
- **Measures:** Whether the canonical URL serves Markdown on request, with correct cache headers.
- **Detect:** Request sampled URLs with `Accept: text/markdown`. Confirm the response type and that `Vary: Accept` is set. Re-request through the CDN to confirm the variants are cached separately.
- **Bands:** 4 — negotiation works and `Vary: Accept` is correct. 2 — negotiation works but `Vary` is missing, which risks serving Markdown to browsers. 0 — not supported.
- **Fix:** Getting `Vary` wrong is worse than not doing this at all. Verify at the edge.

### AIR-7.4 — Alternate representations are edge-cached


- **Points:** 2
- **Applicability:** conditional — requires AIR-1.10 or LI-7
- **Verification:** auto
- **Effort:** small
- **Measures:** Whether generated representations hit the CDN rather than origin.
- **Detect:** Request `/llms.txt` and sampled `.md` URLs twice. Read cache-status headers and compare response times.
- **Fix:** Set TTLs and wire invalidation to the same cache tags as the source content.

## Leading Indicators · 7 items · not scored

Practices with essentially no adoption in 2026. Each is measured and reported
like any other test and contributes nothing to the hundred, because scoring a
practice nobody has adopted subtracts the same points from every site and moves
nobody relative to anybody.

They are numbered within this section rather than inside a dimension. A global
number would imply a place in the score, and they have none.

### LI-1 — HowTo on procedural content


- **Points:** none
- **Not scored**
- **Applicability:** conditional — the site has procedural content
- **Verification:** auto
- **Effort:** small
- **Measures:** Whether step sequences are explicit rather than inferred from prose.
- **Detect:** Identify procedural pages by ordered lists with step-like language in headings. Check for `HowTo` with ordered `HowToStep`.
- **Fix:** Application walkthroughs, enrollment steps, and permit processes are exactly what people ask assistants about.

### LI-2 — speakable on summaries and ledes


- **Points:** none
- **Not scored**
- **Applicability:** always
- **Verification:** auto
- **Effort:** small
- **Measures:** Whether the page flags its own best short answer.
- **Detect:** Extract `speakable` with `cssSelector` or `xpath`. Resolve the selector against the DOM and confirm it matches 1–3 elements containing 40–400 characters. Flag selectors that match the whole body or nothing.
- **Bands:** 4 — resolves to a genuine summary passage. 2 — present but resolves too broadly. 0 — absent.
- **Fix:** Point it at the lede, not the article.

### LI-3 — IndexNow fires on publish and update


- **Points:** none
- **Not scored**
- **Applicability:** always
- **Verification:** auto-partial
- **Effort:** small
- **Measures:** Whether the site pushes change notifications instead of waiting to be crawled.
- **Detect:** Look for an IndexNow key file at the site root (a 32–128 character hex filename returning its own key). Confirm the key resolves. Whether submissions actually fire on publish requires attestation.
- **Bands:** 4 — key file present, valid, and submission confirmed. 3 — key file present and valid, submission unconfirmed. 2 — a key file is referenced but does not resolve or does not contain its own key. 0 — absent or not discoverable.
- **Limitation:** the key file's name *is* the key, and the key is not published. A scanner can only find it where `robots.txt` or a conventional path advertises it, so a zero here means *not discoverable* rather than *absent*. Score it honestly and say which.
- **Fix:** Hook submission into the publish workflow rather than a schedule.

### LI-4 — Editorial policy page referenced via publishingPrinciples


- **Points:** none
- **Not scored**
- **Applicability:** always
- **Verification:** auto
- **Effort:** small
- **Measures:** Whether the site documents how content is produced, reviewed, corrected, and funded.
- **Detect:** Look for `publishingPrinciples` on the `Organization` node and resolve it. For publishers, also check `correctionsPolicy` and `diversityPolicy`.
- **Fix:** Write the page that describes the process the client already follows.

### LI-5 — NLWeb endpoint exposed as an MCP server with an ask method


- **Points:** none
- **Not scored**
- **Applicability:** always
- **Verification:** auto
- **Effort:** large
- **Detect:** Probe known NLWeb paths and `/.well-known/` entries. If found, send a representative question and check for a schema.org-formatted JSON response.
- **Bands:** 4 — endpoint answers correctly and is grounded in site content. 2 — endpoint responds but answers poorly. 0 — absent.

### LI-6 — Domain MCP server over real content APIs


- **Points:** none
- **Not scored**
- **Applicability:** always
- **Verification:** auto-partial
- **Effort:** large
- **Measures:** Whether the client's structured data — course catalog, provider directory, grant database, program finder — is callable as documented tools.
- **Detect:** Discover from `/.well-known/`. Enumerate tools, check each has a description and typed parameters, and call one read-only tool against live data.
- **Bands:** 4 — documented, typed tools returning live data, with auth and rate limiting. 3 — tools present but thinly documented. 2 — a generic query endpoint rather than narrow tools. 0 — none.
- **Note:** The most defensible item in the whole Index. A well-described tool with three parameters beats a generic endpoint.

### LI-7 — /llms-full.txt where full-text inclusion is appropriate


- **Points:** none
- **Not scored**
- **Applicability:** conditional — the client permits full-text inclusion; N/A if Content Signals declare `ai-train=no`
- **Verification:** auto
- **Effort:** small
- **Measures:** Presence of the expanded variant.
- **Detect:** Fetch `/llms-full.txt`. Confirm it inlines content rather than linking. Cross-check against AIR-1.4 for contradiction.
- **Bands:** 4 — present and consistent with the licensing posture. 2 — present but contradicts declared signals. 0 — absent where appropriate.
- **Fix:** Gate this on the licensing decision, not on convenience.

## Best Practices · 9 items · not scored

Real work, worth doing, and invisible to a scan. Whether a channel group exists
in somebody's analytics, whether access logs are retained, whether a build
validates schema before it merges — none of it can be observed from outside the
organization.

They were scored until 11.0, and scoring them was a mistake worth naming.
Across 69 scans of 39 sites every one of them returned the identical result,
because an unverifiable check awarded the same mark to everybody. Sixteen of the
hundred were therefore the same number on every report: they inflated every
score equally, moved nobody relative to anybody, and made the total look more
precise than it was. They are reported as guidance now and left out of the
arithmetic.

Numbered within this section, for the same reason the leading indicators are.

### BP-1 — Critical flows complete without JS-only interactions


- **Points:** none
- **Not scored**
- **Applicability:** conditional — the client names critical flows
- **Verification:** auto-partial
- **Effort:** large
- **Measures:** Whether register, donate, apply, or find-a-clinician can be completed through standard form submissions and real URL transitions.
- **Detect:** For each named flow, walk it with JavaScript disabled. Flag `div`-based controls with click handlers, wizards with no addressable step URLs, and modal-only paths.
- **Bands:** 4 — every flow completes with real URLs at each step. 2 — flows start but cannot complete. 0 — flows are JS-only.
- **Fix:** Run this as a joint engineering and UX audit.

### BP-2 — Stable selectors on critical-flow elements


- **Points:** none
- **Not scored**
- **Applicability:** conditional — requires BP-1 flows to be named
- **Verification:** auto
- **Effort:** medium
- **Detect:** Collect selectors for the key elements in each flow. Classify as stable (`id`, `data-*`, semantic) or fragile (hashed utility classes, build-generated IDs). Diff across two runs to detect churn.
- **Bands:** 4 — stable and unchanged across runs. 2 — mixed. 0 — hashed classes only.
- **First run:** cap at 2 without a baseline run, since band 4 is defined entirely by cross-run stability. See *Tests that need a previous run*.
- **Fix:** Treat these selectors as a contract that survives redeploys.

### BP-3 — GA4 channel group for AI assistant referrers


- **Points:** none
- **Not scored**
- **Applicability:** always
- **Verification:** manual
- **Effort:** small
- **Measures:** Whether the client can see traffic from `chatgpt.com`, `perplexity.ai`, `claude.ai`, `copilot.microsoft.com`, and `gemini.google.com`. By default it lands in Direct and disappears.
- **Detect:** Analytics presence is detectable from the page. The channel group is not — attest it, ideally with a screenshot.
- **Bands:** 4 — group live with a saved report. 2 — analytics present, no grouping. 0 — no analytics.

### BP-4 — Server-side tagging captures stripped referrers


- **Points:** none
- **Not scored**
- **Applicability:** always
- **Verification:** manual
- **Effort:** medium
- **Detect:** A server-side endpoint may be inferable from network requests. Reconciliation against client-side numbers must be attested.

### BP-5 — Access log retention with a queryable store


- **Points:** none
- **Not scored**
- **Applicability:** always
- **Verification:** manual
- **Effort:** medium
- **Measures:** Whether crawler logs survive long enough to show a before-and-after window. Many hosts discard them in days.
- **Bands:** 4 — ≥12 months retained and queryable by user-agent, path, status, date. 2 — retained but not queryable. 0 — discarded.

### BP-6 — Scheduled AI crawler activity report


- **Points:** none
- **Not scored**
- **Applicability:** conditional — requires BP-5
- **Verification:** manual
- **Effort:** medium
- **Measures:** Whether crawler behavior is monitored continuously rather than audited once. A crawler that stops appearing usually means somebody reintroduced a block.

### BP-7 — Fixed prompt panel run monthly across models


- **Points:** none
- **Not scored**
- **Applicability:** always
- **Verification:** manual
- **Effort:** medium
- **Measures:** Whether answer-share is tracked on a stable panel of brand, category, and comparison prompts.
- **Bands:** 4 — fixed panel, baseline stored, monthly runs comparable. 2 — ad-hoc tests. 0 — none.
- **Note:** The panel has to stay fixed. Tuning prompts between runs destroys the trend line.

### BP-8 — Extraction-fidelity baseline captured


- **Points:** none
- **Not scored**
- **Applicability:** always
- **Verification:** manual
- **Effort:** medium
- **Measures:** Whether a before-and-after fidelity score exists — feed the page to a model, ask it the client's key questions, score the accuracy.
- **Note:** The scanner can generate this itself in a later version. Until then, attest it. The before-and-after delta is the best single-slide value demonstration we have.

### BP-9 — Search Console and Bing Webmaster Tools verified


- **Points:** none
- **Not scored**
- **Applicability:** always
- **Verification:** auto-partial
- **Effort:** small
- **Detect:** Look for `google-site-verification` and `msvalidate.01` meta tags, or DNS TXT verification records. Whether anyone reads the alerts must be attested.
- **Bands:** 4 — both verified, sitemaps submitted, alerts routed to a person. 3 — both verified; submission and alert routing unconfirmed. 2 — one verified. 0 — neither.
- **Note:** Bing matters more than its search share suggests. Its index feeds several AI products, and clients frequently have Search Console with no Bing property at all.

## What 100 means

A site scoring 100 has, at minimum:

- Every AI crawler reaching real content at the edge and in `robots.txt`, with a deliberate, documented licensing position.
- A clean, finite URL space with honest freshness signals.
- Every substantive fact present in the server response, in semantic markup, with stable anchors a model can cite.
- Generated `llms.txt` and Markdown companions, cached at the edge.
- One linked entity graph, anchored to external authorities, validated in CI so it cannot rot.
- Credentialed authors, honest timestamps, machine-readable sourcing.
- Documented agent endpoints and flows an agent can actually complete.
- Instrumentation that proves all of it, running on a schedule.

Almost no site will score 100, and that is the point. The Index is a target to move toward, not a bar to clear. A client who moves from 41 to 72 in one engagement has a number they can take to their board — which is worth considerably more to them than a perfect score they will never reach.

## What the Index deliberately does not measure

We keep this list public because leaving it out would be dishonest.

- **Whether the content is any good.** Structure, not substance.
- **Whether the models actually cite the site.** That is answer-share monitoring, a different instrument on a different cadence.
- **Traditional SEO performance.** Rankings, backlinks, and Core Web Vitals are real, and they are not this.
- **Accessibility conformance.** Several tests overlap with WCAG, happily. This is not a WCAG audit and must not be sold as one.
- **Security posture.** We probe bot handling at the edge. We do not test the site's defenses.

## Versioning and change control

The Index version is independent of the scanner version. Both appear in every result file.

Bump the **minor** version for new tests, clarified detection, or corrected band anchors. Bump the **major** version for any weight change, gate change, or scoring-model change — anything that makes two scores non-comparable.

Never silently re-weight. A client's trend line is the most valuable thing the Index produces, and one quiet weight change destroys it. When a major bump lands, re-score the last run under both versions and show the difference.

### Changelog

#### 11.0 — 2026-09-24

**Scores computed under 10.x cannot be compared with scores computed under
11.0.** Seventeen tests left the hundred, one dimension was removed, every
remaining test was renumbered, and eleven points moved. Every number moved.

1. **Seventeen tests were returning the same answer for every site.** Measured,
   not guessed: across every production scan on record, seventeen tests in
   dimensions 2 through 8 had never once produced a different result. A test
   that cannot separate two sites is not measuring anything; it is a constant
   added to every score, and sixteen points of the hundred were being handed out
   for nothing. They are all still run and still reported — as **Leading
   Indicators** (LI-1 through LI-7) and **Best Practices** (BP-1 through BP-9),
   beside the hundred rather than inside it.

2. **The hundred is 51 tests, and every one of them is observable.** The Index
   no longer has a single scored test in `manual` mode. What the scanner cannot
   see for itself is reported as a best practice; what carries points is
   something the crawl establishes. `auto-partial` remains for the fifteen
   points held below band 4 pending a confirmation only the client can give.

3. **Dimension 8, Measurement and instrumentation, was removed.** Whether a
   channel group exists in somebody's analytics, whether their access logs are
   queryable, whether a prompt panel runs monthly — none of it is visible from
   outside, and all of it was scoring the same everywhere. It is now the Best
   Practices family, which is where advice that a scan cannot verify belongs.

4. **Seven dimensions, reweighted.** Primary Indicators 28 → **30**; Rendering
   and Extraction **15**; Structured Data and the Entity Graph 14 → **15**;
   Crawl, Index and Licensing Hygiene 14 → **12**; Authorship, Provenance,
   Freshness 8 → **12**; Agent Interfaces 7 → **8**; Alternate Representations 7
   → **8**. The eleven points freed by the removals went to tests that already
   separate sites rather than into a new one.

5. **AIR-3.8 measures the markup instead of asking about the build.** It used to
   ask whether schema validation ran in CI — unobservable, and so worth two free
   points to everybody. It now parses every structured-data block on every
   sampled page, holds each node against the properties its declared type needs,
   and checks that dates are ISO 8601, that URLs are absolute and that a price
   carries a currency. Worth 2 points now rather than 1, and AIR-3.3 gave up
   the point to pay for it.

6. **Everything is renumbered, with no gaps.** Removing a dimension and moving
   seventeen tests out of the hundred left holes across every family, and an id
   that means nothing is worse than an id that changed. There is no mapping
   table on purpose: the numbers are not comparable across this bump, and a
   table would invite somebody to pretend they are.

#### 10.4 — 2026-09-24

**Two dimensions renamed. No weight, band, gate or scoring-model change, and no
score computed under 10.3 moves.**

Dimension 1 is now **Primary Indicators**, and dimension 9 is now **Leading
Indicators**.

1. **"First contact" named a moment; the dimension measures a condition.** It
   holds the thirteen highest-signal tests, pulled from six other dimensions
   because each can be answered from a single page — which is also what makes
   them the free scan's thirteen. "Primary" says what they are to the reader:
   the first things to get right, and the ones that decide whether anything
   below them can be read at all. The free scan takes the same name, so the
   page somebody lands on and the dimension it reports are no longer two
   products with one set of tests.

2. **"Unscored leading indicators" said the same thing twice.** Leading
   indicators are unscored by definition — that is what the tier means, and the
   dimension's own section has said so since 10.0. The word was carried in the
   name to prevent a misreading that the zero in the weight column already
   prevents.

**Scores are comparable across this bump.** Both dimensions keep their ids,
their weights, their tests and their band anchors. A result scored under 10.3
and the same evidence scored under 10.4 produce the same number; only the two
labels differ. Earlier changelog entries keep the names in use at the time,
because they are a record of what happened rather than a description of the
Index today.

#### 10.3 — 2026-09-21

**Ten more verticals for AIR-3.3, and an explicit "other". No weight, band or
scoring-model change, and no score computed under 10.2 moves.**

AIR-3.3 is scored against the client's declared vertical, and five were
published: higher education, healthcare, nonprofit, government and research.
Every other kind of organization had nothing to declare, so the test left their
denominator as *not declared* — correct, but it meant an ecommerce site or a
manufacturer could never be measured on the one test about modelling what they
actually contain.

Nine are added — ecommerce, financial services, hospitality, local services,
manufacturing, media, professional services, real estate and SaaS — each with
its expected type set above. The five that existed are **byte-for-byte
unchanged**, in both their stored values and their type sets, so a site scored
under 10.2 scores identically under 10.3.

`other` is the tenth, and it is deliberately different: it carries no expected
type set. A client who tells us their industry is not on the list has told us
something true, and marking them against types the Index never defined for them
would be scoring them for our own gap. AIR-3.3 is therefore **not scored** for
`other` and the report says why, with an invitation to tell us the industry so
a set can be added.

A client with no vertical at all is unchanged in every respect: still *not
declared*, still out of the denominator, still the same score.

#### 10.2 — 2026-09-17

**Corrected detection for AIR-1.2. No weight, band or scoring-model change, but a
site's AIR-1.2 result can move, because the old one was measuring the wrong thing.**

The edge probe sent each agent's bare token as its entire User-Agent header — a
literal `User-Agent: GPTBot`. No crawler sends that. Edge protection keys on the
real signature: a browser-shaped prefix, `compatible; <Token>/<version>`, and the
vendor's contact URL. So the probe walked straight past the rules it exists to
find, and a site whose CDN was turning GPTBot away scored as though it were
letting GPTBot in.

Measured on readinessindex.io on 2026-09-17, against an edge running Cloudflare's
AI-crawler blocking: probing as `GPTBot` and `CCBot` was served 200, while the
real strings for the same two agents were served 403 by the same edge on the same
request path. The Index reported 2 of 8 agents blocked. Four were.

AIR-1.2 is a gate. Under-detecting here does not shade a number slightly — it
hands a site the one verdict the Index exists to raise, *AI assistants can reach
you*, when they cannot.

1. **The probe now sends each agent's real User-Agent string**, listed in
   `PROBE_USER_AGENTS` with its provenance: OpenAI and Perplexity publish theirs
   verbatim, and Anthropic, Common Crawl and ByteDance do not, so those three are
   the canonical observed forms and are marked as such.
2. **Findings still name the bare token.** It is what a reader recognises and what
   they would write into robots.txt; the full string belongs on the wire, not in
   the report.
3. **Version drift is expected and harmless.** OpenAI moved GPTBot from 1.2 to 1.4
   without announcement. Edge rules match the shape and the token rather than the
   digits, so a stale version still probes honestly — a bare token does not.

A site that passed AIR-1.2 under 10.1 and fails it under 10.2 was failing all
along. Re-scan before concluding anything from a comparison across this boundary.

#### 10.1 — 2026-09-17

**Clarifications and two scanner switches. No weight, band or scoring change — a
site scored under 10.0 and the same site scored under 10.1 report the same number
from the same evidence.**

1. **Rounding is now stated in the specification**, in *Rounding, and the full
   accounting*. It always happened; it was never written down. Every points figure
   shown to a reader is a whole number, roughly a third of a run's rows round to get
   there, and the score is computed from the exact contributions and rounded exactly
   once at the end — so a report's printed rows do not sum to its printed score.
   Implementers must publish the arithmetic and must never compute a score from
   figures they have already rounded.
2. **Every result carries a full accounting**, reachable from the foot of the report:
   every test, its weight, its band, `weight × band ÷ 4` written out, its exact
   contribution, the figure printed instead, and the final division. It recomputes
   the score from weight and band rather than reprinting the stored figure, so a
   disagreement between the accounting and the report is visible rather than silent.
3. **Resolving external identifiers has a rule of its own**, in *Resolving external
   identifiers*. A third party's availability is never a fact about the client's
   site: a timeout against ORCID records one link unconfirmed, and must neither fail
   the scan nor score the client down. What the audit disclosed, and to whom, is
   recorded in the evidence.

#### 10.0 — 2026-09-17

**Every test is scored.** Nothing outside dimension 9 leaves the denominator any
more, on any run a person is shown. A test the engine cannot band from evidence
resolves to a band by rule instead of dropping out — full marks where the site has
none of the thing, a held band 2 where the answer is invisible from outside, zero
where the test it builds on is not passing. The full model is in *Applicability and
N/A* above.

**Scores are not comparable across this boundary, in either direction.** Both halves
of the fraction moved: the denominator is now the full 100 on every site, and tests
that contributed nothing now contribute a band. The reference fixtures show the size
of it — the clean site reads 84 of 100 where it read 93 of 69, and the weak site 50
where it read 38 of 55. A site that scored under 9.x and re-scores under 10.0 must
have both numbers shown side by side, per the versioning rule, and neither number is
wrong: they are answers to two different questions.

1. **The rule.** Applicability resolution is unchanged — it still decides whether a
   test has anything to find. What changed is what happens to its verdict:
   `precondition_not_met`, `vertical_mismatch` and `dependency_not_applicable` score
   4; `not_declared_in_engagement_profile`, `no_attestation` and
   `insufficient_evidence` are held at 2; `dependency_not_met` scores 0. Two reasons
   are deliberately not in that table — `check_not_implemented` and `scanner_error`
   both mean *we* failed, not the site, and inventing a band for our own defect would
   hide it.
2. **Held is not earned.** A held test carries `band_capped_at: 2` and the cap reason
   *awaiting confirmation from your team*, exactly as an `unattested_ceiling` always
   has, so `earnable_points` falls and confirming it can still move the score. A
   report that showed 100 points available where only 87 could be reached would be
   selling points nobody can buy.
3. **The backlog now ranks by what is reachable, not by raw weight.**
   `points_available` is `earnable − earned` rather than `weight − earned`. This
   fixes an older bug in passing: a test held by a first-run cap or an attestation
   ceiling used to promise its full remaining weight at the top of the backlog, and
   no amount of work on the site could collect it.
4. **Every scored row carries evidence.** A test banded by this rule records what was
   looked for, what came back, and what it scored — so *we found none of this* is
   checkable in the same place as any other finding, rather than being an assertion.
5. **Two exemptions, both for runs nobody is handed.** The quick scan (thirteen
   tests, a fixed published ceiling) and the census (population statistics, where our
   blind spot must not be reported as an industry result) keep the 9.x behavior via
   `score_everything: false`.

**Why now.** The self-serve full scan made the old model's cost impossible to ignore.
Nobody filling in a URL box declares an engagement profile, so a real report came
back with 52 of its 100 points unassessed, 19 of them for want of business facts
alone — and the rows that carried them told the reader neither a result nor anything
to do. The Index is named for a hundred points. It should score them.

#### 9.0 — 2026-09-16

**A new dimension, not a new score.** Every `leading`-tier test — measured
and reported since 4.0, never scored — moves out of the dimension it
happened to be filed under and into its own: **#9 · Unscored Leading
Indicators**, declared at 0 of the 100 points, last in dimension order.
Nothing about *why* a leading test doesn't count changes; what changes is
that it no longer sits inside a dimension whose own heading implies every
one of its tests carries weight. A reader adding up a dimension's own rows
now gets that dimension's own declared total, every time — before this,
seven of those rows didn't, because they were leading tests hiding in
plain sight in dimensions 3, 4, 5, and 6.

1. **Six tests move, ten more renumber to close the gaps they leave.**
   `checks.yaml` stays grouped and numbered by dimension (6.0's own rule),
   so a check moving dimension is a check getting a new id:
   - AIR-3.5 (HowTo) → AIR-9.1, AIR-3.6 (speakable) → AIR-9.2. The four that
     followed them close the gap: AIR-3.7 → AIR-3.5, AIR-3.8 → AIR-3.6,
     AIR-3.9 → AIR-3.7, AIR-3.10 → AIR-3.8.
   - AIR-4.11 (IndexNow) → AIR-9.3. Last in its dimension; nothing behind it
     to renumber.
   - AIR-5.7 (editorial policy) → AIR-9.4. Also last; same reason.
   - AIR-6.1 (NLWeb endpoint) → AIR-9.5, AIR-6.3 (domain MCP server) →
     AIR-9.6. The seven that followed close the gap: AIR-6.2 → AIR-6.1,
     AIR-6.4 → AIR-6.2, AIR-6.5 → AIR-6.3, AIR-6.6 → AIR-6.4, AIR-6.7 →
     AIR-6.5, AIR-6.8 → AIR-6.6, AIR-6.9 → AIR-6.7.
   - The one dependency naming an id that moved is updated to match: the
     test now numbered AIR-6.1 (formerly AIR-6.2) required AIR-6.1 before
     this change; it requires AIR-9.5 now, the new number for the same
     test it always depended on.
2. **No weight, band, or score changes.** Every test kept its own weight;
   every dimension holding only core tests kept its own declared total,
   because removing a test that never contributed to it changes nothing
   about the sum. A site scored under 8.0 and the same site scored under
   9.0 report identical numbers everywhere except which id and dimension
   two handfuls of findings print under — the same guarantee 6.0 made for
   the same reason.
3. **A related bug, caught while building this: a dimension's own
   applicable/earned points included any leading test inside it.** Scoring
   itself was never affected — `score_run`'s own total already excluded
   leading tests — but `_dimension_results` didn't apply the same
   exclusion, so a dimension holding one could show a share-earned number
   diluted by a test that was never part of its denominator. Dimension 9
   would have inherited this as a dimension that could show nonzero earned
   points against a declared weight of zero — negative points "left,"
   on the one dimension where that number most needs to read as exactly
   zero. Fixed in the same change: leading tests are now excluded from a
   dimension's own applicable and earned totals, everywhere they're
   computed, not only in the overall score.

#### 8.0 — 2026-09-16

**A convention change, not a scoring one, bumped major because it touches
the one canonical dimension label everywhere it's used.** A dimension
heading now reads `#1 · First Contact · 13 tests · 28 pts` rather than
`13 checks · 28% weight` — the dimension's own weight, shown as points,
not re-expressed as a percent of the total (every other number on the
site already reads as points, and the percent was the odd one out). The
site's own word for what it audits is "test" now, not "check" — the nav,
the footer, the scan pages. Nothing is re-scored: a check's id, weight,
and band are exactly what they were; only what the page calls it changed.

#### 7.0 — 2026-09-15

**A major bump on both the letter and the spirit of the rule.** Seventeen
checks change weight, and for the nine that were at zero this is not
accounting — a check worth zero points cannot move a score, and now every
one of them can. Re-score before comparing a 7.0 result to anything earlier.

1. **No core check weighs zero.** The nine checks 5.0 rounded to zero on the
   100-point scale — `AIR-5.5`, `AIR-5.6`, `AIR-6.4`, `AIR-6.5`, `AIR-7.4`,
   `AIR-8.1`, `AIR-8.2`, `AIR-8.6`, `AIR-8.7` — each now carry at least 1
   point, funded by moving weight rather than by widening the scale. A check
   in the Index is a claim that something is worth measuring; a check worth
   nothing was a claim the number itself contradicted.
2. **Three dimensions could not fund their own floor.** *Agent interfaces*
   (9 checks) and *Measurement and instrumentation* (7 checks) each totaled
   fewer points than they have checks — mathematically, some of them had to
   be zero. Both dimensions' declared weight rises to match their own check
   count (5 → 7, and 3 → 7), every check in each now weighing exactly 1.
   *Alternate representations* funds its own zero-weight check from within
   its existing 5, plus 2 more for the same reason (5 → 7), landing at
   1/2/1/2/1 across its five checks rather than the flat 1 the other two
   dimensions reach.
3. **First contact funds the difference: 36 → 28.** The 8 points the other
   three dimensions needed come from here, not from *Structured data* or
   *Crawl, index and licensing hygiene* — First contact is still by far the
   heaviest dimension, and the two checks it gates on (`AIR-1.1`, `AIR-1.2`)
   keep their weight of 5 each unchanged; the reduction is absorbed by the
   other eleven, heaviest first, none dropping below 1.
4. **Dimensions are still declared in descending weight order** — see
   *Weights* — which is why *Alternate representations* lands at 7 rather
   than the 5 its own checks would otherwise need: dimension 7 must weigh at
   least as much as dimension 8, and dimension 8's own floor is 7.

#### 6.0 — 2026-09-11

**A major bump, on the letter of the versioning rule rather than its spirit,
same as 5.0.** No check's detection, band anchor, gate, or weight moved.
What changed is how a check is named.

1. **Every check is renumbered so its id prefix names its own dimension.**
   `AIR-1.1` through `AIR-1.13` are Dimension 1, `AIR-2.1` through `AIR-2.6`
   are Dimension 2, and so on through `AIR-8.7` — this reverses 5.0's
   decoupling (above), which let a check's id and its `dimension` field
   disagree. checks.yaml is now grouped and ordered by dimension, and
   `_validate_ids` enforces the ascending order that follows from it.
2. **The id prefix itself changes, from `AR-` to `AIR-`.** Every check id,
   every anchor (`#air-1.1`, not `#ar-1.1`), and every mention of a check
   anywhere the site or this document names one, uses the new prefix.
3. **A dimension has one canonical label, used everywhere one is named:**
   `#N · Title Case Name · N checks · W% weight` — for example,
   `#1 · First Contact · 13 checks · 36% weight`. The methodology page's own
   dimension headings and the checks list both follow it now.
4. **Nothing is re-scored.** A site scored under 5.0 and the same site scored
   under 6.0 differ only in which id and dimension label each finding is
   printed under, not in any check's verdict, weight, or band.

#### 5.0 — 2026-09-10

**A major bump, on the letter of the versioning rule rather than its spirit.**
Nothing here changes what a check detects or how a band is assigned — no
detection logic moved, no band anchor moved, no gate moved. What changed is
the accounting underneath the checklist, which is exactly the kind of change
the rule says must never happen silently.

1. **Weights are integers, on the same 100-point scale.** Every dimension
   total and every check weight is now a whole number, and the eight
   dimension totals still sum to exactly 100. That forces nine checks —
   `AIR-7.4`, `AIR-5.5`, `AIR-5.6`, `AIR-6.4`, `AIR-6.5`, and four of the five
   checks in *Measurement and instrumentation* — to round to zero: each was
   already worth less than half a point, and a 100-point scale has no
   smaller unit to give them. A check worth zero points cannot move a score.
   That is a known, accepted cost of staying on 100 rather than moving to a
   finer-grained scale, not an oversight — recorded here rather than left
   for someone to discover. The leading tier's own weights (measured, still
   never scored) round the same way, with a floor of 1 rather than 0: they
   describe a real practice worth naming even at the bottom of the scale.
   (Fixed in 7.0, below — every core check now floors at 1 too, by moving
   weight rather than by widening the scale.)
2. **A dimension is a field, not a naming convention.** The registry used to
   validate that a check's id prefix matched its `dimension` — `AR-2.3` (now
   `AIR-1.7`) had to live in dimension 2. That coupling is gone. A check's id
   is permanent from the day it is written; which dimension it is grouped
   under can move, and now does. (Reversed in 6.0, below — the id prefix is
   coupled to the dimension again, this time by renumbering the check rather
   than by validating against it.)
3. **First contact is a new dimension: exactly the 13 checks the free quick
   scan runs**, pulled out of six different dimensions they used to sit in.
   The quick scan's own weight and the number this document states for
   Dimension 1 are now the same claim — they used to be two different sums
   that happened to disagree. `AIR-4.1`–`AIR-4.4` (RSL, Web Bot Auth,
   `X-Robots-Tag`) — what was left of the old Dimension 1 once its
   quick-scan checks moved out — folded into the old Dimension 2, renamed
   *Crawl, index and licensing hygiene*, so the Index still has 8 dimensions,
   not 9.
4. **Dimensions are ordered by weight, descending; checks within a dimension
   the same way, gates pinned first.** Both orders are computed from the
   registry, never a hand-maintained list — see *Weights*, and every
   `### AIR-x.y` section in this document, for the result.
5. **Nothing is re-scored.** Every scan already on record keeps its `4.0`
   stamp and renders exactly as it did the day it ran; the Census stays on
   4.0 until it is next run. `docs/AI-READINESS-INDEX.md`'s own rule above —
   re-score before comparing across a major bump — still applies to anyone
   who wants to line up a 4.0 result next to a 5.0 one; nothing here does
   that automatically, because nothing here changed what a re-score would
   find. A site scored under 4.0 and the same site scored under 5.0 differ
   only by integer rounding, not by any check's verdict.

#### 4.0 — 2026-09-08

**A major bump, and the versioning rule applies in full.** Any site scored under
3.x must be re-scored under 4.0 before its number is compared to anything.

1. **A gate failure no longer caps the score.** `total` is the site's real, uncapped
   score, computed exactly as in *Scoring*; `gated` and `gate_failures` are reported
   beside it as a separate fact, never blended into it. The old model produced one
   number standing for two facts and comparable to neither — a gated site with
   excellent structure scored identically to a gated site with none. `ungated_score`,
   `ungated_grade`, and the `Gated` grade are gone rather than kept as duplicates of
   what `total` and `gated` already say between them. See *Gates*.
2. **A leading tier exists for checks with essentially no 2026 adoption.** Six checks
   move out of the hundred points entirely — scoring them would subtract the same
   points from every site, which moves nobody relative to anybody. They are still run
   and still reported, as `leading_adopted` of `leading_total` (adopted meaning band 2
   or better), so a site gets credit for being early without it changing anyone's
   score.

#### 3.0 — 2026-09-07

**A major bump, and the versioning rule applies in full.** Any site scored under
2.x must be re-scored under 3.0 before its number is compared to anything.

1. **The crawler obeys the `robots.txt` group addressed to it.** Until now the
   scanner read `robots.txt` only to score it, and crawled regardless of what it
   said about the scanner itself. It now honors a `Disallow` addressed to its own
   product token, and a site that turns it away is recorded as exactly that —
   `robots.txt` is still fetched, so AIR-1.1 still resolves and the gate still
   answers, while every other check reports insufficient evidence. **This changes
   which URLs are sampled, and therefore what a site can score, which is why the
   version is 3.0.**
2. **A blanket `User-agent: *` disallow does not stop the crawler**, and the
   reasoning is now published rather than implied. `Crawl-delay` under `*` is
   honored regardless. See *Consent*.
3. **The AIR-1.2 probe is documented rather than glossed.** The behavior has not
   changed — the probe has always carried other crawlers' product tokens — but
   the specification previously said we "do not pretend to be somebody else's
   bot", which was not an accurate description of the wire. It now states what
   the requests carry and the three constraints that make it measurement rather
   than impersonation. No score moves; the honesty of the document does.
4. **A survey population is a first-class concept**, with the sourcing and
   citation rules a published measurement needs. See *Populations*.
5. **Every result carries the Index's effective date** alongside its version, so
   an old report cannot silently start meaning something new.

#### 2.0 — 2026-09-01

**A major bump, and the versioning rule applies in full.** The one-time exemption
taken at 1.1 is spent. Any site scored under 1.x must be re-scored under 2.0 before
its number is compared to anything.

1. **AIR-1.1's bands are anchored on agent class rather than agent count**, and the
   gate now turns on the answer-serving class. Blocking every training crawler costs 2
   of 5 points and does not gate; blocking one answer-serving agent does. Under 1.x the
   table could not tell `PerplexityBot` from `CCBot` — both scored band 3 — which meant
   a site invisible to an answer engine and a site excluded from a corpus were graded
   identically. **This changes which sites are gated, which is why the version is 2.0.**
2. **`critical` severity widens to band 0 or 1** on checks weighing 3.0 or more. A
   6-point check at band 1 is not a `major` finding sitting alongside a missing table
   caption. Ordering and points are unchanged; only the label moves.
3. **AIR-4.11, AIR-3.1 and AIR-8.7 gain a band 3**, and their unattested ceilings rise from
   2 to 3 to match. Their tables previously skipped 3, so with a ceiling at 2 the whole
   observable range was `{0, 2}` and there was no way to say *mostly right*.

#### 1.2 — 2026-09-01

Refinements found while implementing all 68 checks. None changes a weight or a gate.
Two change how a band is reached and are recorded here so a re-score is explicable.

1. **`@id` coverage in AIR-1.11 counts entity nodes only.** A `ListItem` inside a
   `BreadcrumbList`, a `PostalAddress` on an `Organization` and a `SearchAction`'s
   `EntryPoint` are value objects: they have no identity and nothing should reference
   them. Counting them penalized a site for marking breadcrumbs up correctly.
2. **AIR-1.12 resolution is opt-in.** Band 4 requires `sameAs` links to resolve, which
   means requesting Wikidata, ROR and LinkedIn — hosts the client never allowlisted,
   and hosts that learn which site is being audited. The scanner classifies by
   authority tier from the markup alone and declares an observed ceiling of band 3
   until told to fetch.
3. **AIR-6.6 expects an `autocomplete` token only where the field type implies one**
   (`email`, `tel`, `url`, `password`). A search box has no sensible token and must
   not be marked down for lacking one.
4. **"Widespread" in AIR-4.6 needs a count as well as a share.** One three-hop chain on
   a two-page site is a finding, not a systemic failure.
5. **Every band ceiling now states its reason** in the result and in the report:
   awaiting a baseline run, awaiting confirmation from the client, limited by the band
   attested, or needing review by a person. A client told a number was held down is
   owed the reason.
6. **AIR-4.11's detection method is weaker than this document implied.** The IndexNow key
   file is named after the key, and the key is not published, so a third party can only
   find it where the site advertises it. A negative result is *not discoverable*, which
   is weaker than *absent*, and the check says so.
7. **AIR-5.6 is not implemented.** Reading C2PA manifests means parsing image and PDF
   bytes. The check reports that it had no evidence and leaves the denominator rather
   than scoring a site zero for a measurement nobody took.

#### 1.1 — 2026-09-01

Twelve gaps found while building the first scanner against 1.0. Several of these would ordinarily force a major bump, because they change how a score is computed. They are landing as a minor bump because **no site has ever been scored under 1.0** — the scanner did not exist yet, so there is no trend line to break. This exemption applies once. Every change after this one follows the rule above.

1. **Automated coverage** now counts `auto-partial` at half weight. Under the previous wording the metric could not exceed 0.83, while this document's own example result showed 0.87 — arithmetically unreachable.
2. **Attestations may lift an `auto-partial` band ceiling**, never set or lower a band. Four checks (AIR-4.11, AIR-1.10, AIR-3.1, AIR-8.7) had a top band this document itself says requires attestation, and no mechanism to supply one.
3. **Cross-run checks cap at band 3 on a first run** (AIR-6.9 at 2) and lift with a supplied baseline. AIR-4.8, AIR-2.4 and AIR-6.9 previously required a second audit with no definition of what run one does.
4. **The ten AI agents are sorted into three classes.** AIR-1.1's bands referenced categories the document never defined, and AIR-1.2 asked us to probe two user-agents that no crawler ever sends.
5. **Money-page double weighting is defined** as counting 2 in both sides of the coverage rubric. Previously asserted and never operationalized.
6. **A gated run grades as `Gated`** and gains an `ungated_grade` field. The bands table offered two labels for the same run.
7. **"Requires AIR-X.Y to pass" means band ≥ 2**, with N/A cascading to dependents. Previously undefined against a 0–4 scale.
8. **Business-fact applicability comes from a declared engagement profile**, never inference. Undeclared scores N/A with a reason.
9. **A 25-URL representation sub-sample** bounds AIR-4.4, AIR-7.2, AIR-7.4 and AIR-7.5, which together would otherwise triple the request count against a client's origin.
10. `confidence` is an enum: `high`, `medium`, `low`.
11. AIR-2.1 coverage clamps to 1.0 and excludes empty renders; a missing `robots.txt` scores AIR-1.1 band 4; `score_pct` is `null` where nothing applies; backlog ties break on check ID; `site.cms` is best-effort and nullable.
12. **Rounding happens once**, and the file records that per-check points will not sum exactly to the total.

## Appendix — Test index

| Test | Points |
|---|---|
| AIR-1.1 | 5 |
| AIR-1.2 | 5 |
| AIR-1.3 | 1 |
| AIR-1.4 | 3 |
| AIR-1.5 | 1 |
| AIR-1.6 | 1 |
| AIR-1.7 | 2 |
| AIR-1.8 | 2 |
| AIR-1.9 | 2 |
| AIR-1.10 | 2 |
| AIR-1.11 | 2 |
| AIR-1.12 | 2 |
| AIR-1.13 | 2 |
| AIR-2.1 | 6 |
| AIR-2.2 | 3 |
| AIR-2.3 | 2 |
| AIR-2.4 | 2 |
| AIR-2.5 | 1 |
| AIR-2.6 | 1 |
| AIR-3.1 | 2 |
| AIR-3.2 | 2 |
| AIR-3.3 | 4 |
| AIR-3.4 | 2 |
| AIR-3.5 | 1 |
| AIR-3.6 | 1 |
| AIR-3.7 | 1 |
| AIR-3.8 | 2 |
| AIR-4.1 | 1 |
| AIR-4.2 | 1 |
| AIR-4.3 | 1 |
| AIR-4.4 | 1 |
| AIR-4.5 | 1 |
| AIR-4.6 | 1 |
| AIR-4.7 | 2 |
| AIR-4.8 | 2 |
| AIR-4.9 | 1 |
| AIR-4.10 | 1 |
| AIR-5.1 | 2 |
| AIR-5.2 | 2 |
| AIR-5.3 | 2 |
| AIR-5.4 | 2 |
| AIR-5.5 | 2 |
| AIR-5.6 | 2 |
| AIR-6.1 | 2 |
| AIR-6.2 | 2 |
| AIR-6.3 | 2 |
| AIR-6.4 | 2 |
| AIR-7.1 | 2 |
| AIR-7.2 | 2 |
| AIR-7.3 | 2 |
| AIR-7.4 | 2 |

**51 tests · 100 points**
