ReadinessIndex.io Logo, orange and white bar charts on a purple background

AIRBot · Index 11.0

You found us in your logs

Last updated September 2026. Questions or corrections: hello@readinessindex.io.

AIRBot reads public web pages and scores them against the AI Readiness Index — an open standard for whether machines can find, read, trust and use a website. If you are here because you saw it in your access logs, this page tells you exactly what it did and how to stop it.

How to recognize it

Every request carries this user agent:

AIRBot/2.0 (+https://readinessindex.io/bot)

Most requests also carry a From header with a real address you can write to. Grep for AIRBot and you have found all of it.

What it does

It fetches your home page, follows your sitemap and navigation, and samples a small number of pages — around 30 for a survey run, up to a couple of hundred for a site we have been asked to audit. It reads robots.txt, llms.txt, your sitemap, a handful of well-known paths, and the HTML of the pages it sampled. It renders pages in a headless browser to compare what a crawler sees against what a person sees.

It does not submit forms, follow links behind a login, or fetch anything it was not pointed at. It reads. That is all it does.

How hard it crawls

One request every few seconds, per site, with one worker. It honors Crawl-delay as a floor and never goes faster than its own limit. It backs off on 429 and 503, and abandons a site rather than retrying into a wall. There is a hard cap on total requests per run, and hitting it aborts the run rather than continuing.

If AIRBot ever looks like load to you, that is a bug on our side. Tell us and we will fix it.

How to turn it off

Name it in your robots.txt and we stop:

User-agent: AIRBot
Disallow: /

We record that we were turned away — which still scores AIR-1.1, because robots.txt is the one file we always fetch — and take nothing else. No appeal, no exception, no "but this one is different."

One thing we do that deserves saying plainly

A blanket User-agent: * disallow does not stop AIRBot. You have to name it.

That is a deliberate choice and a strict reading of RFC 9309 would call it wrong, so here is the reasoning rather than a footnote. A * rule is aimed at the anonymous scrape — the thing with no name, no published policy, and nobody to write to. AIRBot is none of those: it identifies itself, it publishes this page, it takes one page every few seconds, and it will stop the moment you say its name. We would rather be turned off by a decision than by a default nobody revisited since 2011.

We still honor Crawl-delay under *. Politeness costs us nothing, and a rate limit is not an exclusion.

Why you may see other crawler names from us

This is the part most worth reading, because it looks worse in a log file than it is.

One of the things the Index measures is whether AI crawlers can actually reach your site — AIR-1.2. Your robots.txt states an intention; your CDN states a fact, and the two disagree more often than anyone expects. A firewall rule can turn away a crawler your robots.txt welcomes, and you would never know.

The only way to measure that is to ask. So on up to ten of the pages we already fetched, we send a request carrying each AI crawler's product token — GPTBot, ClaudeBot, OAI-SearchBot, PerplexityBot and four others — and record what comes back. Status code, response size, whether we got a challenge page instead of your content.

Three things about that, and we hold ourselves to all three:

  • Those requests still identify us. They carry X-Audit-Operator: AIRBot/2.0 and the same From header as everything else. If you look at the whole request rather than the user agent alone, we are not hiding.
  • They only ever go to pages we were already allowed to fetch. The probe runs over the sample, and the sample is filtered by your robots.txt first. We can never reach a page under somebody else's name that we declined to reach under our own.
  • If you turn AIRBot off, there is no sample, and therefore no probe. Naming us in robots.txt stops all of it.

We could avoid this by not measuring AIR-1.2 at all. We think a readiness score that skipped "can an AI crawler actually reach this page" would be worth less than the awkwardness of this paragraph.

Why you, specifically

Two reasons, and it is one or the other.

Somebody asked. A person typed your URL into the box on our home page, or you are a TEN7 client and we are auditing your site. Anyone can scan any public site, so this does not mean much about who asked.

Or you are in the Readiness Index Census. Periodically we scan a defined population — every US state portal, the largest universities, health systems, nonprofits — and publish what we find. Nobody asks us to, and we do not ask permission first, the same way a survey of the web's accessibility does not ask permission first. How the population was chosen, and the source list it came from, is published with the results. If you are in it, you can have your own full report free, on request.

The results are yours to check

The methodology is published in full under CC BY-SA 4.0 — every test, every point, every band. You can read what we measured, disagree with it, implement the Index yourself and score your own site without us. See the methodology, the 51 tests and licensing.

Every report we publish records which version of the Index it was scored against and the date that version took effect, so an old result stays readable rather than quietly meaning something new.

If we got something wrong

Write to hello@readinessindex.io. Include the URL, and the report link if you have one.

We keep the raw responses behind every scan, so we can re-run a test against exactly what we saw rather than guessing. Three things can come out of that, and we tell you which one it was: our scanner was wrong and we fix it and re-scan you; the test itself was wrong and we change it for everybody; or the finding stands and we show you the evidence it stands on. A change that would move existing scores gets a version bump, so nobody's old report silently changes meaning.

And if you would simply rather we did not: name AIRBot in your robots.txt. That is enough, and we will not treat it as a problem to be solved.

Is your site ready for AI?

Scan a page, get a number you can verify. Free.