AIRBot · Index 11.0
You found us in your logs
Last updated September 2026. Questions or corrections: hello@readinessindex.io.
AIRBot reads public web pages and scores them against the AI Readiness Index — an open standard for whether machines can find, read, trust and use a website. If you are here because you saw it in your access logs, this page tells you exactly what it did and how to stop it.
How to recognize it
Every request carries this user agent:
AIRBot/2.0 (+https://readinessindex.io/bot)
Most requests also carry a From header with a real address you can
write to. Grep for AIRBot and you have found all of it.
What it does
It fetches your home page, follows your sitemap and navigation, and samples a
small number of pages — around 30 for a survey run, up to a couple of hundred for
a site we have been asked to audit. It reads robots.txt,
llms.txt, your sitemap, a handful of well-known paths, and the HTML
of the pages it sampled. It renders pages in a headless browser to compare what a
crawler sees against what a person sees.
It does not submit forms, follow links behind a login, or fetch anything it was not pointed at. It reads. That is all it does.
How hard it crawls
One request every few seconds, per site, with one worker. It honors
Crawl-delay as a floor and never goes faster than its own limit. It
backs off on 429 and 503, and abandons a site rather than retrying into a wall.
There is a hard cap on total requests per run, and hitting it aborts the run
rather than continuing.
If AIRBot ever looks like load to you, that is a bug on our side. Tell us and we will fix it.
How to turn it off
Name it in your robots.txt and we stop:
User-agent: AIRBot
Disallow: /
We record that we were turned away — which still scores
AIR-1.1, because robots.txt
is the one file we always fetch — and take nothing else. No appeal, no exception,
no "but this one is different."
One thing we do that deserves saying plainly
A blanket User-agent: * disallow does not stop AIRBot.
You have to name it.
That is a deliberate choice and a strict reading of
RFC 9309 would call it wrong, so here is the reasoning
rather than a footnote. A * rule is aimed at the anonymous
scrape — the thing with no name, no published policy, and nobody to write to.
AIRBot is none of those: it identifies itself, it publishes this page, it takes
one page every few seconds, and it will stop the moment you say its name. We
would rather be turned off by a decision than by a default nobody revisited
since 2011.
We still honor Crawl-delay under *. Politeness costs
us nothing, and a rate limit is not an exclusion.
Why you may see other crawler names from us
This is the part most worth reading, because it looks worse in a log file than it is.
One of the things the Index measures is whether AI crawlers can actually reach
your site — AIR-1.2. Your
robots.txt states an
intention; your CDN states a fact, and the two disagree more often than anyone
expects. A firewall rule can turn away a crawler your robots.txt
welcomes, and you would never know.
The only way to measure that is to ask. So on up to ten of the pages we already
fetched, we send a request carrying each AI crawler's product token —
GPTBot, ClaudeBot, OAI-SearchBot,
PerplexityBot and four others — and record what comes back. Status
code, response size, whether we got a challenge page instead of your content.
Three things about that, and we hold ourselves to all three:
- Those requests still identify us. They carry
X-Audit-Operator: AIRBot/2.0and the sameFromheader as everything else. If you look at the whole request rather than the user agent alone, we are not hiding. - They only ever go to pages we were already allowed to fetch.
The probe runs over the sample, and the sample is filtered by your
robots.txtfirst. We can never reach a page under somebody else's name that we declined to reach under our own. - If you turn AIRBot off, there is no sample, and therefore
no probe. Naming us in
robots.txtstops all of it.
We could avoid this by not measuring AIR-1.2 at all. We think a readiness score that skipped "can an AI crawler actually reach this page" would be worth less than the awkwardness of this paragraph.
Why you, specifically
Two reasons, and it is one or the other.
Somebody asked. A person typed your URL into the box on our home page, or you are a TEN7 client and we are auditing your site. Anyone can scan any public site, so this does not mean much about who asked.
Or you are in the Readiness Index Census. Periodically we scan a defined population — every US state portal, the largest universities, health systems, nonprofits — and publish what we find. Nobody asks us to, and we do not ask permission first, the same way a survey of the web's accessibility does not ask permission first. How the population was chosen, and the source list it came from, is published with the results. If you are in it, you can have your own full report free, on request.
The results are yours to check
The methodology is published in full under CC BY-SA 4.0 — every test, every point, every band. You can read what we measured, disagree with it, implement the Index yourself and score your own site without us. See the methodology, the 51 tests and licensing.
Every report we publish records which version of the Index it was scored against and the date that version took effect, so an old result stays readable rather than quietly meaning something new.
If we got something wrong
Write to hello@readinessindex.io. Include the URL, and the report link if you have one.
We keep the raw responses behind every scan, so we can re-run a test against exactly what we saw rather than guessing. Three things can come out of that, and we tell you which one it was: our scanner was wrong and we fix it and re-scan you; the test itself was wrong and we change it for everybody; or the finding stands and we show you the evidence it stands on. A change that would move existing scores gets a version bump, so nobody's old report silently changes meaning.
And if you would simply rather we did not: name AIRBot in your
robots.txt. That is enough, and we will not treat it as a problem to
be solved.
Is your site ready for AI?
Scan a page, get a number you can verify. Free.