Internet Archive

3 crawlers operated by Internet Archive

Who Internet Archive is, for site owners

Internet Archive operates 3 crawlers in our directory, filed under Other Agents. Of those 3, 3 rated caution for crawl behaviour.

Blocking Internet Archive carries an SEO impact score of 0/10 across its crawlers. Every crawler here carries the same blocking impact: Varies - Evaluate before blocking.

Our recommendation for Internet Archive, set by its least reassuring crawler (archive.org_bot, rated caution): Consider blocking based on your content strategy.

Internet Archive at a glance

OperatorInternet Archive
Crawlers in directory3
CategoriesOther Agents
SEO impact if blocked0/10
Crawl behaviour3 rated caution

Crawlers operated by Internet Archive

How to block or allow Internet Archive in robots.txt

These rules cover all 3 crawlers Internet Archive operates. Add them to the robots.txt file at the root of your domain.

Block Internet Archive completely
User-agent: archive.org_bot User-agent: ia_archiver User-agent: ia_archiver-web.archive.org Disallow: /
Allow Internet Archive explicitly
User-agent: archive.org_bot User-agent: ia_archiver User-agent: ia_archiver-web.archive.org Allow: /

robots.txt is a request, not an enforcement mechanism. A crawler that ignores it needs a block at the server or CDN layer. Rules are case-sensitive on the path and matched against the user-agent token, not the full user-agent string.

Frequently asked questions about Internet Archive

What crawlers does Internet Archive run?
Internet Archive operates 3 crawlers in our directory: archive.org_bot (user-agent archive.org_bot), ia_archiver (user-agent ia_archiver), ia_archiver-web.archive.org (user-agent ia_archiver-web.archive.org).
Should I block Internet Archive?
Consider blocking based on your content strategy. That guidance follows this operator's least reassuring crawler, archive.org_bot, rated caution, because allowing an operator allows every crawler it runs. The SEO impact score for blocking is 0/10, where 10 means blocking costs you the most visibility.
How do I block Internet Archive in robots.txt?
Add the following to robots.txt at your domain root: User-agent: archive.org_bot then User-agent: ia_archiver then User-agent: ia_archiver-web.archive.org then Disallow: /. One Disallow line covers all 3 user-agent tokens.
Does blocking Internet Archive hurt my search rankings?
Every crawler here carries the same blocking impact: Varies - Evaluate before blocking. Google's own ranking crawler is Googlebot, which this operator does not run, so blocking Internet Archive does not by itself remove you from Google's index. What it can change is whether AI assistants and analytics tools can read your pages.

Other operators in the same categories