Changelog

Version history and updates for AI Crawler Check.

v6.0 Major Release

The 2026 Accuracy Release: Honest Scoring, Bot Intelligence & New Builder Tools

Our biggest accuracy and honesty upgrade. We recalibrated the entire AI Visibility Score against what vendors actually document in 2026, added a live bot-intelligence layer that flags deprecated tokens and stealth crawling, and shipped four new free tools. Every one of our 146 articles was re-verified for 2026 accuracy.

Why this release exists. Through 2025 most AI visibility tooling, including ours, quietly inherited assumptions that stopped being true. Scores rewarded publishing an llms.txt file as though it were a confirmed ranking input. Checkers kept treating anthropic-ai as a live crawler years after Anthropic replaced it. Opt-out preference tokens were counted as bots. The result was a familiar failure mode: a site scores well, the owner relaxes, and ChatGPT still cannot read a single page. v6.0 is a deliberate correction. Where the evidence is strong we score it heavily; where it is a reasonable hedge we say so and weight it lightly; and where nothing can be measured from the outside we tell you that rather than inventing a number.

Scoring, rebuilt from the evidence

  • SCORE v6 (70/15/15): Bot Access now carries 70 of 100 points, AI Infrastructure an honest 15, and a brand-new Technical Readiness category the remaining 15 (sitemap 6, structured data 6, HTTPS 3). The reasoning: crawler access is binary and verifiable. If GPTBot is disallowed, your content cannot be cited, and no amount of schema markup compensates. So access dominates the score. The previous weighting let a site with excellent AI files but a blocked GPTBot outrank a plain site that was fully readable, which is exactly backwards. Every weight, and the argument behind it, is now documented on the methodology page.
  • LLMS.TXT DEMOTED, DELIBERATELY: Publishing llms.txt used to be worth far more. It is now a small part of a 15-point category. No major AI vendor has confirmed llms.txt as a retrieval or ranking input. It remains a sensible, low-cost hedge and we still check it strictly, but inflating its weight produced falsely reassuring scores. We would rather lose the neat narrative than overstate the evidence.
  • WEIGHTED BY REAL IMPACT: Blocking GPTBot, ClaudeBot, PerplexityBot or Googlebot costs materially more than blocking a minor scraper, and the score reflects that instead of counting every user-agent equally.

2026 bot intelligence

  • DEPRECATED TOKEN DETECTION: We flag retired user-agents such as anthropic-ai and claude-web that are still sitting in robots.txt files. This is the most common silent failure we see. An operator renames its crawler, the old rule keeps matching nothing, and the site owner believes they are still governing a bot that no longer exists under that name. Worse, the intended policy is not being applied at all. We name the stale token and give you the current replacement.
  • OPT-OUT TOKENS ARE NOT CRAWLERS: Google-Extended and Applebot-Extended are now correctly labelled as training-preference signals rather than fetchers. Neither ever requests a page. Disallowing Google-Extended opts you out of Gemini model training but does not remove you from AI Overviews, and blocking it will not reduce your crawl traffic by a single request. Tools that list it as a bot lead people to the wrong conclusion about what they just changed.
  • FETCHERS THAT IGNORE ROBOTS.TXT: Some user-triggered fetchers do not consult robots.txt at all, and we now say which, instead of implying a rule will stop them. When a person pastes your URL into an assistant, that request is often treated as user-initiated rather than crawling. A robots.txt rule is the wrong tool there; edge or WAF rules are. Telling you a disallow line works when it does not is worse than saying nothing.
  • STEALTH CRAWLING FINDINGS: Where independent research has documented undeclared or misrepresented crawling, most prominently the Perplexity findings, the report surfaces it so your policy accounts for traffic that may not honour your rules.
  • AGENTIC BROWSERS, HONESTLY UNDETECTABLE: ChatGPT Atlas, Comet and Copilot Actions drive a real browser session on a user's behalf. They largely present as ordinary browser traffic, so no robots.txt audit, ours included, can confirm or deny their access. We list them as a known blind spot rather than scoring a category we cannot measure. Naming the limit is more useful than a confident guess.
  • HONEST WAF HANDLING: If a firewall blocks our probe we report exactly that, across 15+ recognised providers, instead of recording a passing grade we did not verify.

Four new free tools

  • LLMS.TXT GENERATOR: The llms.txt builder gives you live preview, a correctly formed Optional section, and starting templates for SaaS, blogs, e-commerce and agencies. Most hand-written llms.txt files we encounter are malformed in ways that make them useless to parse, so the generator enforces valid structure as you type.
  • SCHEMA GENERATOR: The JSON-LD generator covers 8 types: Organization, LocalBusiness, Article, FAQPage, Product, WebSite, BreadcrumbList and HowTo. Structured data is one of the few levers that demonstrably helps machines read entities and relationships correctly.
  • BOT COMPARISONS: The /compare hub ships 20 side-by-side matchups such as GPTBot vs ClaudeBot, each with user-agent strings, safety rating, SEO impact and copy-paste blocking rules, for the common case of allowing one operator while blocking another.
  • SCORE BADGE: After any check you can embed a live SVG score badge in a site footer or README, with one-click HTML and Markdown embed codes.

Standards and content

  • WEB BOT AUTH: We now detect the IETF Web Bot Auth directory at /.well-known/http-message-signatures-directory. This is the emerging answer to a real problem: user-agent strings are trivially spoofed, so allow-listing by name has never been trustworthy. Cryptographic signatures let a crawler prove which operator it belongs to. Adoption is early, which is precisely why it is worth knowing whether the operators reaching you support it.
  • 146 ARTICLES RE-VERIFIED: Every guide was reread against 2026 reality, not spot-checked. That means honest llms.txt framing, deprecated-token warnings, corrected user-triggered fetcher behaviour, the Perplexity compliance findings and recalibrated score references throughout. Advice that was accurate in 2024 is now actively misleading in places, and stale guidance from a tool that claims to measure accuracy is a contradiction we were not willing to ship.
  • 196 CRAWLERS, ONE SOURCE OF TRUTH: The homepage checker, directory, validator, generator and batch checker all continue to report the same verified 196 bots across 8 categories, now including agentic entries like Google-Agent (Project Mariner) and OAI-AdsBot.

Verified numbers, documentation and tool comparisons

A tool whose pitch is honest measurement had been mis-stating its own figures: different pages claimed different bot totals, and the downloadable copy of the bot database had drifted to 154 entries while the real count was 196. Editing the text would have fixed nothing, because the cause was that any number could be typed by hand anywhere. The numbers now come from one source and a build gate refuses to ship the site when a claim disagrees with the database.

  • One source of truth for every figure. Bot counts, page counts, category totals and the batch limit all come from one file, and the build fails if any page disagrees with it.
  • The downloadable bot database is generated, not maintained by hand. It is rebuilt from the live database on every build, so it cannot fall behind again.
  • New documentation section. Step-by-step guides for running a scan, reading your score, fixing a blocked crawler, and auditing client sites.
  • New comparison section. How this tool compares with the alternatives, including the rows where the other tool wins and who should choose it instead.
  • Better crawlability and accessibility. Corrected robots.txt rules, redirects for renamed pages, a real 404 page, AA link contrast and semantic landmarks throughout.
  • Safer outbound links. External links no longer leak referrer data or expose the originating tab.

Mobile, rebuilt where it was actually broken

Most of this release was about being honest with numbers. This part is about being honest that the site was noticeably worse on a phone than on a desktop, which is where a majority of visitors read it. Every fix below was measured at real handset widths rather than judged by eye in a resized browser window, and desktop layout is deliberately untouched.

  • THE MENU NOW OPENS AND CLOSES PROPERLY: The mobile navigation was rebuilt as a slide-in panel with every tool reachable in one tap. The old panel could leave the page unable to scroll after it closed, which reads to a visitor as a frozen site rather than a stuck menu. It also added invisible width to every page, so the whole site could be dragged sideways a little on every screen.
  • THE MENU BUTTON IS VISIBLE ON EVERY PAGE: On the GEO audit tool the button was there and worked, but rendered as a completely blank square. The three lines of a menu button come from an icon font. That page was not loading the icon font at all, so all of its icons were invisible while the markup around them was correct. Nothing errored and nothing looked broken in the code, which is why it survived so long. The build now refuses to ship a page that mounts the navigation without the icons it needs.
  • TAPPING A TOOL LANDS ON THE TOOL: Choosing a destination from the mobile menu used to land part-way into the section, with the heading scrolled up out of view.
  • THE FIRST SCREEN IS THE FIRST SCREEN: On a phone the opening section now fills the screen exactly, instead of leaving a strip of the next section showing underneath. The first thing a visitor read below the main call to action was the start of an unrelated headline further down the page. It now tracks the live viewport, so it stays correct when the browser's address bar slides away.
  • READABLE AT 320px: Text, tap targets and the comparison tables were checked at the narrowest phone widths still in use, not just at modern handset sizes.

Every page is now measurable, and privacy copy is tighter

  • ANALYTICS COVERAGE ON EVERY PAGE: Two pages, the GEO audit tool and the not-found page, were loading no analytics at all. This is the quietest kind of fault. There is no error, nothing looks wrong, and the reporting does not warn you: it simply shows fewer pages, which is indistinguishable from those pages having no visitors. A release check now verifies that all 492 listed pages plus the unlisted tool and error pages carry analytics and the cookie banner, and separately confirms in a real browser that the tag stays silent until consent is given and genuinely starts once it is. Both error pages are included, because a not-found page is exactly the one you want in reporting: every visit to it is a broken link worth fixing.
  • CONSENT STILL COMES FIRST: No analytics cookie is set and no tracking script we load runs until you accept. Rejecting also clears anything previously set, and the Global Privacy Control browser signal is honoured without needing a click. One exception is now written into the cookie policy rather than left unsaid: our hosting provider adds a cookieless visitor counter to every page after our code has run, so the banner cannot switch it off. It stores nothing on your device and cannot recognise you on a later visit or on any other site. It is disclosed in full, including how to block it.
  • FEWER TECHNICAL DETAILS IN THE POLICIES: Our privacy and cookie pages still name every processor we use, which is what the disclosure is for, but no longer print internal account identifiers alongside them. Those identifiers gave a reader nothing useful and are not required by any disclosure rule. A release check now scans published pages for credentials, tokens and internal identifiers so none can reach a visitor by accident.

Upgrade note: because the weighting changed, a score from before v6.0 is not directly comparable with one after it. Sites that were blocking a major AI crawler while publishing tidy AI files will generally see a lower, more accurate number. That drop is the fix working, not a regression. Re-run your check to get a current baseline, and see the methodology page for the full formula.

v5.0

The Big Sync: 196 Bots Everywhere, Gamified Learning & a Verified Single Source of Truth

Our largest consistency and content update yet. We rebuilt the bot database from the ground up so that every single page (homepage checker, robots.txt validator, generator, batch checker, full directory, and all 146 articles) now reports the exact same number: 196 verified AI and web crawlers. We also turned every blog article into an interactive learning experience and expanded the directory with 35 brand-new AI bot profiles.

  • ONE SOURCE OF TRUTH: Eliminated every count mismatch across the site. The checker database, the crawler directory (196 detail pages summed across 8 categories), the validator, and the marketing copy are now perfectly synchronized at 196 bots. We also removed 2 duplicate Meta crawler entries that were inflating the count.
  • 35 NEW AI BOT PROFILES: Added full directory pages for the newest agents and search crawlers, including ChatGPT Agent, OpenAI Operator, Google NotebookLM, Gemini CLI, Project Mariner, AWS Bedrock Bot, Azure AI Search, Kimi, Manus, Devin, Firecrawl, Exa, Tavily, Brave Leo, and more. Each ships with user-agent strings, safety ratings, robots.txt rules, FAQ schema, and related-bot links.
  • GAMIFIED ARTICLES: Every one of our 146 guides now features a live reading-progress bar, XP-style milestone badges, an interactive knowledge-check quiz, an animated action checklist, count-up statistics, Myth vs Fact cards, one-click copy buttons on code blocks, and an auto-generated table of contents. Learning AI SEO has never been this engaging.
  • 89 OPERATOR PAGES: Expanded operator coverage from 72 to 89 companies, adding profiles for Moonshot AI, Cognition, Exa, Tavily, Linkup, Kagi, Phind, WRTN, Brave, iAsk, AddSearch, Klaviyo, Zhipu AI, Andi, aiHit, and Microsoft Azure.
  • AUTHOR FIX: Corrected the author profile link across all 146 articles to point to Brian Ho’s verified Horatos profile, strengthening E-E-A-T signals for AI search citations.
  • SITEMAP REFRESH: The XML sitemap now includes all 196 bot pages, 8 categories, 89 operators, and 146 articles so AI crawlers and search engines discover every new page immediately.
v4.6

Ultimate Unified Report: GEO + AI Crawler Check Combined

  • UNIFIED REPORT: All 30+ GEO audit checks merged directly into the main AI Crawler Check report: one domain check, one comprehensive result covering 196 bots + full GEO analysis
  • GEO SCORE: Overall GEO Score (0–100) with 5 weighted categories: Technical SEO (18%), Content Quality (22%), AI Readiness (30%), Security & Trust (15%), UX (15%)
  • 22 AUDIT CHECKS: SSL, sitemap, canonical, hreflang, schema markup, semantic HTML, llms.txt, AI content readiness, JS blocking, E-E-A-T-A scoring, CDN detection, and more: all inline in the report
  • ZERO EXTRA FETCHES: GEO analysis reuses homepage HTML from Step 0, robots.txt from Step 1, and llms.txt from Step 2: only sitemap requires an additional fetch
  • ENHANCED EXPORTS: PDF & Excel reports now include full GEO Audit page with category scores, 22-check table, and actionable recommendations
  • VISUAL DASHBOARD: Conic-gradient score ring, category progress bars, status-colored check cards, and prioritized recommendation list in the unified report
v4.5

GEO Audit Tool: Standalone Mode

  • GEO AUDIT TOOL: Full GEO audit tool available at /geo-audit (now superseded by unified report in v4.6)
  • SCHEMA FIX: All social media removed from Person schema: LinkedIn only (linkedin.com/in/ho-bao)
v4.4

Horatos.ai Rebrand & 163 Bots Live

  • REBRAND: AI Crawler Check is now developed by Horatos.ai: an AI Search Optimization (AISEO) & Generative Engine Optimization (GEO) consultancy. All author links, legal pages, schema markup, and branding updated
  • 163 BOTS DEPLOYED: All 8 new bots from v4.3 now live in production: OAI-AdsBot, MistralAI-Index, Google-NotebookLM, Google-Read-Aloud, Google-Site-Verification, FeedFetcher-Google, GoogleProducer, Amazon-Bedrock-AgentCore
  • LEGAL PAGES: Privacy Policy, Terms of Use, and Cookie Policy updated to reflect Horatos.ai as operating entity
v4.3

30 Blog Posts: E-E-A-T, GEO vs SEO & AI Bot Troubleshooting

  • 30 BLOG POSTS: Published 3 new E-E-A-T focused articles: E-E-A-T for AI Search, GEO vs SEO Complete Comparison, Why AI Bots Can’t Crawl Your Website
  • 12 NEW IMAGES: 3 hero images + 9 inline diagrams covering trust pyramids, GEO/SEO frameworks, diagnostic flowcharts, and robots.txt fixes
  • E-E-A-T AUTHORITY: Deep-dive into Experience, Expertise, Authoritativeness & Trustworthiness signals for AI search citation. Original frameworks and expert analysis
  • INTERNAL LINKS: 80+ strategic internal links across 3 posts. Heavy tool CTAs, cross-referencing to crawler guides, robots.txt generator, schema markup, and audit checklist
  • 163 BOTS: Added 9 new bots: OAI-AdsBot (ChatGPT Ads, Apr 2026), MistralAI-Index (Vibe search index), Google-NotebookLM, Google-Read-Aloud, Google-Site-Verification, FeedFetcher-Google, GoogleProducer, Amazon-Bedrock-AgentCore. Upgraded Amazonbot to Tier 2 (Nova AI training confirmed)
  • WAF FIX: Fixed false-positive bot blocking when target sites use Cloudflare WAF. Multi-UA fallback strategy with WAF detection
v4.2

3 New Blog Posts: Google-Agent, WebMCP & llms.txt Guide

  • 27 BLOG POSTS: Published 3 new in-depth articles: Google-Agent & Project Mariner, WebMCP Guide, How to Create llms.txt
  • 12 NEW IMAGES: 3 hero images + 9 inline diagrams covering agentic browsing, WebMCP architecture, and llms.txt structure
  • TOPICS: Agentic Engine Optimization (AEO), WebMCP Tool Contracts, user-triggered-agents.json, Web Bot Auth, llms.txt templates for 4 website types
v4.1

Google-Agent, 24 Blog Posts & Microsoft Clarity

  • NEW BOT: Google-Agent (Project Mariner): Google’s brand-new agentic AI crawler that navigates sites, fills forms, and takes actions on behalf of users. Monitor its activity to measure how important WebMCP is for your site
  • 155 BOTS: Bot database expanded from 154 → 155. Google-Agent added to Google Bots category with full directory page, FAQs, and structured data. Now 293 SEO-optimized directory pages (161 bot + 8 category + 89 operator)
  • 24 BLOG POSTS: Published 12 new articles (posts 13–24) covering robots.txt creation, Applebot-Extended, Meta-ExternalAgent, AI SEO Audit Checklist, Cohere AI Crawler, and more
  • E-E-A-T FOCUS: All new posts emphasize Experience, Expertise, Authoritativeness, and Trustworthiness with author bylines, expert recommendations, and data-backed analysis
  • ANCHOR TEXT: Strategic homepage anchor text distribution: “AI crawler checker free” (14x), “AI crawler checker online” (12x), “web crawler tool free” (6x), “AI crawlers” (4x)
  • CLARITY: Microsoft Clarity analytics added to all pages: heatmaps, session recordings, and user behavior tracking across homepage, blog, tools, directory, and legal pages
  • 48 NEW IMAGES: 12 hero images + 36 inline diagrams/infographics for blog posts 13–24. Custom illustrations for each article topic
v4.0

Focused AI Bot Checker: Simplified & Faster

  • FOCUS: Removed homepage signals parsing (Schema, H1, Meta Description, Canonical, OG tags): tool now focuses solely on AI bot access & infrastructure
  • SIMPLIFIED SCORE: New 2-component scoring: Bot Access (65pts) + AI Infrastructure (35pts) = 100pts. Removed Content (25) and Technical (10) components
  • REMOVED: Browser Rendering API fallback, WAF/bot protection detection, homepage HTML parsing, extractSchemas(), extractMeta(), countTag(): eliminates all false-negative complaints from WAF-blocked sites
  • NO MORE WAF ISSUES: Since homepage is no longer scanned, sites behind Cloudflare/LiteSpeed/Wordfence WAFs get accurate results with zero workarounds needed
  • SMARTER RECS: Only actionable recommendations: unblock AI bots, create llms.txt/llms-full.txt, upgrade partial to full access. No more schema/meta/author suggestions
  • CTA: “Want a Full AI Visibility Audit?” card replaces Homepage Signals section: directs users to paid GEO audit for deeper analysis
  • FASTER: Removed homepage fetch + retry + Browser Rendering: checks complete ~3s faster on average
v3.9

Browser Rendering Fallback: Bypass WAF with Headless Chrome

  • BREAKTHROUGH: When WAF blocks homepage scan, tool now falls back to Cloudflare Browser Rendering API: headless Chromium renders the page like a real browser, bypassing CAPTCHAs and WAF challenges
  • RESULT: WAF-protected sites (LiteSpeed, Cloudflare, Wordfence, Sucuri) now get full homepage analysis: Schema, H1, Meta Description, Canonical, OG tags all detected correctly
  • 3-TIER FETCH: Homepage fetch now uses three strategies: (1) Chrome UA fetch → (2) Safari UA retry → (3) Headless Browser Rendering fallback
  • UX: Green “Browser Rendering Used” badge shows when headless browser was needed to bypass WAF
  • TYPED: Added Env type bindings for CF_API_TOKEN and CF_ACCOUNT_ID environment variables
v3.8.2

Honest WAF Reporting & UX Overhaul

  • HONESTY: WAF-blocked signals now show ⚠ “Unable to verify” with yellow triangles instead of pretending they pass or fail: no more misleading results
  • UX: New 2-column layout in WAF warning: “✓ Accurate Results” vs “⚠ Unable to Verify”: users instantly see what’s reliable
  • NEW: Links to Google Rich Results Test & Schema.org Validator in WAF warning: guides users to verify schema manually
  • SMART: Recommendations now skip “Add FAQ Schema” and “Add Author signal” when WAF blocks scan: avoids recommending things the site may already have
  • UX: WAF-blocked signal rows show impact as “N/A” in yellow instead of misleading Critical/High labels
  • CLEANUP: Removed temporary debug endpoint /api/debug-signals
v3.8.1

Enhanced WAF Detection & Smarter Scoring

  • ENHANCED: WAF/Bot detection now covers 15+ providers: Cloudflare, LiteSpeed, Wordfence, Sucuri, Imperva, Akamai, StackPath, AWS WAF, hCaptcha, reCAPTCHA
  • NEW: Auto-retry with Safari UA when initial Chrome UA triggers bot protection: bypasses simpler WAF rules
  • IMPROVED: Bot-protected sites now score 19/25 content (was 13/25): assumes H1, meta desc, schema exist (standard for CMS sites behind WAF)
  • UX: Homepage signals show yellow ⚠ warning icons for WAF-blocked items instead of misleading red ✗ marks
  • IMPROVED: Homepage fetch now sends full browser headers (Accept, Sec-Fetch-*, Accept-Language): reduces WAF trigger rate
v3.8

Schema Detection Fix & Bot Protection Awareness

  • FIXED: Schema detection now traverses @graph arrays (WordPress/RankMath/Yoast): previously only checked root-level @type
  • FIXED: extractSchemas() rewritten as recursive collectSchemaTypes(): handles nested @graph, mainEntity, and arrays at any depth
  • NEW: Bot Protection Detection: identifies WAF/CAPTCHA pages (LiteSpeed, Cloudflare, Wordfence) and flags botProtection: true in results
  • IMPROVED: Bot-protected sites now receive partial content credit (13/25) instead of 0: no longer penalized for CAPTCHA pages
  • PERF: Homepage HTML cached from STEP 0 probe: eliminated duplicate fetch, saving ~200ms per check
  • IMPROVED: extractMeta() now handles multi-line attributes (WordPress meta tags with newlines between properties)
v3.7

AI Files Detection Overhaul & Score Rebalance

  • Fixed False Positives: llms.txt and llms-full.txt detection now uses strict validation: redirect: manual prevents 301/302 homepage redirects from being counted as “Found”
  • Content Validation: AI files must return 200 OK with non-HTML content. Responses containing HTML (custom 404 pages, redirected homepages) are correctly rejected
  • Smart Redirect Follow: Follows one redirect only if the target URL ends with the same filename (e.g., shopify.com/llms.txt → www.shopify.com/llms.txt). Homepage redirects are rejected
  • Removed ai.txt: /ai.txt has been removed from all checks: no actual standard exists for this file. Only llms.txt and llms-full.txt (official llmstxt.org standard) are checked
  • Score Rebalanced: Bot Access (40), AI Infrastructure (25), Content (25), Technical (10). Removed 5pts from ai.txt, redistributed to Content quality signals
  • Social Preview: All pages now have OG image + Twitter Card meta tags for rich social previews when sharing on Facebook, LinkedIn, X/Twitter, Slack
v3.6

Bot Database Accuracy & Versioned Bots

  • 154 Bots: Added 3 new versioned AI bots: ChatGPT-User/2.0, MistralAI-User-1.0, Perplexity-User-1.0: matching CrawlerCheck.com’s 34-bot AI directory
  • Deprecated Bot Labels: Claude-Web and anthropic-ai are now marked as Deprecated per Anthropic’s official docs (replaced by Claude-User, Claude-SearchBot, ClaudeBot)
  • ClaudeBot Relabeled: Changed from “Claude AI (Legacy)” to “Claude AI Training” to accurately reflect its official purpose per Anthropic docs
  • Versioned Bot Matching: New parser logic handles versioned user-agents (e.g., ChatGPT-User/2.0 inherits ChatGPT-User rules from robots.txt). Same for -1.0 suffixes
  • Cross-Verified: AI bot list verified against official sources: OpenAI Bots docs, Anthropic Crawler docs, CrawlerCheck.com directory, Search Engine Journal’s 2025 list
v3.5

Partial Access & Report Naming

  • Partial = Accessible: Bots with partial access (e.g., /admin blocked but / allowed) now count as “accessible” in overview. Only fully BLOCKED bots are counted as denied. This prevents misleading “0% allowed” for sites that only block admin paths
  • Category Overview Rework: Shows “X full · Y partial” breakdown instead of just “X/Y allowed”. Percentage now reflects real accessibility
  • Bot Card Labels: Status now shows “FULLY ALLOWED”, “PARTIALLY ALLOWED”, or “BLOCKED” for clarity
  • Report Naming: PDF/Excel exports renamed from “AI Visibility Report” to “AI Crawler Check Report”: accurate to the tool’s function
  • Score Adjustment: PARTIAL bots now receive 75% score credit (was 50%). Reflects that partial access still means content is crawlable
v3.4

Deep Verification & Accuracy Guarantee

  • Double Verification: robots.txt is now fetched twice and results compared. If content changed between fetches (CDN cache, etc.), the freshest version is used
  • Enforced Backend Processing: Minimum 8-second server-side processing time ensures all checks genuinely complete. No shortcuts, no cached guesses
  • Frontend Sync: Animation steps now match real backend processing. 10+ second minimum display with step-by-step progress visible to the user
  • Error UX Improved: Invalid domains show immediately clear error messages. No more fake results for non-existent websites
  • Version Tag: Reports now include processingTime to prove thorough analysis
v3.3

Domain Validation & Integrity

  • Critical Fix: Added mandatory domain reachability check (STEP 0). Invalid or non-existent domains (e.g., example.com.sgsg) now return a clear error instead of fake results
  • DNS Validation: The tool now verifies the website actually responds before analyzing. DNS failures, connection timeouts, and Cloudflare 5xx errors are properly detected
  • robots.txt Validation: Content is now validated to ensure it’s actually a robots.txt file (not a custom error page returning 200). Per RFC 9309: 404 = all allowed, 403 = all blocked
  • Minimum Display Time: Results always show for at least 6 seconds with progressive step animation to confirm thorough analysis
  • Error UX: Clear, actionable error messages for unreachable domains, DNS failures, and timeout scenarios
v3.2

Accuracy Fix: robots.txt Only

  • Critical Fix: Completely removed HTTP simulation. robots.txt is now the SOLE source of truth for bot access status. This eliminates ALL false positives caused by WAF/CDN/IP-based blocking (Cloudflare, Akamai, etc.)
  • Why: HTTP 403 from datacenter IPs does NOT mean the site owner blocks a bot: it's often generic bot-protection. Only robots.txt reflects the site owner's explicit intent
  • Partial Explained: Added tooltip explaining what "Partial" means on every bot card
  • Bot Purpose Labels: Every bot card shows its purpose: Training, Indexing, User Request, Both, Scraping, Monitoring, Social
  • Check Timing: Increased analysis animation to 14+ seconds for user trust
  • Directory UI Fix: Fixed broken layout in Web Crawler & Bot Directory pages
v3.1

Accuracy & UX Overhaul

  • Critical Fix: Resolved false-positive blocking reports caused by WAF/CDN detection (e.g., Cloudflare-protected sites returning HTTP 403 for all datacenter IPs)
  • Accuracy: robots.txt is now the canonical source of truth. HTTP simulation only upgrades severity, never overrides a robots.txt ALLOWED status
  • Bot Purpose Labels: Every bot now shows its purpose: Training, Indexing, User Request, Both (Training + User Request), Scraping, Monitoring, Social
  • Partial Explained: PARTIAL status now includes an inline explanation showing which paths are blocked vs allowed
  • New Logo: Custom SVG logo for AI Crawler Check replaces old BH branding in header
  • Full Navigation: Directory, Tools, Generator, Validator, Batch Checker links on every page
  • Check Animation: Progressive step-by-step loading with realistic timing for better user trust
  • Changelog: This page: track all updates and version history
v3.0

SEO Tools Suite

  • Robots.txt Generator: Create custom robots.txt with 196+ bot presets, 6 one-click presets, per-bot toggles
  • Robots.txt Validator: Paste or fetch robots.txt, analyze against 196+ bots, SEO Safety Score
  • Batch URL Checker: Check up to 20 URLs at once with CSV export and formatted reports
  • Bot Directory: 293 SEO-optimized pages: 196 bot detail pages, 8 category pages, 89 operator pages
  • SEO: JSON-LD structured data, FAQ rich snippets, XML sitemap with 248+ URLs
v2.0

151 Bots & 8 Categories (Initial)

  • Expanded to 196 bots across 8 categories initially (AI bots, search engines, Google bots, SEO tools, social bots, scrapers, cloud services, other agents)
  • PDF export with branded report
  • Meta-robots and X-Robots-Tag analysis
  • AI Visibility Score with 4-component breakdown
v1.0

Initial Launch

  • First release with basic robots.txt checking
  • Core AI bot database and analysis engine