669 sites
aren't ready.
Real data from 669 public scans. Updated hourly.
78% of sites miss robots.txt — the most commonly failed check.
Eight checks.
Each one a signal.
AI agents don't browse — they parse. Here's what each check detects and why it matters.
The emerging standard for telling AI agents exactly what your site does, what pages matter, and what to do with them. Without it, agents guess — and guess wrong.
Machine-readable metadata that tells agents your entity type, name, description, and relationships. Used by every major LLM training pipeline and retrieval system.
AI agents retrieve your pages via headless readers, not browsers. If your content is buried in JS-heavy templates, ads, or infinite scroll, agents see noise.
GPTBot, ClaudeBot, and Googleother all read robots.txt before crawling. Explicit allow rules signal you actively want AI traffic; silence is treated as ambiguity.
A sitemap gives agents a deterministic list of canonical pages. Without it, agents crawl by following links — missing deep content and paginated sections.
og:title and og:description are extracted by retrieval systems when displaying source citations. They define how your content appears in agent responses.
The title and description are the first text an agent reads. They anchor the page's purpose and determine whether it gets retrieved for a query.
Agents deduplicate content by canonical URL. Without a canonical tag, the same article at multiple paths gets indexed multiple times — diluting signal and wasting crawl budget.
The difference
is visible.
Two domains. Same web. Completely different results for AI agents.
See where your site stands.
Scan your site →