Scan a site
Why It Matters
08/23/2026

669 sites
aren't ready.

Real data from 669 public scans. Updated hourly.

Total scans
669
Avg score
61.4 / 100
Most common grade
Good
Pass rate per check — sorted by most failed
robots.txt
22%
/llms.txt
42%
JSON-LD
55.5%
Canonical
68.9%
OpenGraph
76.7%
sitemap.xml
77.3%
Title+meta
85.2%
Clean content
86.5%

78% of sites miss robots.txt — the most commonly failed check.

Eight checks.
Each one a signal.

AI agents don't browse — they parse. Here's what each check detects and why it matters.

/llms.txt file
+25 pts

The emerging standard for telling AI agents exactly what your site does, what pages matter, and what to do with them. Without it, agents guess — and guess wrong.

The single highest-weight check. Missing it is like removing your homepage from every AI index.
42% pass
{}
JSON-LD structured data
+15 pts

Machine-readable metadata that tells agents your entity type, name, description, and relationships. Used by every major LLM training pipeline and retrieval system.

Agents that can't parse your type skip your content during indexing.
55.5% pass
Clean content
+15 pts

AI agents retrieve your pages via headless readers, not browsers. If your content is buried in JS-heavy templates, ads, or infinite scroll, agents see noise.

A site that looks great to humans can be completely blank to agents.
86.5% pass
robots.txt AI rules
+10 pts

GPTBot, ClaudeBot, and Googleother all read robots.txt before crawling. Explicit allow rules signal you actively want AI traffic; silence is treated as ambiguity.

Some agents opt out of ambiguous sites entirely.
22% pass
sitemap.xml
+10 pts

A sitemap gives agents a deterministic list of canonical pages. Without it, agents crawl by following links — missing deep content and paginated sections.

Agents with sitemaps index significantly more deep content.
77.3% pass
OpenGraph tags
+10 pts

og:title and og:description are extracted by retrieval systems when displaying source citations. They define how your content appears in agent responses.

No OG tags → agents cite your page without context, reducing trust.
76.7% pass
Title + meta description
+10 pts

The title and description are the first text an agent reads. They anchor the page's purpose and determine whether it gets retrieved for a query.

Missing or generic titles cause agents to misclassify your content topic.
85.2% pass
Canonical URL
+5 pts

Agents deduplicate content by canonical URL. Without a canonical tag, the same article at multiple paths gets indexed multiple times — diluting signal and wasting crawl budget.

A silent killer: duplicate content suppresses ranking in agent retrieval.
68.9% pass

The difference
is visible.

Two domains. Same web. Completely different results for AI agents.

Before — unoptimised
blog-example.com
20Poor
/llms.txt
JSON-LD
Clean content
robots.txt
sitemap.xml
OpenGraph
Title+meta
Canonical
Our traffic from AI-driven referrals was near zero. We had no /llms.txt, no structured data, and our JS-heavy template couldn't be read by agent crawlers.
After — fully optimised
stripe.com
90Excellent
/llms.txt
JSON-LD
Clean content
robots.txt
sitemap.xml
OpenGraph
Title+meta
Canonical
Stripe's developer docs have always been machine-readable by design. The /llms.txt was a natural extension of their existing technical documentation culture.

See where your site stands.

Scan your site →