Crawl Surface
The machine-readable surface of this site — discovery files, AI context endpoints, and structured data — and which crawlers access each layer. Updated every 4 hours from the crawl log; owner traffic excluded.
2,258
Total requests
10
Distinct crawlers
10
Endpoint types
2h ago
Last access
AhrefsBotClaudeBotChromeMajesticBotSemrushBotMozBot
Discovery layer
Standard files every crawler checks first
Crawl rules + Content-Signal permissions
ClaudeBot 307
AhrefsBot 97
MajesticBot 96
SemrushBot 88
MozBot 69
Chrome 61
OpenAI SearchBot 48
Other bot 47
+23 more crawlers
Canonical sitemap (linked from robots.txt)
AhrefsBot 512
ClaudeBot 308
GPTBot 33
Bingbot 33
Chrome 13
PetalBot 12
Other bot 5
GoogleOther 2
+10 more crawlers
AI context layer
Structured hints for LLMs and AI agents
LLM-readable site index (llmstxt.org spec)
Chrome 8
PetalBot 7
Other bot 5
AhrefsBot 4
Firefox 4
GoogleOther 3
Google (by ASN) 3
MozBot 2
+9 more crawlers
Machine-readable endpoint catalog (JSON-LD)
Chrome 29
Other bot 4
PetalBot 4
AhrefsBot 4
MozBot 2
Safari 2
Firefox 2
Applebot 1
+5 more crawlers
Entity map — people, concepts, and organizations (EntityMap v1.0)
Chrome 28
PetalBot 6
Other bot 4
AhrefsBot 4
Safari 3
Firefox 2
MozBot 1
Meta (Scraper) 1
+7 more crawlers
Raw content layer
.md / .mdx page requests — AI agents preferring plain text
44 requests
PetalBot 7
Safari 6
Chrome 5
Other bot 4
Firefox 4
AhrefsBot 4
SemrushBot 2
MozBot 2
+9 more crawlers
23 requests
Firefox 4
curl 3
GoogleOther 2
AhrefsBot 2
Safari 2
PetalBot 2
Other bot 2
GPTBot 1
+5 more crawlers
21 requests
Firefox 4
curl 3
AhrefsBot 2
PetalBot 2
Other bot 2
GPTBot 1
GoogleOther 1
Screaming Frog 1
+5 more crawlers
17 requests
Firefox 3
Chrome 2
AhrefsBot 2
PetalBot 2
Other bot 2
curl 1
MajesticBot 1
GPTBot 1
+3 more crawlers
All crawlers — combined across every endpoint
AhrefsBot 631
ClaudeBot 619
Chrome 147
MajesticBot 100
SemrushBot 97
MozBot 77
Other bot 75
PetalBot 72
Safari 61
OpenAI SearchBot 49
How this works: Every request to this site passes through a Cloudflare Worker that logs path, User-Agent, IP, and CF metadata to D1. A periodic snapshot job aggregates those rows into KV, which this page reads. Bot classification uses UA strings, CF reverse-DNS verification, and ASN fallbacks.
What each layer means: The discovery layer is what every crawler checks before indexing — robots.txt signals intent, sitemaps provide the URL map. The AI context layer is purpose-built for LLMs: llms.txt gives a plain-text orientation and page list, ai-catalog.json enumerates machine-readable endpoints, and entitymap.json declares structured entity relationships (people, concepts, organizations) following the EntityMap v1.0 spec. The raw content layer (.md requests) is agents bypassing HTML parsing entirely.
What each layer means: The discovery layer is what every crawler checks before indexing — robots.txt signals intent, sitemaps provide the URL map. The AI context layer is purpose-built for LLMs: llms.txt gives a plain-text orientation and page list, ai-catalog.json enumerates machine-readable endpoints, and entitymap.json declares structured entity relationships (people, concepts, organizations) following the EntityMap v1.0 spec. The raw content layer (.md requests) is agents bypassing HTML parsing entirely.