--- site: "Crawlytics.app" url: https://crawlytics.app/c/claude/ publisher: "Crawlytics" author: "Crawlytics Team" lastUpdated: 2026-08-25 pagesIncluded: 96 pagesTotal: 96 generatedAt: 2026-08-26T02:21:47.319Z --- # Crawlytics.app > Full markdown bundle of the top 96 pages on https://crawlytics.app/c/claude/, ranked by content quality, freshness, and importance. ## About this site **Publisher:** Crawlytics **Author:** Crawlytics Team **Last updated:** 2026-08-25 **Total pages indexed:** 96 ## Pages in this bundle 1. [AI Bot Tracking + llms.txt Generator + WebMCP — Crawlytics](https://crawlytics.app/c/claude/) 2. [AI Search Optimization: The AEO, GEO & LLMO Framework (2026)](https://crawlytics.app/c/claude/resources/ai-search-optimization) 3. [What Is llms.txt? The Complete Reference + Generator](https://crawlytics.app/c/claude/resources/llms-txt) 4. [Crawlytics Blog: AI Search, llms.txt, Bot Tracking, WebMCP](https://crawlytics.app/c/claude/blog) 5. [WebMCP Snippet: Let AI Agents Transact on Your Site](https://crawlytics.app/c/claude/features/webmcp-snippet) 6. [How to Manage AI Crawlers (Allow, Block, Monitor) — 2026 Guide](https://crawlytics.app/c/claude/resources/manage-ai-crawlers) 7. [Complete List of AI Crawler Bots: User-Agents + robots.txt (2026)](https://crawlytics.app/c/claude/resources/ai-bots-list) 8. [Crawlytics vs Google Analytics for AI Traffic](https://crawlytics.app/c/claude/blog/crawlytics-vs-google-analytics) 9. [Crawlytics vs Cloudflare Markdown for Agents: Honest Comparison](https://crawlytics.app/c/claude/blog/crawlytics-vs-cloudflare-markdown-for-agents) 10. [ChatGPT Traffic Shows as "Direct" in GA — Here Are 3 Fixes](https://crawlytics.app/c/claude/blog/chatgpt-direct-traffic-fix) 11. [How to Create an llms.txt File (and Test It) in 2026](https://crawlytics.app/c/claude/blog/what-is-llms-txt-guide) 12. [What Is WebMCP? AI Agent Actions Explained (2026)](https://crawlytics.app/c/claude/blog/webmcp-explained-ai-agent-actions) 13. [How to Track AI Citations (ChatGPT, Claude, Perplexity) 2026](https://crawlytics.app/c/claude/blog/how-to-track-ai-citations) 14. [Crawlytics vs Profound: AI Brand Visibility Tools Compared (2026)](https://crawlytics.app/c/claude/blog/crawlytics-vs-profound) 15. [What Schema Markup Still Matters in the AI Search Era](https://crawlytics.app/c/claude/blog/schema-markup-ai-search) 16. [AEO vs SEO vs GEO: Real Differences and Which to Invest in for 2026](https://crawlytics.app/c/claude/blog/aeo-vs-seo-vs-geo) 17. [How to Add llms.txt to WordPress (Plugin and Manual Methods)](https://crawlytics.app/c/claude/blog/wordpress-llms-txt-guide) 18. [How to Add llms.txt to Shopify (Step-by-Step Guide for 2026)](https://crawlytics.app/c/claude/blog/shopify-llms-txt-guide) 19. [AI Search and the SEO Funnel: New Conversion Paths for 2026](https://crawlytics.app/c/claude/blog/ai-search-changes-seo-funnel) 20. [How to Add WebMCP to Shopify Without Custom Code](https://crawlytics.app/c/claude/blog/shopify-webmcp-install) 21. [Default-Deny AI Crawlers: Why Reuters and Publishers Are Switching](https://crawlytics.app/c/claude/blog/default-deny-ai-crawlers) 22. [AI Agent Transactions: Chrome Auto-Browse Hits 200M+ Phones](https://crawlytics.app/c/claude/blog/ai-agent-transactions) 23. [Blended Retrieval: Gemini Fuses Web + Private Context](https://crawlytics.app/c/claude/blog/blended-retrieval) 24. [AI Share of Voice Is a Made-Up Number — Measure This Instead](https://crawlytics.app/c/claude/blog/ai-share-of-voice) 25. [Shopify AI Search Visibility: Five Fixes to Get Found](https://crawlytics.app/c/claude/blog/shopify-ai-search-visibility) 26. [Google AI Search Opt-Out Is Live — What Publishers Are Missing](https://crawlytics.app/c/claude/blog/google-ai-search-opt-out) 27. [What Is the Agentic Web? AI Agents Now Change Your Traffic](https://crawlytics.app/c/claude/blog/what-is-the-agentic-web) 28. [WebMCP Security: How to Deploy Agent Tools Safely](https://crawlytics.app/c/claude/blog/webmcp-security) 29. [Selling to AI Agents: Visa Cards Are Now Inside ChatGPT](https://crawlytics.app/c/claude/blog/ai-agent-commerce) 30. [Microsoft Web IQ: Why AI Agents Read Your Site Differently](https://crawlytics.app/c/claude/blog/microsoft-web-iq) 31. [WebKit Opposes WebMCP: Browser Fragmentation and What to Do](https://crawlytics.app/c/claude/blog/webkit-webmcp-browser-support) 32. [Google's llms.txt Guidance: What It Permits in 2026](https://crawlytics.app/c/claude/blog/google-llms-txt-guidance) 33. [Otterly vs Peec vs Crawlytics: Which AI Visibility Tool Wins?](https://crawlytics.app/c/claude/blog/otterly-vs-peec-vs-crawlytics) 34. [How to Get Your Products Into ChatGPT Shopping (2026)](https://crawlytics.app/c/claude/blog/chatgpt-shopping-product-feed-guide) 35. [llms.txt for AI Agents: Navigation, Not Discovery (2026)](https://crawlytics.app/c/claude/blog/llms-txt-agent-navigation) 36. [Perplexity Merchant Program: Complete Setup Guide (2026)](https://crawlytics.app/c/claude/blog/perplexity-merchant-program-guide) 37. [Sell in ChatGPT From a WooCommerce Store (No Shopify Required)](https://crawlytics.app/c/claude/blog/woocommerce-chatgpt-shopping) 38. [7 Best AI Visibility Tools Under $50/Month (2026)](https://crawlytics.app/c/claude/blog/best-ai-visibility-tools-under-50) 39. [Cloudflare Agent Readiness Score: What It Checks and Misses](https://crawlytics.app/c/claude/blog/cloudflare-agent-readiness-score) 40. [How to Add llms.txt to Squarespace (Yes, It's Possible)](https://crawlytics.app/c/claude/blog/squarespace-llms-txt-guide) 41. [Fix What ChatGPT Says About Your Brand — Step by Step](https://crawlytics.app/c/claude/blog/fix-what-chatgpt-says-about-your-brand) 42. [Retrieval vs Citation in AI Search: What's the Difference](https://crawlytics.app/c/claude/blog/retrieval-vs-citation) 43. [Agentic Checkout Readiness: The 2026 Find/Read/Buy Checklist](https://crawlytics.app/c/claude/blog/agentic-checkout-readiness-checklist) 44. [97% of llms.txt Files Got No AI Requests. Here's the Full Story.](https://crawlytics.app/c/claude/blog/llms-txt-no-traffic-data) 45. [ChatGPT Agent Blocked From Your Site? 6 Causes and Fixes](https://crawlytics.app/c/claude/blog/chatgpt-agent-cant-access-website) 46. [AI Bot Traffic Cost: Which Crawlers Are Worth It?](https://crawlytics.app/c/claude/blog/ai-bot-traffic-cost) 47. [AI Search Visibility Audit: Check If You're AI-Ready](https://crawlytics.app/c/claude/blog/ai-search-visibility-audit) 48. [Best AI Bot Tracking Tools for 2026](https://crawlytics.app/c/claude/blog/best-ai-bot-tracking-tools) 49. [Crawlytics vs Ahrefs Bot Analytics (2026)](https://crawlytics.app/c/claude/blog/crawlytics-vs-ahrefs-bot-analytics) 50. [Do You Need Ahrefs or Semrush for AI Visibility?](https://crawlytics.app/c/claude/blog/do-you-need-ahrefs-semrush-for-ai-visibility) 51. [Best AI Brand Monitoring Tools for 2026](https://crawlytics.app/c/claude/blog/best-ai-brand-monitoring-tools) 52. [How to Track Which AI Bots Crawl Your Site (2026)](https://crawlytics.app/c/claude/blog/how-to-track-ai-bots-crawling-your-site) 53. [Is llms.txt Worth It? What Skeptics Get Wrong (2026)](https://crawlytics.app/c/claude/blog/is-llms-txt-worth-it) 54. [How Much Do AI Visibility Tools Cost? (2026)](https://crawlytics.app/c/claude/blog/ai-visibility-tracking-cost) 55. [7 Best Profound Alternatives for 2026](https://crawlytics.app/c/claude/blog/best-profound-alternatives) 56. [Crawlytics vs Scrunch AI: Honest Comparison (2026)](https://crawlytics.app/c/claude/blog/crawlytics-vs-scrunch-ai) 57. [Best AI Visibility Tools for Agencies (2026)](https://crawlytics.app/c/claude/blog/ai-visibility-tools-for-agencies) 58. [Best AI Readiness Tools for 2026](https://crawlytics.app/c/claude/blog/best-agent-readiness-tools) 59. [Best ChatGPT Brand Monitoring Tools (2026)](https://crawlytics.app/c/claude/blog/best-chatgpt-brand-monitoring-tools) 60. [Best GEO Tools for 2026 (Generative Engine Optimization)](https://crawlytics.app/c/claude/blog/best-geo-tools) 61. [Best llms.txt Generators (2026): 6 Tools Compared](https://crawlytics.app/c/claude/blog/best-llms-txt-generators) 62. [How to Prove GEO / AI-SEO ROI to Clients (2026)](https://crawlytics.app/c/claude/blog/prove-geo-roi-to-clients) 63. [White-Label AI Search Reports for Clients (2026)](https://crawlytics.app/c/claude/blog/white-label-ai-search-reports) 64. [Crawlytics vs Knowatoa: AI Visibility Compared (2026)](https://crawlytics.app/c/claude/blog/crawlytics-vs-knowatoa) 65. [Safari Just Gave AI Agents a Browser Window: What Site Owners Should Know](https://crawlytics.app/c/claude/blog/safari-mcp-server-ai-agents) 66. [The Silent Funnel: The AI Agent Traffic You Can't See](https://crawlytics.app/c/claude/blog/silent-funnel-ai-agent-traffic) 67. [AI Agents Can't Read Your Pricing Page (And It's Costing You Deals)](https://crawlytics.app/c/claude/blog/ai-agent-pricing-page) 68. [Why Specific Content Gets Cited by AI (And How to Prove It With Your Own Data)](https://crawlytics.app/c/claude/blog/specific-content-ai-citations) 69. [ai-catalog.json: Put Your Site on the AI Agent Discovery Map](https://crawlytics.app/c/claude/blog/ai-catalog-json) 70. [Cloudflare's New AI Bot Rules Go Live September 15 — What Site Owners Must Do Now](https://crawlytics.app/c/claude/blog/cloudflare-ai-bot-rules-september-2026) 71. [Cloudflare's New Bot Detection Is Smarter Than Ever. Is It Blocking AI Agents You Want?](https://crawlytics.app/c/claude/blog/cloudflare-precursor-bot-detection) 72. [The Crawl-to-Referral Ratio: The AI-Era Metric Every Site Owner Needs](https://crawlytics.app/c/claude/blog/crawl-to-referral-ratio) 73. [Is Your Site AI-Agent Ready? Free Scorers Are Only a Snapshot](https://crawlytics.app/c/claude/blog/ai-agent-ready-website-audit) 74. [AI Agents Will Search More Than Humans This Year: Is Your Site Ready?](https://crawlytics.app/c/claude/blog/ai-agent-search-traffic-exa) 75. [AI Bot Traffic 2026: What the Cloudflare Report Means](https://crawlytics.app/c/claude/blog/ai-bot-traffic-2026) 76. [Cloudflare Just Made It Real: How to Charge AI Agents for Your Content with x402](https://crawlytics.app/c/claude/blog/charge-ai-agents-x402) 77. [AI Bot Tracking: Detect GPTBot, ClaudeBot, PerplexityBot](https://crawlytics.app/c/claude/features/llm-tracking) 78. [AI Referral Tracking: ChatGPT, Claude, Perplexity Clicks](https://crawlytics.app/c/claude/features/ai-attribution) 79. [Block GPTBot or Allow It? The 2026 AI Crawler Decision Guide](https://crawlytics.app/c/claude/blog/block-gptbot-decision-guide) 80. [How to Get Cited by ChatGPT: A Practical Playbook for 2026](https://crawlytics.app/c/claude/blog/how-to-get-cited-by-chatgpt) 81. [Optimize Blog Posts for AI Citations: The 8-Edit Checklist](https://crawlytics.app/c/claude/blog/optimize-blog-posts-for-ai-citations) 82. [Google's Open Knowledge Format: The Next llms.txt?](https://crawlytics.app/c/claude/blog/open-knowledge-format) 83. [Agentic Commerce for SaaS: When AI Agents Buy Your Plan](https://crawlytics.app/c/claude/blog/agentic-commerce-for-saas) 84. [WebMCP Agent Support: Which AI Agents Invoke Tools (2026)](https://crawlytics.app/c/claude/resources/webmcp-agent-support) 85. [Crawlytics Privacy Policy: What We Store and Don't](https://crawlytics.app/c/claude/privacy) 86. [llms.txt Generator + Per-Page Markdown for AI Bots](https://crawlytics.app/c/claude/features/llms-txt-generator) 87. [Enterprise Security: Encrypted, Private, GDPR-Ready](https://crawlytics.app/c/claude/features/security) 88. [Multi-Site AI Bot Tracking + Portfolio Dashboard](https://crawlytics.app/c/claude/features/multi-site) 89. [Crawlytics Terms of Service: Plans, Billing, Acceptable Use](https://crawlytics.app/c/claude/terms) 90. [Crawlytics Pricing: AI Bot Tracking + llms.txt from $29.99/mo](https://crawlytics.app/c/claude/pricing) 91. [Crawlytics Resources: AI Search Guides, Bot Lists, Grader](https://crawlytics.app/c/claude/resources) 92. [What is Crawlytics? 60-Second Explainer Video](https://crawlytics.app/c/claude/explainer) 93. [AI Bot Traffic Dashboard: GPTBot & ClaudeBot Analytics](https://crawlytics.app/c/claude/features/analytics-dashboard) 94. [Crawlytics Features: AI Bot Tracking, llms.txt & WebMCP](https://crawlytics.app/c/claude/features) 95. [AI Bot Analytics Demo: Live Sample Dashboard](https://crawlytics.app/c/claude/demo) 96. [Agent-Ready Grader: Free AI-Readiness Scanner for Any Website](https://crawlytics.app/c/claude/agent-ready) --- title: "AI Bot Tracking + llms.txt Generator + WebMCP — Crawlytics" type: [Organization, FAQPage, WebSite] canonical: https://crawlytics.app/c/claude/ category: homepage wordCount: 739 readingTime: 4 min crawledAt: 2026-08-19 13:05:24 lastVerified: 2026-08-25 13:09:57 site: https://crawlytics.app/c/claude/ --- # AI Bot Tracking + llms.txt Generator + WebMCP — Crawlytics ## Key facts - When ChatGPT or Google's agent lands on your site, it arrives with a job: book the appointment, buy the product, request the quote. - No code changes to your site, no DNS, no reverse proxy. - no reverse proxy · no DNS · one `. That's it. The loader registers your configured tools with navigator.modelContext on browsers that support WebMCP, and silently no-ops on browsers that don't. No CMS plugin, no build step. ### Which AI agents support WebMCP? WebMCP is the draft web spec exposing navigator.modelContext. Currently supported in Chrome 146+ Canary (which means Gemini Live, in-browser Claude artifacts, ChatGPT browser-mode, and Perplexity's Comet browser can invoke tools). Safari and Firefox have not shipped support yet. Crawlytics feature-detects before doing anything — zero risk to non-supporting browsers. ### Does WebMCP work in Safari? Not yet. WebMCP is a draft web spec and Safari has not announced support. The Crawlytics snippet feature-detects navigator.modelContext before doing anything, so Safari visitors see no behavior change. The conversion-attribution half of the snippet does run in every browser (it watches Stripe's ?session_id= on redirect-back), so you still get attribution from Safari-routed purchases. ### What is WebMCP? WebMCP is a draft web spec — currently in Chrome 146+ Canary preview — that exposes navigator.modelContext, letting a page register tools an in-browser AI agent can invoke. The snippet is your one-step way to register tools without writing browser-API code yourself. ### Does it require Chrome 146 Canary to work? The agent-action half does. On every other browser the snippet silently no-ops — it feature-detects navigator.modelContext before doing anything, so there is zero risk to real visitors. The conversion-attribution half runs in every browser (it just watches the success URL on Stripe redirect-back). ### Do I need to change my checkout? No. Conversion attribution works by detecting Stripe's ?session_id=cs_… on your success page — same page your customers already land on. Zero customer setup, no webhook, no API key. For cryptographically verified amounts you can optionally add a Stripe webhook later. ### What about CMS plugins? There aren't any and there won't be. The snippet is one script tag that drops into any HTTPS page — Shopify, Wix, Squarespace, custom Next.js, WordPress. No CMS-specific code anywhere. ### Where do API secrets live? On your server, never in the DB or browser. The snippet config stores the NAME of an env var (e.g. SITE_42_SHOPIFY_TOKEN, where 42 is the site id); Crawlytics resolves the value at invocation time via process.env. Names must match the per-site SITE__ pattern so a user can't name a server-internal env var as their "auth ref" and exfiltrate the value. The dashboard shows a green/red dot so you can confirm the var is wired without ever seeing the value. ### Can agents enter card details? No. PCI compliance and Stripe's sandboxed iframes make this impossible — by design. The agent collects intent, your endpoint creates a Stripe Checkout session, the agent hands the URL to the user. The user completes payment on Stripe's hosted page. No card data ever touches Crawlytics or the snippet. --- title: "How to Manage AI Crawlers (Allow, Block, Monitor) — 2026 Guide" type: [Organization, TechArticle, BreadcrumbList, FAQPage, WebSite] author: Crawlytics publisher: Crawlytics datePublished: 2026-06-03 dateModified: 2026-06-03 canonical: https://crawlytics.app/c/claude/resources/manage-ai-crawlers category: docs wordCount: 1518 readingTime: 8 min crawledAt: 2026-06-21 16:40:27 lastVerified: 2026-08-25 13:09:28 site: https://crawlytics.app/c/claude/ --- # How to Manage AI Crawlers (Allow, Block, Monitor) — 2026 Guide ## Summary A practical guide to managing AI crawlers on your site: when to block, when to allow, robots.txt patterns, CDN bot rules, and how to measure the impact. ## Key facts - Before you paste anything into robots. - For each AI bot, ask three questions: - An increasingly popular third path: don't block AI bots, but serve them clean markdown instead of your full HTML. - Three things to check: - If you want a "good enough" starting position without overthinking it: ## Start with the framework, not the config Before you paste anything into robots.txt, decide what you're trying to achieve. There are really only four positions a site can take on AI crawlers: 1. **Allow everything.** You want to be cited by every AI assistant. Default for most marketing sites, SaaS, content sites, ecommerce. 2. **Allow but track.** You allow AI traffic but want to know who's reading what so you can optimize. Most sites belong here once they get curious. 3. **Allow some, block others.** Allow the ones that drive measurable referral traffic (Perplexity, ChatGPT search), block the ones that just train models without sending visitors (CCBot, anthropic-ai). Selective. 4. **Block everything.** You're behind a paywall, your content is proprietary, or you're philosophically opposed to AI training on your work. Rare in commercial contexts; common for publishers fighting copyright issues. Most teams default to position 1 without thinking about it. The interesting question is whether _your_ position should be position 2, 3, or 4 — and you can't answer that without data on what AI crawlers are actually doing on your site. ## The "should I allow this bot?" decision For each AI bot, ask three questions: 1. **Does it drive referral traffic?** Perplexity and ChatGPT search produce real human visits to cited sites. Pure training crawlers (CCBot, Bytespider for ByteDance's internal use, Applebot-Extended) don't drive direct traffic — they feed a model whose output may or may not include your site. 2. **Does it serve your customers?** If your audience uses Claude or Gemini, having those models trained on your content means your customers get accurate answers about your product. Blocking means accuracy drops. 3. **Is it scraping you in a way you'd consider harmful?** Some publishers care about copyright; some don't. Some care about competitive intelligence (e.g., pricing pages being scraped by competitors masquerading as AI bots); some don't. Two answers of "yes" to questions 1 or 2 generally means allow. Two answers of "yes" to question 3 means block. Mixed answers mean monitor for 30-60 days first. ## robots.txt — the polite signal robots.txt is the gentleman's agreement of the web. Well-behaved bots honor it; bad actors ignore it. All the major AI companies (OpenAI, Anthropic, Google, Meta, Apple) honor robots.txt for their named bots — they have legal teams who care. ### Block all AI bots ``` User-agent: GPTBot Disallow: / User-agent: ChatGPT-User Disallow: / User-agent: OAI-SearchBot Disallow: / User-agent: ClaudeBot Disallow: / User-agent: claude-web Disallow: / User-agent: anthropic-ai Disallow: / User-agent: PerplexityBot Disallow: / User-agent: Perplexity-User Disallow: / User-agent: Google-Extended Disallow: / User-agent: Bytespider Disallow: / User-agent: CCBot Disallow: / User-agent: Meta-ExternalAgent Disallow: / User-agent: Amazonbot Disallow: / User-agent: Applebot-Extended Disallow: / User-agent: GrokBot Disallow: / User-agent: cohere-ai Disallow: / ``` This blocks the major LLM crawlers but allows traditional search bots (Googlebot, Bingbot) — you don't lose SEO. Note: `Google-Extended` is specifically Google's AI opt-out token; blocking it removes you from Gemini training and AI Overviews _without_ removing you from Google Search. ### Block training but allow live-fetch (let users get fresh answers about your site) ``` User-agent: GPTBot Disallow: / User-agent: ClaudeBot Disallow: / User-agent: Google-Extended Disallow: / User-agent: Bytespider Disallow: / User-agent: CCBot Disallow: / User-agent: Applebot-Extended Disallow: / # Allow live-fetch agents — these fire when a user asks the AI about your page # User-agent: ChatGPT-User (not listed = allowed) # User-agent: Perplexity-User (not listed = allowed) # User-agent: claude-web (not listed = allowed) ``` This is a defensible middle ground: your content isn't used to train new models, but a user asking ChatGPT "what does example.com say about X?" still gets a fresh fetch of your page. ### Block specific paths only ``` User-agent: GPTBot Disallow: /pricing Disallow: /customers Disallow: /case-studies User-agent: ClaudeBot Disallow: /pricing Disallow: /customers Disallow: /case-studies ``` Useful when you want AI assistants to recommend your product (so allow the homepage, features, docs) but you don't want them quoting your exact pricing or customer logos out of context. AI assistants regularly misquote prices because they trained on outdated cache; blocking `/pricing` from the training crawlers forces the model to either skip pricing or fetch it live. ## CDN bot rules — the enforced signal robots.txt is a request. CDN bot rules are enforcement. If you have Cloudflare, Fastly, or Vercel in front of your site, you can return 403/429 to specific bot fingerprints and they have no choice. ### Cloudflare Cloudflare's Bot Management tier lets you write rules in the Web Application Firewall. A typical block looks like: ``` (cf.client.bot) and (http.user_agent contains "GPTBot") ``` Set the action to **Block** (or **Challenge** if you want to be less aggressive). Cloudflare also ships a free "AI Scrapers and Crawlers" managed rule you can toggle in one click, which covers most of the bots in this list. Cloudflare's recently-shipped [Content Signals](https://crawlytics.app/c/claude/blog/crawlytics-vs-cloudflare-markdown-for-agents) mechanism is a more nuanced version of this — you declare whether each path may be used for training, search, or inference, and crawlers self-comply. Worth enabling alongside hard blocks. ### Vercel Bot Management (Edge Network) ``` // middleware.ts export function middleware(req: Request) { const ua = req.headers.get('user-agent') || ''; const aiBots = /GPTBot|ClaudeBot|PerplexityBot|Bytespider|CCBot/i; if (aiBots.test(ua)) { return new Response('Forbidden', { status: 403 }); } } ``` ### nginx ``` map $http_user_agent $is_ai_bot { default 0; ~*GPTBot 1; ~*ClaudeBot 1; ~*PerplexityBot 1; ~*Bytespider 1; ~*CCBot 1; } server { if ($is_ai_bot) { return 403; } } ``` ## Allow + serve markdown (the "agent-friendly" approach) An increasingly popular third path: don't block AI bots, but serve them clean markdown instead of your full HTML. The benefits: - You stay cited (good for visibility) - You save bandwidth (markdown is ~1/5 the size of HTML) - You get better AI summaries because the bot reads structured content, not nav/footer noise - You can inject per-LLM UTM tags into outbound links for [attribution recovery](https://crawlytics.app/c/claude/blog/chatgpt-direct-traffic-fix) Two ways to do this: 1. **Stable URLs:** publish `/llms.txt`, `/llms-full.txt`, and `/md/` markdown files at predictable URLs. AI bots that know the convention fetch them directly. This is what [Crawlytics generates](https://crawlytics.app/c/claude/features/llms-txt-generator) for you. 2. **Content negotiation:** when an AI bot sends `Accept: text/markdown`, return markdown instead of HTML for the canonical URL. This is what [Cloudflare's Markdown for Agents](https://blog.cloudflare.com/markdown-for-agents/) ships. Both approaches work; the first reaches more bots (most clients don't send `Accept: text/markdown` yet), the second is lower-latency. [Full comparison here](https://crawlytics.app/c/claude/blog/crawlytics-vs-cloudflare-markdown-for-agents). ## Measuring whether your config is working Three things to check: ### 1\. Are blocked bots actually blocked? Run this from a test environment: ``` curl -A "GPTBot" https://yoursite.com/ ``` If your block rule fires you should see 403. If you see 200, your robots.txt is being honored but your CDN isn't enforcing — fine if that's intentional, a problem if you meant to hard-block. ### 2\. Are allowed bots still visiting? Grep your server logs for the User-Agents you allowed: ``` grep -iE 'PerplexityBot|ChatGPT-User|claude-web' /var/log/nginx/access.log | tail -50 ``` If the count is climbing over time, your allow list is working as intended. If it dropped to zero after a config change, you accidentally blocked something. ### 3\. Are you actually getting referral traffic from AI assistants? This is the bottom-line question. Blocking and allowing only matter if they translate to (or away from) human visits. Two ways to measure: - **Free:** grep your logs for Referer values matching `chat.openai.com`, `perplexity.ai`, `claude.ai`. You'll miss most in-app browser sessions (they strip Referer) but the desktop traffic shows up. - **Full coverage:** install [Crawlytics' attribution layer](https://crawlytics.app/c/claude/features/ai-attribution), which injects per-LLM UTM tags into the AI-Optimized HTML bots fetch — so when ChatGPT cites your URL, the UTM travels with it and your analytics see `chatgpt` as the source even when Referer is stripped. ## What about bots that ignore robots.txt? They exist. Scrapers masquerading as legitimate AI bots, abandoned crawlers running on autopilot, and a handful of named bots from less-reputable companies. For these: - **Rate-limit by IP** at the CDN layer. AI training crawlers often run from concentrated IP ranges. - **Use Cloudflare's bot fight mode** (or Fastly's equivalent) — they detect headless browsers, mismatched UA/fingerprint pairs, and known-bad IPs without needing custom rules. - **Honeypot pages** — pages disallowed in robots.txt that legitimate bots respect. Anything hitting them is by definition ignoring your robots, so you can ban the IP immediately. ## A reasonable default for most sites If you want a "good enough" starting position without overthinking it: 1. Allow all AI bots in robots.txt (don't block anything for the first 30 days) 2. Install bot tracking — [Crawlytics](https://crawlytics.app/c/claude/features/llm-tracking) or grep your own logs 3. Observe for 30 days: which bots are visiting, how much volume, which pages they prefer, whether the cited-by-AI traffic appears in your analytics 4. Make blocking decisions based on the data — block the bots that consume bandwidth without sending visits, allow the ones that drive measurable referrals 5. Generate `/llms.txt` to make the allowed bots' job easier and get better citations This is more work than "block everything" but it's also the only way to make a decision that aligns with your actual business outcomes instead of a vibes-based reaction. ## Related ## Frequently Asked Questions ### How do I block GPTBot from crawling my website? Add the following to your robots.txt: User-agent: GPTBot then Disallow: /. Repeat for ChatGPT-User and OAI-SearchBot if you want to block live-fetch and search-index bots too. For hard enforcement (not just polite request), add a CDN bot rule in Cloudflare, Fastly, or Vercel that returns 403 to that User-Agent. ### How do I block all AI crawlers at once? List each major bot explicitly in robots.txt: GPTBot, ChatGPT-User, OAI-SearchBot, ClaudeBot, claude-web, anthropic-ai, PerplexityBot, Perplexity-User, Google-Extended, Bytespider, CCBot, Meta-ExternalAgent, Amazonbot, Applebot-Extended, GrokBot, cohere-ai. Note Google-Extended is Google's AI opt-out token, blocking it removes you from Gemini training and AI Overviews without affecting your Google Search ranking. ### Should I block AI crawlers or allow them? Allow them if you want to be cited by AI search and AI assistants because blocking the training crawler means your content is absent from the model's knowledge. Block them if your content is paywalled or proprietary. A common middle ground: block pure training crawlers (CCBot, Bytespider, Applebot-Extended, Google-Extended), allow live-fetch agents (ChatGPT-User, Perplexity-User, claude-web) so user-initiated questions about your site still get fresh content. ### Does Google-Extended affect my Google Search ranking? No. Google-Extended is a separate token Google introduced specifically as an AI opt-out signal. Blocking Google-Extended in robots.txt removes you from Gemini training and Google AI Overviews, but Googlebot and Googlebot-News still crawl normally and your traditional Google Search ranking is unaffected. ### Do AI crawlers honor robots.txt? The major ones do. OpenAI, Anthropic, Google, Meta, Apple, and Perplexity all honor robots.txt for their named bots because they have legal teams that care. A handful of scrapers and less reputable bots ignore robots.txt entirely, for those you need CDN bot rules, rate limiting, or honeypot pages that ban any IP that fetches them. --- title: "Complete List of AI Crawler Bots: User-Agents + robots.txt (2026)" type: [Organization, TechArticle, BreadcrumbList, FAQPage, WebSite] author: Crawlytics publisher: Crawlytics datePublished: 2026-06-03 dateModified: 2026-06-03 canonical: https://crawlytics.app/c/claude/resources/ai-bots-list category: docs wordCount: 1245 readingTime: 6 min crawledAt: 2026-08-25 13:10:41 lastVerified: 2026-08-25 13:10:41 site: https://crawlytics.app/c/claude/ --- # Complete List of AI Crawler Bots: User-Agents + robots.txt (2026) ## Summary Every major AI crawler: GPTBot, ClaudeBot, PerplexityBot, Google-Extended, Bytespider, and 20 more. User-Agent patterns, purpose, robots.txt directives. ## Key facts - AI crawlers fall into three categories: - Depends on your business. - We've intentionally left out: - If you have raw access logs (nginx, Apache, Vercel, Cloudflare), this command will surface AI bot traffic for the bots in the list above: - New AI bots show up roughly monthly. If you've been getting unexpected traffic from User-Agents you don't recognize, this is probably the table you wanted. We maintain this list by watching real ingest traffic across the Crawlytics customer base — when a new AI crawler shows up in the wild, we add the signature here and in the production classifier. 25 bots across 19 companies as of June 2026. ## Why these bots exist AI crawlers fall into three categories: 1. **Training crawlers** — fetch your content to use in model training. These visit periodically (weekly to monthly), don't fire JavaScript, and won't show up in your front-end analytics. Examples: GPTBot, ClaudeBot, Bytespider, Applebot-Extended. 2. **Live-fetch agents** — fire when a user asks the AI a question that requires fetching a specific URL right now. Lower volume but real-time. Examples: ChatGPT-User, Perplexity-User, claude-web. 3. **Search-index crawlers** — feed AI search products (SearchGPT, You.com, DuckAssist, Kagi). Behave more like traditional search crawlers — frequent, broad, indexed for retrieval. Examples: OAI-SearchBot, PerplexityBot, YouBot. Most production AI assistants use multiple bots from this list in combination — training plus live-fetch plus index. Blocking one but not the others usually doesn't get you the result you want. ## Should you block AI crawlers? Depends on your business. The short version: - **If you want to be cited by ChatGPT, Claude, Perplexity, and AI search:** allow them. Blocking the training crawler means your content is missing from the model's knowledge; blocking the live-fetch agent means user-initiated queries about your page can't pull fresh content. - **If your content is paywalled or proprietary and being scraped without compensation:** block them. Use robots.txt for compliant bots and CDN bot rules (Cloudflare, Fastly, Vercel Bot Manager) for the rest. - **If you're not sure:** install [Crawlytics](https://crawlytics.app/c/claude/) first to see what they're actually doing on your site. Then decide based on data instead of vibes. For a deeper walkthrough of allow/block strategy, see [how to manage AI crawlers](https://crawlytics.app/c/claude/resources/manage-ai-crawlers). ## The full bot table ### OpenAI | Bot name | Purpose | robots.txt | | --- | --- | --- | | GPTBotofficial docs | Training crawl | User-agent: GPTBot Disallow: / | | ChatGPT-Userofficial docs | Live user-initiated fetch | User-agent: ChatGPT-User Disallow: / | | OAI-SearchBotofficial docs | SearchGPT index | User-agent: OAI-SearchBot Disallow: / | ### Anthropic | Bot name | Purpose | robots.txt | | --- | --- | --- | | ClaudeBotofficial docs | Training crawl | User-agent: ClaudeBot Disallow: / | | claude-webofficial docs | Live user-initiated fetch | User-agent: claude-web Disallow: / | | anthropic-aiofficial docs | Legacy / general | User-agent: anthropic-ai Disallow: / | ### Perplexity | Bot name | Purpose | robots.txt | | --- | --- | --- | | PerplexityBotofficial docs | Index for Perplexity answers | User-agent: PerplexityBot Disallow: / | | Perplexity-Userofficial docs | Live user-initiated fetch | User-agent: Perplexity-User Disallow: / | ### Google | Bot name | Purpose | robots.txt | | --- | --- | --- | | Google-Extendedofficial docs | Gemini training + AI Overviews opt-out signal | User-agent: Google-Extended Disallow: / | ### ByteDance | Bot name | Purpose | robots.txt | | --- | --- | --- | | Bytespider | Training crawl for Doubao + TikTok AI | User-agent: Bytespider Disallow: / | ### Common Crawl | Bot name | Purpose | robots.txt | | --- | --- | --- | | CCBotofficial docs | Open crawl corpus used by many AI labs | User-agent: CCBot Disallow: / | ### Meta | Bot name | Purpose | robots.txt | | --- | --- | --- | | Meta-ExternalAgent | Meta AI / Llama training | User-agent: Meta-ExternalAgent Disallow: / | | FacebookBotofficial docs | Public sharing previews — overlaps AI use | User-agent: FacebookBot Disallow: / | ### Amazon | Bot name | Purpose | robots.txt | | --- | --- | --- | | Amazonbotofficial docs | Alexa training + Amazon AI | User-agent: Amazonbot Disallow: / | ### Apple | Bot name | Purpose | robots.txt | | --- | --- | --- | | Applebot-Extendedofficial docs | Apple Intelligence training | User-agent: Applebot-Extended Disallow: / | ### Microsoft | Bot name | Purpose | robots.txt | | --- | --- | --- | | CopilotBot | Microsoft 365 Copilot crawl | User-agent: CopilotBot Disallow: / | ### xAI | Bot name | Purpose | robots.txt | | --- | --- | --- | | GrokBot | Grok training | User-agent: GrokBot Disallow: / | ### Mistral | Bot name | Purpose | robots.txt | | --- | --- | --- | | MistralAI-User | Le Chat live fetch | User-agent: MistralAI-User Disallow: / | ### Cohere | Bot name | Purpose | robots.txt | | --- | --- | --- | | cohere-ai | Cohere training | User-agent: cohere-ai Disallow: / | ### You.com | Bot name | Purpose | robots.txt | | --- | --- | --- | | YouBot | You.com search + AI | User-agent: YouBot Disallow: / | ### Phind | Bot name | Purpose | robots.txt | | --- | --- | --- | | PhindBot | Phind developer search | User-agent: PhindBot Disallow: / | ### DuckDuckGo | Bot name | Purpose | robots.txt | | --- | --- | --- | | DuckAssistBot | DuckAssist (AI answers) | User-agent: DuckAssistBot Disallow: / | ### Kagi | Bot name | Purpose | robots.txt | | --- | --- | --- | | KagiBot | Kagi search + AI features | User-agent: KagiBot Disallow: / | ### Diffbot | Bot name | Purpose | robots.txt | | --- | --- | --- | | Diffbotofficial docs | Knowledge graph extraction | User-agent: Diffbot Disallow: / | ### AI2 | Bot name | Purpose | robots.txt | | --- | --- | --- | | ai2bot | Allen Institute for AI research | User-agent: ai2bot Disallow: / | ## What's not in this list We've intentionally left out: - **Googlebot, Bingbot, traditional search crawlers.** They predate the AI category and are well-documented elsewhere. Blocking them is almost always a bad idea regardless of your AI stance. - **Generic scrapers** with no clear AI affiliation (e.g., random Python `requests` User-Agents). We classify those as "unknown" traffic, not AI. - **Image-only crawlers** (ImageSift, etc.) unless they participate in AI training, which most don't currently. - **RSS/feed readers** and uptime monitors that some sites mistake for AI traffic. ## Detecting these bots in your own logs If you have raw access logs (nginx, Apache, Vercel, Cloudflare), this command will surface AI bot traffic for the bots in the list above: ``` grep -iE 'GPTBot|ChatGPT-User|OAI-SearchBot|ClaudeBot|claude-web|anthropic-ai|PerplexityBot|Perplexity-User|Google-Extended|Bytespider|CCBot|Meta-ExternalAgent|FacebookBot|Amazonbot|Applebot-Extended|CopilotBot|GrokBot|MistralAI-User|cohere-ai|YouBot|PhindBot|DuckAssistBot|KagiBot|Diffbot|ai2bot' /var/log/nginx/access.log | wc -l ``` That gives you a count. Drop the `| wc -l` for the full list of requests. For an actual dashboard with per-bot per-page breakdowns and historical trends, [install Crawlytics](https://crawlytics.app/c/claude/features/llm-tracking) — it does this in real time across 19 providers. ## This list will get out of date New AI bots show up roughly monthly. We update this page on a similar cadence — the "Last updated" date at the top is the source of truth. If you're looking at this 6+ months past that date, expect there to be additions we haven't shipped yet. If you spot an AI crawler in your logs that's not on this list, [email us](https://crawlytics.app/c/claude/cdn-cgi/l/email-protection#076f626b6b6847647566706b7e736e647429667777387472656d6264733a496270223537464e223537656873223537746e60696673727562) — we add new bot patterns within a few days of seeing them in the wild. ## Related ## Frequently Asked Questions ### What is GPTBot? GPTBot is OpenAI's training crawler. It visits public websites a few times per week to collect content for training future versions of ChatGPT. It does not execute JavaScript, does not show up in Google Analytics, and respects robots.txt. To block it, add User-agent: GPTBot then Disallow: / to your robots.txt. ### What is the difference between GPTBot and ChatGPT-User? GPTBot is OpenAI's training crawler that runs on a schedule. ChatGPT-User is the live-fetch agent that fires only when a real user asks ChatGPT to read a specific page right now. OAI-SearchBot is a third bot, OpenAI's SearchGPT index crawler. Each can be allowed or blocked independently in robots.txt. ### How do I see which AI bots are crawling my site? Three options: (1) grep your raw server access logs for known User-Agent patterns (GPTBot, ClaudeBot, PerplexityBot, Bytespider, CCBot, etc.); (2) check your CDN dashboard if you use Cloudflare or Fastly; (3) install a dedicated tracker like Crawlytics, which classifies 25+ AI crawlers in real time and shows per-page per-bot crawl frequency. ### Should I block AI bots from crawling my site? Depends on your goal. Block them if your content is paywalled, proprietary, or being scraped without compensation. Allow them if you want to be cited by ChatGPT, Claude, Perplexity, and AI search results, because blocking the training crawler means your content is absent from the model's knowledge. A common middle ground: block pure training crawlers like CCBot and Bytespider, allow live-fetch agents like ChatGPT-User and Perplexity-User. ### How often do AI crawlers visit a website? Varies widely. Training crawlers like GPTBot and ClaudeBot typically hit a site a few times per week per page. Live-fetch agents like ChatGPT-User and Perplexity-User only fire when a real user asks a question that requires reading that specific URL. High-traffic pages or pages with frequent updates get crawled more often. --- title: "Crawlytics vs Google Analytics for AI Traffic" type: [Organization, Article, BreadcrumbList, WebSite, FAQPage] author: Crawlytics Team publisher: Crawlytics datePublished: 2026-06-03 dateModified: 2026-06-03 canonical: https://crawlytics.app/c/claude/blog/crawlytics-vs-google-analytics category: blog wordCount: 1283 readingTime: 6 min crawledAt: 2026-06-21 16:40:12 lastVerified: 2026-08-25 13:09:19 site: https://crawlytics.app/c/claude/ --- # Crawlytics vs Google Analytics for AI Traffic ## Summary Google Analytics filters out bot traffic and can't see what AI crawlers do on your site. Here is where Crawlytics complements GA, and where each tool wins. ## Key facts - Where it gets interesting is the overlap zone — AI assistants driving real human visits to your site. - Google Analytics 4 is excellent at: - The cleanest setup uses both: - Cloudflare Radar shows aggregated, industry-wide AI bot traffic distribution. - You probably need both if any of these are true: ## The short answer **You should run both.** Google Analytics measures human behavior; Crawlytics measures AI behavior. They answer different questions and they complement, not compete. Where it gets interesting is the overlap zone — AI assistants driving real human visits to your site. GA misses most of this (it ends up in "(direct)"). Crawlytics catches it. If you only run GA, you have a large and growing blind spot in your acquisition data. ## What GA does well Google Analytics 4 is excellent at: - Tracking human sessions — page views, events, conversions, funnels - Attribution across paid, organic, social, email, and referral channels - Audience segmentation and demographics - Cross-device user journeys (via Google's signals) - Integration with Google Ads, Search Console, BigQuery, Looker Studio If you're optimizing your funnel for human conversion, GA is the right tool. None of what's below is a criticism of that use case. ## Where GA falls short on AI traffic Three concrete gaps: ### 1\. GA filters out bot traffic by default GA4 has a setting called "exclude all hits from known bots and spiders" and it's on by default. Even if you turn it off, GA's JavaScript tag only fires in real browsers. AI crawlers (GPTBot, ClaudeBot, PerplexityBot, etc.) don't execute JavaScript. The tag never loads. The bot visit never lands in GA. Net result: **Google Analytics shows you ~0% of your AI crawler traffic**, because the data pipeline can't see it. The full detection playbook — which UAs to grep for, the live-fetch vs training-crawler distinction, and benchmarks for what a healthy bot-to-human ratio looks like — is in our piece on [how to track AI citations](https://crawlytics.app/c/claude/blog/how-to-track-ai-citations). ### 2\. AI assistant referrals get mis-attributed to "Direct" Even when an AI assistant drives a real human visitor to your site, GA usually fails to attribute it correctly. Why: ChatGPT, Claude, and Perplexity's mobile and in-app browsers strip the Referer header on outbound clicks. GA sees a visit with no source and buckets it as `(direct) / (none)`. This is the most under-discussed measurement problem in marketing right now. A typical mid-size SaaS site might have 5-15% of its "direct traffic" actually originating from AI assistants. [Full write-up here](https://crawlytics.app/c/claude/blog/chatgpt-direct-traffic-fix). ### 3\. No visibility into AI-specific behavior Even if GA could see AI traffic, it doesn't have the right schema for it. GA knows about sessions, page views, conversions, and channels but not: - Which AI bots fetched which pages - How often each bot returns - Whether your `/llms.txt` is being fetched and by whom - How long since GPTBot last cached your /pricing page - Whether your content shows up in the AI-Optimized HTML bots actually get served These questions don't map to GA's data model at all. You can't build a custom report or property to answer them, because the underlying events never enter GA's pipeline. ## What Crawlytics captures that GA can't | Question | GA4 | Crawlytics | | --- | --- | --- | | How many sessions did real human visitors have last week? | ✓ Excellent | — | | Which content is converting? | ✓ Excellent | — | | What % of visits came from Google organic vs paid vs social? | ✓ Excellent | — | | Which AI bots are crawling my site right now? | ✗ Not captured | ✓ Real-time, per-bot | | How often is GPTBot fetching my pricing page? | ✗ Not captured | ✓ Per-page time series | | Did Perplexity drive any real human visitors this month? | ~ Shows in "direct" — mis-attributed | ✓ Per-LLM UTM attribution | | Which AI assistants are citing me most? | ✗ Not captured | ✓ Per-source breakdown | | Is my /llms.txt being fetched? | ✗ Not captured | ✓ Per-bot fetch log | | How does my AI bot traffic trend compare to last month? | ✗ Not captured | ✓ Date-range compare | | Are AI agents transacting (checkout, leads, bookings)? | ~ Sees the conversion, can't attribute to agent | ✓ WebMCP-level conversion attribution | ## How they work together The cleanest setup uses both: 1. **GA4 stays primary** for human funnel optimization — conversions, paid channels, organic traffic, audience targeting. 2. **Crawlytics handles AI** — bot crawl frequency, AI referral attribution (via UTM injection that flows back _into_ GA), llms.txt fetches, WebMCP agent activity. 3. **Crawlytics' UTM injection feeds GA**, so your AI referral traffic shows up in GA's standard channels report as `chatgpt / ai_referral`, `claude / ai_referral`, etc. — recoverable in any GA report that respects UTM params. This means you don't lose anything by adding Crawlytics. Your GA reports get more accurate (because previously-"direct" AI traffic now shows its real source), and you gain a whole new analytics surface for AI-specific behavior that GA was never going to provide. ## What about Cloudflare's free AI bot tracking? Cloudflare Radar shows aggregated, industry-wide AI bot traffic distribution. It's a great public reference. But it's not per-customer — you can't see which pages on _your_ site GPTBot is reading, how often, or whether the trend is up or down. For per-customer analytics with the same depth that GA provides for human traffic, you need a dedicated tool. [Full Cloudflare comparison here](https://crawlytics.app/c/claude/blog/crawlytics-vs-cloudflare-markdown-for-agents). ## Pricing comparison | | GA4 | Crawlytics Visibility | Crawlytics Commerce | | --- | --- | --- | --- | | Monthly price | Free (with 360 upgrade ~$150k/yr enterprise) | $29.99 | $49.99 | | Human session analytics | ✓ | — | — | | AI bot tracking | — | ✓ | ✓ | | AI referral attribution | ~ broken on in-app | ✓ | ✓ | | llms.txt generation | — | ✓ | ✓ | | WebMCP agent commerce | — | — | ✓ | ## When you actually need Crawlytics over GA You probably need both if any of these are true: - Your "direct traffic" has been growing without an obvious cause - You publish content that you suspect AI assistants are citing but can't measure it - You want to know which AI providers prefer your content - You're considering blocking AI crawlers and want to make the decision based on data - You run an ecommerce site and want AI agents to convert (WebMCP) - You're optimizing for AI search and need to measure what's working If none of those apply, GA alone is probably fine for another quarter or two but the trend lines say "another quarter or two" is roughly the lifespan of that statement. ## Related Written by Crawlytics Team. Crawlytics tracks AI bots, generates llms.txt, and powers WebMCP commerce, all from one snippet on any stack. [See how it works →](https://crawlytics.app/c/claude/) ## Frequently Asked Questions ### What about Cloudflare's free AI bot tracking? Cloudflare Radar shows aggregated, industry-wide AI bot traffic distribution. It's a great public reference. But it's not per-customer — you can't see which pages on your site GPTBot is reading, how often, or whether the trend is up or down. For per-customer analytics with the same depth that GA provides for human traffic, you need a dedicated tool. Full Cloudflare comparison here. ### Will Crawlytics replace my Google Analytics? No. Crawlytics doesn't track human page views, sessions, conversions, or any of the things GA4 is good at. It tracks AI-specific events GA4 can't see. Use both. ### Will adding Crawlytics affect my GA data? Yes, in a good way. Crawlytics' UTM injection means previously-"direct" AI traffic starts showing up in GA with proper source attribution (chatgpt, claude, etc.). Your channels report gets more accurate. Nothing else about your GA setup changes. ### Can I just turn off "exclude bots" in GA to see AI bots? No. The exclude-bots setting is unrelated. GA can't see AI bots because they don't run JavaScript, so the GA tag never fires. Toggling the setting won't help. ### Does Crawlytics support GA4 Looker Studio integration? Not directly today. The export-to-CSV/JSON endpoint that would feed Looker is on the roadmap. For now, the in-product dashboard is the main reporting surface. --- title: "Crawlytics vs Cloudflare Markdown for Agents: Honest Comparison" type: [Organization, Article, BreadcrumbList, WebSite, FAQPage] author: Crawlytics Team publisher: Crawlytics datePublished: 2026-06-03 dateModified: 2026-06-03 canonical: https://crawlytics.app/c/claude/blog/crawlytics-vs-cloudflare-markdown-for-agents category: blog wordCount: 2690 readingTime: 13 min crawledAt: 2026-06-21 16:40:12 lastVerified: 2026-08-25 13:09:19 site: https://crawlytics.app/c/claude/ --- # Crawlytics vs Cloudflare Markdown for Agents: Honest Comparison ## Summary Cloudflare converts HTML to markdown on demand. Crawlytics serves AI-Optimized HTML to every bot, plus stable llms.txt. Different format bets — and ChatGPT cannot read markdown. An honest decision guide. ## Key facts - Three things, all at the network edge: - Crawlytics' Visibility tier ($29. - I'm not going to pretend Cloudflare's offering is weak. - Cloudflare Radar shows aggregated industry-wide bot traffic. - I'll be transparent: Cloudflare's feature has patterns Crawlytics should match, and the team is working on them: Quick answer Cloudflare's Markdown for Agents is a free edge feature that converts HTML to markdown on demand — but only when an AI agent sends `Accept: text/markdown`, which most don't. It's a clean primitive, but markdown isn't universal: ChatGPT-User, the fetcher that fires when someone pastes a link into ChatGPT, discards `text/markdown` as unreadable. Crawlytics instead serves **AI-Optimized HTML** — clean, chrome-free, with JSON-LD — to every AI bot by routing on the User-Agent, so there's nothing to negotiate and every fetcher can read it. It also publishes stable `/llms.txt` markdown, per-customer bot analytics, ChatGPT/Claude/Perplexity referral attribution, and WebMCP agent commerce — none of which Cloudflare ships. **Use Cloudflare if you're on Pro+ and only need on-demand markdown for the agents that ask for it. Use Crawlytics if you want universal HTML coverage, analytics, attribution, llms.txt, or you're not on Cloudflare. Run both to layer Cloudflare's edge markdown under Crawlytics' measurement.** Cloudflare quietly shipped [Markdown for Agents](https://blog.cloudflare.com/markdown-for-agents/) in beta a few months ago. If you're on Cloudflare Pro, Business, Enterprise, or SSL for SaaS, it's free. It converts your HTML to clean markdown on the fly, at the edge, whenever an AI agent asks for it. That overlaps with what Crawlytics does — both make your site readable by AI bots. So the obvious question: **if Cloudflare is free, why would anyone pay Crawlytics $29.99/mo for the Visibility tier?** This post is the honest answer. Cloudflare's offering is real, well-built, and free, and there are absolutely sites that should just use it. But the two products make a different bet on _format_, and that bet decides how many AI clients can actually read your content. Cloudflare converts your HTML to **markdown**, on demand, when an agent negotiates for it. Crawlytics serves **AI-Optimized HTML** — clean semantic HTML with JSON-LD, no nav or chrome — to every AI bot by routing on the User-Agent, and still publishes stable `/llms.txt` markdown alongside it. The difference matters more than the price tag, because of one inconvenient fact: the most-used live fetcher can't read markdown at all. I'm going to compare them feature-by-feature without weasel words. Then I'll give you a clean decision tree at the end. ## What Cloudflare's Markdown for Agents actually does Three things, all at the network edge: 1. **HTTP content negotiation.** When an AI agent sends `Accept: text/markdown` in the request header, Cloudflare's edge fetches the HTML from your origin, converts it to markdown on the fly, and serves it back. The agent sees the same URL but a different format. 2. **Token-count signal.** The response includes an `x-markdown-tokens` header showing the estimated token count of the markdown payload. Useful for agents budgeting their context window. 3. **Aggregated public analytics.** Cloudflare Radar now shows the distribution of content types returned to AI agents and crawlers — visible to anyone, useful for industry trend tracking. Not per-customer, not per-page. That's it. It's a clean, well-executed primitive. If you're a Cloudflare customer on Pro or above, you get it free with one toggle. ### What Cloudflare's feature does _not_ do - It does not pre-generate `/llms.txt` or `/llms-full.txt` at stable URLs. Conversion only happens when the agent specifically requests markdown via the Accept header. - It does not give you a per-customer dashboard showing which bots visited which pages on your site. - It does not tag outbound links with per-LLM UTM parameters for attribution. - It does not expose agent-callable tools (WebMCP) for commerce, leads, or bookings. - It does not work on sites that aren't behind Cloudflare's edge. None of those are criticisms. Cloudflare scoped the feature deliberately. They built the smallest useful thing and shipped it free to existing customers. Smart. ## What Crawlytics does that overlaps Crawlytics' Visibility tier ($29.99/mo) overlaps with Cloudflare's Markdown for Agents in exactly one area: **making your site's content readable by AI bots**. Same goal, different format bet. Both strip nav, footer, scripts, and boilerplate, and both return clean structured content an LLM context window can ingest cheaply. The split is in _what_ they hand back and _who has to ask for it_. Cloudflare returns **markdown**, but only to agents that send `Accept: text/markdown`. Crawlytics returns **AI-Optimized HTML** — clean semantic HTML with JSON-LD — to every AI bot, routed automatically on the User-Agent, and also publishes the same content as markdown at stable `/llms.txt` URLs (with raw markdown still available per page via `?format=md`). | Aspect | Cloudflare Markdown for Agents | Crawlytics | | --- | --- | --- | | Format served to bots at the page URL | Markdown | AI-Optimized HTML (clean semantic HTML + JSON-LD) | | How a bot gets the AI version | Must send Accept: text/markdown — content negotiation | Crawlytics detects the bot's User-Agent and routes it automatically — no header required | | Coverage of AI fetchers | Only agents that negotiate and can read markdown | Every fetcher reads HTML — universal, including ChatGPT-User | | Stable markdown URLs | None — conversion is per-request only | /llms.txt, /llms-full.txt, and per-page /md (?format=md for raw markdown) | | Token count returned | x-markdown-tokens response header | Not yet (on the roadmap — see below) | | Hosting requirement | Must be on Cloudflare Pro+ | Any host — Vercel, Netlify, WordPress, nginx, Apache, raw HTML | Two of these decide whether your content is actually reachable. ### AI-Optimized HTML is universal — and the most-used fetcher can't read markdown Here's the wedge, and it's the whole reason format strategy matters. **ChatGPT-User — the fetcher that fires the moment someone pastes your link into ChatGPT — discards `text/markdown` as unreadable.** It expects HTML. So even on a site where Cloudflare's feature is enabled _and_ a client negotiates for markdown, the single most important live fetcher gets content it throws away. Markdown isn't dead — Claude's fetcher, for example, tolerates it fine — but it isn't universal, and the gap lands on exactly the traffic you most want to win. AI-Optimized HTML sidesteps the whole problem: every AI fetcher reads HTML. Crawlytics serves it by detecting the bot's User-Agent and returning a clean, chrome-free version of the page (with JSON-LD) at the same URL a human would visit — humans still get your full site. Nothing has to be negotiated, and nothing gets discarded. Cloudflare's markdown is a clean primitive, but it only reaches agents that both ask for markdown and can read it; Crawlytics reaches all of them. ### Stable /llms.txt URLs matter because not every AI client fetches per-page The `llms.txt` standard ([llmstxt.org](https://llmstxt.org/), or our [full setup guide here](https://crawlytics.app/c/claude/blog/what-is-llms-txt-guide)) emerged because LLM crawlers needed a predictable place to look. The convention is: put a markdown file at `/llms.txt`, mention your top pages, and AI systems will discover it. Tools like ChatGPT, Claude, Perplexity, and AI Overviews increasingly fetch this file directly — and here markdown is the right format, because the clients that consume `/llms.txt` expect it. Cloudflare's Markdown for Agents has no stable URL at all; conversion only happens per request when an agent negotiates for it. So there's no single place an AI system can go to discover your site's structure. Crawlytics pre-generates `/llms.txt` and `/llms-full.txt` and keeps them current with a daily re-crawl. Between AI-Optimized HTML at every page URL and markdown at a stable `/llms.txt`, you cover both the fetchers that crawl pages and the clients that look for the index. Cloudflare covers neither without the agent asking first. ### Host independence matters because most sites aren't on Cloudflare Pro Cloudflare's Markdown for Agents is free, but only if you're already a Pro+ customer — that's $25/mo to start. If you're on Vercel, Netlify, plain GitHub Pages, WordPress.com, Squarespace, or any of the other ~70% of sites that aren't proxied through Cloudflare, the feature doesn't exist for you. Crawlytics is host-agnostic. The snippet runs as a Cloudflare Worker, a Vercel middleware, a WordPress plugin, an Express middleware, or a static nginx/Apache log shipper. Pick what matches your stack. ## Where Cloudflare clearly wins I'm not going to pretend Cloudflare's offering is weak. Three things they do better: ### 1\. Edge latency on the markdown path When an agent does negotiate for markdown, Cloudflare converts your live origin at the edge during the request — no cron, no re-crawl lag. Crawlytics' AI-Optimized HTML at every page URL is rendered per request from the most-recently-crawled version of that page — the same crawled content as `/llms.txt`, just rendered to HTML instead of markdown — so it carries the same crawl lag: publish at 9:00 a.m. and the page may not reflect it until the next crawl. Cloudflare reads the current origin every time, so it's genuinely fresher for just-published pages, and that edge applies across the board, not just on a markdown path. It's a real freshness tradeoff — Crawlytics gives up a little currency in exchange for UA-routed universal HTML, per-page analytics, and attribution — but on raw freshness, Cloudflare wins here. ### 2\. Cost — if you're already a Cloudflare customer If you're already paying $25/mo for Cloudflare Pro, Markdown for Agents is bundled at no additional cost. You toggle it on and you're done. Crawlytics' Visibility tier is $29.99/mo on top of whatever else you're paying for hosting. ### 3\. The `x-markdown-tokens` header This is a small but real quality-of-life feature. AI agents that fetch your markdown can budget their context window without having to count tokens themselves. It's the kind of signal that becomes a de facto standard once enough hosts return it. Crawlytics doesn't return this header today — it should, and it will (more below). ## Where Crawlytics clearly wins ### 1\. Per-customer bot analytics Cloudflare Radar shows aggregated industry-wide bot traffic. It's a great public reference, but it doesn't tell you which pages on _your_ site GPTBot is reading, when, or which paths it's ignoring. Crawlytics gives you that as a real-time dashboard — per-bot, per-page, per-day, with time-series charts and a 14-day projection. If you're trying to optimize for AI search (figuring out which content is getting cited, which is being ignored, which deserves a refresh), you need per-customer data. Cloudflare's free tier doesn't provide it. ### 2\. Per-LLM UTM attribution This is the biggest functional gap. ChatGPT, Claude, and Perplexity's in-app browsers strip the Referer header on outbound clicks. So when a user taps a citation in ChatGPT mobile, your Google Analytics logs the visit as "Direct / None." Cloudflare can't fix this — they don't touch your outbound links. Crawlytics solves this by injecting per-LLM UTM tags (`utm_source=chatgpt`, `utm_medium=ai_referral`) into the links inside the AI-Optimized HTML each bot fetches. When ChatGPT cites your URL, the UTMs travel with it. Your analytics see `chatgpt` as the source, not `(direct)`. [Full write-up here.](https://crawlytics.app/c/claude/blog/chatgpt-direct-traffic-fix) This is the kind of feature that exists because someone went looking for the problem. Nobody at Cloudflare has shipped it. Probably nobody will until the problem gets loud enough. ### 3\. WebMCP commerce WebMCP is the draft web spec (currently in Chrome 146+ Canary) that exposes `navigator.modelContext`, letting a page register tools an in-browser AI agent can invoke. Crawlytics' Commerce tier ($49.99/mo) ships a one-tag snippet that registers your tools (search, checkout, book, lead-capture) and attributes conversions back to the agent that drove them. Cloudflare has nothing in this category. They convert HTML to markdown. They don't expose action surfaces, they don't handle conversion attribution, they don't integrate with Stripe/Paddle/Lemon Squeezy for revenue tracking. If you're running an ecommerce site and you want AI agents to actually _buy_ things on your site, Cloudflare can't help. Crawlytics can. ### 4\. Multi-host support Already covered above but worth repeating: Crawlytics works on every stack. Cloudflare's feature only works for Cloudflare-proxied traffic. ### 5\. Audit + scoring Crawlytics scores each page on six signals (sitemap priority, URL depth, category, word count, recency, has-meta-description) and surfaces an agent-readiness score so you know which content is winning and which needs work. Cloudflare just converts whatever you have. ## Things Crawlytics should adopt from Cloudflare I'll be transparent: Cloudflare's feature has patterns Crawlytics should match, and the team is working on them: 1. **Return the `x-markdown-tokens` header.** When Crawlytics serves raw markdown (via `/llms.txt` or `?format=md`), it should include the token-count header so agents can budget their context window without counting themselves. Cheap to add, useful, and on its way to becoming a de facto standard. 2. **Honor `Accept: text/markdown` on the page URL too.** Crawlytics already routes AI bots to AI-Optimized HTML by User-Agent, which covers the fetchers that can't read markdown. For the agents that explicitly _prefer_ markdown and say so via the Accept header (Claude's fetcher, for instance), it would be a nice touch to return markdown inline rather than making them hit `?format=md`. The UA routing is the safe default; the Accept header is the courtesy. 3. **Support Content Signals.** Cloudflare's proposed mechanism for declaring whether content may be used for AI training, search indexing, or inference. Open spec, no reason not to support it as a dashboard toggle. None of those affect what Crawlytics charges for — they're table stakes the whole industry is moving toward. ## The decision tree Here's the clean answer. Pick the path that matches you: ### Use Cloudflare Markdown for Agents (skip Crawlytics) if: - You're already on Cloudflare Pro+ and don't want to pay anything additional - You don't care about per-customer bot analytics — Cloudflare Radar's aggregated view is enough for your decision-making - You're not trying to recover ChatGPT/Claude/Perplexity referral attribution from "direct" traffic in GA - You're not doing ecommerce / lead-gen and don't need WebMCP - You're fine reaching only the agents that send `Accept: text/markdown` and can read it — accepting that ChatGPT-User and other HTML-only fetchers fall through ### Use Crawlytics if: - You're not on Cloudflare (Vercel, Netlify, WordPress, nginx, anything else) - You need per-customer bot analytics — which pages, which bots, when, trending up or down - You're losing ChatGPT/Claude/Perplexity referral attribution to "(direct)" in Google Analytics - You want every AI fetcher — including ChatGPT-User, which can't read markdown — to get clean, readable content without negotiating for it - You want stable `/llms.txt` URLs any AI client can discover, not just per-request conversion - You want WebMCP agent commerce (Commerce tier) - You manage multiple sites and want a portfolio dashboard ### Use both if: - You're on Cloudflare Pro+ and want its edge markdown for the agents that negotiate for it, _and_ you want Crawlytics' universal AI-Optimized HTML, analytics, attribution, and llms.txt on top. They layer cleanly — Cloudflare handles the markdown-on-demand path, Crawlytics covers every other fetcher and measures all of it. For most sites I talk to, the answer is "use Crawlytics" because they're not on Cloudflare Pro and they want the dashboard. For Cloudflare-native shops that don't care about per-customer attribution and are comfortable reaching only markdown-capable agents, just turn on Cloudflare's free feature and skip the bill. ## Does Cloudflare's free offering change Crawlytics' pricing? Short answer: **no**. Cloudflare commoditized one layer (on-demand HTML → markdown conversion) for a subset of the market (their existing Pro+ customers) and a subset of agents (the ones that negotiate for markdown and can read it). That's not what Crawlytics charges for. Crawlytics charges for the layer above: universal AI-Optimized HTML that every fetcher can read, stable `/llms.txt` URLs, per-customer analytics, per-LLM attribution, and WebMCP commerce. Cloudflare doesn't compete in any of those. If anything, Cloudflare's launch validates the category. Two years ago, "AI bot tracking" wasn't a phrase anyone used. Now Cloudflare ships a feature with that pitch, attached to one of the biggest infrastructure brands on the internet. That's a signal that the market is real, growing, and worth investing in. If you want to see what a per-customer dashboard looks like, [walk through the live demo](https://crawlytics.app/c/claude/demo). If you just need free HTML→markdown and you're on Cloudflare, go enable their feature — you don't need us for that. ## Related Written by Crawlytics Team. Crawlytics tracks AI bots, generates llms.txt, and powers WebMCP commerce, all from one snippet on any stack. [See how it works →](https://crawlytics.app/c/claude/) ## Frequently Asked Questions ### Does Cloudflare's free offering change Crawlytics' pricing? Short answer: no. Cloudflare commoditized one layer (on-demand HTML → markdown conversion) for a subset of the market (their existing Pro+ customers) and a subset of agents (the ones that negotiate for markdown and can read it). That's not what Crawlytics charges for. Crawlytics charges for the layer above: universal AI-Optimized HTML that every fetcher can read, stable /llms.txt URLs, per-customer analytics, per-LLM attribution, and WebMCP commerce. Cloudflare doesn't compete in any of those. If anything, Cloudflare's launch validates the category. Two years ago, "AI bot tracking" wasn't a phrase anyone used. Now Cloudflare ships a feature with that pitch, attached to one of the biggest infrastructure brands on the internet. That's a signal that the market is real, growing, and worth investing in. If you want to see what a per-customer dashboard looks like, walk through the live demo. If you just need free HTML→markdown and you're on Cloudflare, go enable their feature — you don't need us for that. --- title: "ChatGPT Traffic Shows as \"Direct\" in GA — Here Are 3 Fixes" type: [Organization, Article, BreadcrumbList, WebSite, FAQPage] author: Crawlytics Team publisher: Crawlytics datePublished: 2026-05-28 dateModified: 2026-05-28 canonical: https://crawlytics.app/c/claude/blog/chatgpt-direct-traffic-fix category: blog wordCount: 1337 readingTime: 7 min crawledAt: 2026-06-21 16:40:18 lastVerified: 2026-08-25 13:09:21 site: https://crawlytics.app/c/claude/ --- # ChatGPT Traffic Shows as "Direct" in GA — Here Are 3 Fixes ## Summary Mobile and in-app browsers strip the Referer header on ChatGPT clicks, so GA logs them as "direct." Here is why it happens and how to recover the attribution. ## Key facts - Your web server logs the Referer header for every request before any browser-side analytics runs. - If you control where the links to your site appear (your own social posts, your newsletter, a partner site), you can add UTM parameters at the source: `? - This is the approach Crawlytics ships. - Attribution is downstream of detection. - Written by Crawlytics Team. If you've checked your Google Analytics in the past year, you've probably noticed your "Direct / None" channel growing. Some of that is people typing your URL. Most of it isn't. The boring truth: **a large and growing fraction of your "direct" traffic is actually AI assistants — ChatGPT, Claude, Perplexity, Copilot — whose in-app browsers don't pass a Referer header on outbound clicks.** GA sees a visit with no source, drops it into Direct, and you're none the wiser. Here's why it happens and three ways to start recovering the attribution. ## What's actually happening Imagine the path: 1. A user opens ChatGPT on their phone. 2. They ask "best Airbnb pricing tool for hosts" or whatever's relevant to your site. 3. ChatGPT answers and includes a citation linking to your `/pricing` page. 4. The user taps the citation. 5. The link opens in ChatGPT's in-app browser — a sandboxed WebView, not Safari, not Chrome. At step 5, your server receives a normal HTTP request. The request has: - A path: `/pricing` - A User-Agent that looks like generic mobile Safari - An **empty Referer header** That last one is the problem. The in-app browser strips Referer for privacy reasons. Apple, Google, Meta, and basically everyone else who ships an in-app browser does the same thing for outbound links. ChatGPT's app is not unique here. Google Analytics' default attribution rules see "no Referer, no UTM" and bucket the visit into `(direct) / (none)`. So does Mixpanel. So does Plausible. So does Fathom. You ranked in ChatGPT. ChatGPT cited you. A user clicked. You got the traffic. You got **none** of the credit. ## Why this matters more every month Two trends collide: 1. **AI search is growing as a discovery channel.** ChatGPT was at 700M weekly active users by mid-2025 and climbing. Perplexity, Claude, and Copilot all ship search-aware modes that cite sources. People are increasingly getting recommendations from AI before they Google anything. 2. **In-app browsing is the default.** Phone users don't tap "open in Safari." They tap the link and read in the app. The Referer strip is built into every major in-app browser. The result: a growing share of your real traffic comes from AI assistants, and a growing share of that traffic is invisible in your analytics. If you're optimizing your content strategy or your SEO based on what GA tells you, you're optimizing against a blind spot. ## Fix 1: Server-side log analysis (free, partial) Your web server logs the Referer header for every request before any browser-side analytics runs. If a visit _does_ have a Referer (some AI clients still send one — Perplexity desktop, Claude desktop in some configs), it lands in your raw access logs. You can grep for known AI assistant hosts: ``` grep -E 'chat\.openai\.com|chatgpt\.com|perplexity\.ai|claude\.ai|copilot\.microsoft\.com' /var/log/nginx/access.log ``` What this catches: desktop browser sessions where the Referer survives. What it misses: every mobile in-app browser click — which is most of them. **Coverage:** maybe 20-30% of AI assistant traffic. Better than nothing. Free. ## Fix 2: Manual UTM tagging at link source (high effort, brittle) If you control where the links to your site appear (your own social posts, your newsletter, a partner site), you can add UTM parameters at the source: `?utm_source=newsletter`, etc. That works for your owned channels. It doesn't work for AI citations, because **you don't control how ChatGPT or Claude links to you**. They cite the canonical URL they found during crawling. Whatever URL they have, that's what they share. Some teams try to game this by submitting their pages to LLMs with pre-tagged URLs. It doesn't stick. The models re-crawl, find the un-tagged canonical, and use that instead. You can't manually UTM your way out of this problem. **Coverage:** ~0% of AI assistant traffic. Don't bother. ## Fix 3: UTM injection at the AI-Optimized HTML layer (high coverage, automatic) This is the approach Crawlytics ships. The idea: 1. When an LLM bot fetches a page from your site — say GPTBot crawls `/pricing` — Crawlytics' middleware detects the bot from the User-Agent and serves **AI-Optimized HTML** instead of the standard browser page (clean semantic HTML + JSON-LD, no nav clutter or tracking scripts). 2. Before returning it, Crawlytics rewrites every internal link to append per-LLM UTM tags: `?utm_source=chatgpt&utm_medium=ai_referral&utm_campaign=crawlytics` (for GPTBot), or `utm_source=perplexity` for PerplexityBot, etc. 3. When ChatGPT later cites your page, it cites the URL it fetched — UTM params and all. 4. A user taps the citation in ChatGPT iOS. The in-app browser strips Referer (still). But the URL itself has `utm_source=chatgpt` in it, so Google Analytics, Mixpanel, Plausible — anything that respects UTMs — sees `chatgpt` as the source. The attribution lives in the URL, not in the Referer header. The in-app browser can't strip it. **Coverage:** 100% of citations crawled from now on. Doesn't recover anything that was crawled before the middleware was installed (you can't retroactively change a URL ChatGPT has memorized) — but going forward, every fresh re-crawl tags the page and every fresh citation carries the UTM. ## What does the "fixed" data look like? Before: ``` Channel Sessions Organic Search 18,432 Direct / None 12,108 ← AI traffic hiding here Referral 2,847 Social 1,203 ``` After (a few weeks of UTM injection running): ``` Channel Sessions Organic Search 18,432 Direct / None 8,742 AI Referral 3,366 ← chatgpt + claude + perplexity + gemini ├── chatgpt 1,847 ├── perplexity 812 ├── claude 497 └── gemini 210 Referral 2,847 Social 1,203 ``` You don't suddenly get _more_ traffic — you just see where it was actually coming from. Which means you can: - Know which AI assistants are sending you the most visits (often Perplexity over-indexes here) - Know which of your pages are getting cited in AI answers (top landing pages by AI source) - Make content decisions based on real attribution instead of guessing - Justify the time you spend writing AI-friendly content ## Where to start Attribution is downstream of detection. Before you can fix where AI traffic is bucketed, you have to confirm AI is fetching and citing your site in the first place — the [AI citation detection playbook](https://crawlytics.app/c/claude/blog/how-to-track-ai-citations) covers the server-log and prompt-test side. Detection tells you whether you're showing up; attribution tells you whether the visits convert. If you want to see this working before paying for anything, the [live demo dashboard](https://crawlytics.app/c/claude/demo) shows the AI Referrals panel running on synthetic data — same component the real customer dashboard renders. The full [AI attribution feature page](https://crawlytics.app/c/claude/features/ai-attribution) walks through the install flow per stack (Cloudflare Worker, Vercel middleware, nginx, Express, WordPress). Or just [start a trial](https://crawlytics.app/c/claude/checkout?plan=visibility&billing=monthly&bundle=solo) — it's $29.99/mo for Visibility, which includes the attribution layer plus bot tracking and llms.txt generation. ## Related Written by Crawlytics Team. Crawlytics tracks AI bots, generates llms.txt, and powers WebMCP commerce, all from one snippet on any stack. [See how it works →](https://crawlytics.app/c/claude/) ## Frequently Asked Questions ### What does the "fixed" data look like? Before: ### Does this affect SEO? No. Googlebot is not in the bot list and is never served the tagged AI-Optimized HTML — it gets your normal browser page with normal internal links. Search engines see your site exactly as before. ### Will the UTM params show in the user's address bar? Yes — same as any UTM tag from a paid channel. Most marketers consider that acceptable. If you don't, you can strip the params client-side after recording the visit (one line in your analytics layer). ### What about Bing's Copilot? Apple Intelligence? The mapping handles them: utm_source=copilot for Microsoft Copilot bots, utm_source=apple_intelligence for Applebot-Extended. Same pattern for every detected LLM provider — currently 12 mapped sources covering OpenAI, Anthropic, Perplexity, Google Gemini, Microsoft Copilot, Meta AI, ByteDance Doubao, You.com, Cohere, xAI Grok, Apple, and Mistral. ### Does this replace Google Analytics? No. It feeds GA (and Mixpanel, Plausible, Fathom — anything that reads UTM params). Crawlytics has its own dashboard for AI-specific surfaces (per-bot crawl frequency, llms.txt fetches, WebMCP tool invocations) but the referral attribution layer is designed to make your existing analytics smarter, not replace them. --- title: "How to Create an llms.txt File (and Test It) in 2026" type: [Organization, Article, BreadcrumbList, WebSite, FAQPage] author: Crawlytics Team publisher: Crawlytics datePublished: 2026-06-05 dateModified: 2026-06-05 canonical: https://crawlytics.app/c/claude/blog/what-is-llms-txt-guide category: blog wordCount: 1961 readingTime: 10 min crawledAt: 2026-06-21 16:40:13 lastVerified: 2026-08-25 13:09:19 site: https://crawlytics.app/c/claude/ --- # Acme Tools ## Summary A step-by-step llms.txt setup for 2026: generate the markdown index, add the right sections, host it at /llms.txt, and confirm AI crawlers actually read it. ## Key facts - The proposal came from Jeremy Howard (Answer. - The official spec is short — under 50 lines of normative text — but a few conventions have hardened in practice that aren't on llmstxt. - The three files do different jobs and AI clients fetch them at different times. - For a small site — under 30 URLs — open a text editor, write the file by hand, drop it at the root. - Direct answer: not in the classic Google-ranking sense. If you've heard the phrase `llms.txt` in the past six months, it was probably from a Vercel changelog, an Anthropic doc page, or someone on r/SEO insisting it's the new `robots.txt`. None of those tell you the full story. This guide does — what the file actually is, who reads it, how to ship one on any host, and whether the time investment pays off. The short version: `llms.txt` is a markdown file at the root of your site that gives AI assistants a clean, structured index of what's worth reading. It's not a ranking signal. It's a delivery format. And as more AI clients start fetching it by default, the cost of not having one is going up. ## What llms.txt actually is (and isn't) `llms.txt` is a plain-text markdown file served at the root of your domain — `https://yoursite.com/llms.txt`. Inside it, you list the pages on your site that matter most to a reader who's trying to understand what you do, organized into sections with descriptions. Here's a stripped-down example: ``` > Acme builds open-source CLI utilities for inspecting Docker images. ## Docs - [Getting Started](https://acme.dev/docs/start): install in 30 seconds, scan your first image - [API Reference](https://acme.dev/docs/api): every command, every flag, every exit code ## Blog - [Why we rewrote our scanner in Rust](https://acme.dev/blog/rust-rewrite): 11x faster, 90% less memory - [The case for SBOMs in 2026](https://acme.dev/blog/sbom-2026) ``` That's the whole thing. A heading with your site name, a one-sentence summary, then markdown lists grouped by section. AI systems fetch it, parse it, and use it to decide what to read next. What it _isn't_: - **It isn't `robots.txt`.** Robots.txt is exclusion — telling crawlers what not to fetch. llms.txt is inclusion — telling them what's worth fetching first. - **It isn't a sitemap.** A sitemap lists every URL with metadata for search engines. llms.txt is a curated, human-edited index for LLM ingestion. - **It isn't an SEO ranking factor.** Google has not confirmed it reads `llms.txt`. As of mid-2026, AI Overviews still pull from web search, not from `llms.txt` directly. - **It isn't required.** No client will fail to read your site without it. The fallback is HTML scraping, which is messier and costs the agent more tokens. ## Why the format exists, and who's actually adopting it The proposal came from Jeremy Howard (Answer.AI, fast.ai) in September 2024. The pitch was simple: LLM context windows are expensive, HTML is noisy, and there should be a way for a site owner to hand a clean markdown index to any AI client that wants it. He set up [llmstxt.org](https://llmstxt.org/) with the spec and a directory. The adoption curve looked like most open conventions: a few months of "is this a real thing?", followed by enough notable sites shipping it that the question became "why don't you have one yet?" As of mid-2026, you'll find `llms.txt` live on Anthropic's docs, Vercel, Cursor, Stripe, Supabase, Vue.js, Astro, dbt Labs, and thousands of independent sites. The directory at llmstxt.org tracks public adopters. Cloudflare's Markdown for Agents and OpenAI's developer docs both reference the convention. What the AI clients actually do with it varies. ChatGPT and Claude fetch `llms.txt` opportunistically when you give them a URL or ask about a site by name. Perplexity prefers `llms-full.txt` when available. Custom GPTs and Claude Projects use it as a seed index. Codegen tools (Cursor, Windsurf, Continue) pull it to pre-warm their context when you point them at a library. ## The format: structure, sections, and the rules nobody documents The official spec is short — under 50 lines of normative text — but a few conventions have hardened in practice that aren't on llmstxt.org. Here's what works: ### Required structure 1. **H1 with the site name.** One line. Just the brand. 2. **Blockquote with a one-sentence description.** What the site is, in plain English. AI assistants quote this verbatim when summarizing. 3. **Optional explanatory paragraphs.** Anything that helps an agent understand context — what you sell, who you serve, what's out of scope. 4. **H2 sections.** One per topic area. Common headings: Docs, Guides, API, Blog, Examples, About. 5. **Markdown list of links under each H2.** Format: `- [Page title](https://full-url): short description`. The description is what tips an agent toward fetching that URL. 6. **Optional H2 named "Optional".** Pages that are nice-to-have but not core. Agents on a token budget can skip this section. ### Hard rules - Use absolute URLs, not relative. Agents may not know your origin. - Use full sentences in descriptions, not keyword fragments. The description teaches the LLM what's in the doc. - Keep the whole file under 30k tokens (roughly 100KB of text). Beyond that, agents start truncating in the middle of sections. - Don't put HTML or JavaScript in the file. Strict markdown only. ### Soft conventions that emerged from real adopters - Order sections by importance. The top of the file gets read most. - Lead each section with the page agents are most likely to need first — a "Getting Started" or "Overview" page typically. - If you have a search interface, link to it. Agents will use it. - Update the file when content changes materially. Stale `llms.txt` wastes agent fetches. ## llms.txt vs llms-full.txt vs robots.txt The three files do different jobs and AI clients fetch them at different times. The table makes the distinction concrete: | File | What it contains | Who reads it | When | | --- | --- | --- | --- | | /robots.txt | Crawl rules (allow / disallow / sitemap pointer) | All crawlers and most AI bots | Before fetching anything else | | /llms.txt | Curated index of URLs with descriptions | AI assistants and code agents | When deciding what to fetch from your site | | /llms-full.txt | Full concatenated markdown of your site | AI clients that want everything in one shot | For one-fetch ingestion (often by code agents) | You don't have to choose. The right move for most sites is to ship all three: `robots.txt` for crawl control, `llms.txt` as the curated index, `llms-full.txt` as the bulk download option. Crawlytics generates the latter two automatically. [Here's how it differs from Cloudflare's edge approach.](https://crawlytics.app/c/claude/blog/crawlytics-vs-cloudflare-markdown-for-agents) ## Three ways to generate llms.txt (and when to pick each) ### Path 1 — Hand-write it For a small site — under 30 URLs — open a text editor, write the file by hand, drop it at the root. Total time: 15 minutes for a focused site, an hour for one with a lot of categories. Pick this if you have a stable site that doesn't change weekly, or if you want full editorial control over which pages the AI sees first. Documentation sites with a clean structure (10-20 top pages plus an API reference) often do this and never touch the file again. Downside: you have to remember to update it. A stale `llms.txt` sends agents to dead URLs and old content. ### Path 2 — Let Crawlytics generate and host it Crawlytics crawls your sitemap nightly, scores each URL on six signals (depth, recency, word count, sitemap priority, meta description, category), groups by section, and writes `llms.txt` and `llms-full.txt` to stable URLs that work on any host — Vercel, Netlify, WordPress, raw nginx, anything. You add a snippet, the file regenerates daily, you get a dashboard showing which AI bots are fetching it and from where. Pick this if you have a fast-moving content site (blog, docs that ship often, ecommerce catalog), if you want analytics on bot fetches, or if you don't want to maintain the file by hand. [The Visibility tier is $29.99/mo](https://crawlytics.app/c/claude/pricing) and includes the generator plus per-bot analytics. ### Path 3 — Generate it at build time If you run a static site (Astro, Next.js, Hugo, Eleventy), you can write a build-time script that walks your content collections, formats markdown, and writes `/public/llms.txt` before the build finishes. Vercel publishes a reference script. Astro and Next plugins exist. Pick this if you're already comfortable with custom build steps and you don't want a hosted dependency. Downside: no per-bot analytics, no fetch logging, no UTM injection for attribution. The file just exists. ## Does llms.txt help your SEO? Direct answer: not in the classic Google-ranking sense. Google has not confirmed that Googlebot reads `llms.txt` as a ranking signal, and AI Overviews still pull from the regular web index, not from `llms.txt` directly. What it does help is _AI search_ — the layer of ChatGPT, Claude, Perplexity, Gemini, and the dozens of vertical assistants that fetch sites directly when answering questions. In that channel: - Sites with `llms.txt` get cited more often because the agent doesn't have to decide what to read — you already told it - Token efficiency matters — agents working under a context budget will skip messy HTML in favor of clean markdown, which means your site gets fully read instead of partially scraped - Code agents (Cursor, Continue, Windsurf) prefer `llms-full.txt` when building features against your API, because they can load the whole reference in one fetch The way to think about it: `llms.txt` isn't an SEO tactic, it's an AEO (Answer Engine Optimization) primitive. If you care about being cited in AI answers, ship it. If you only care about Google's blue-link rankings, it's neutral — you won't be penalized for having one, you won't be rewarded. For the broader playbook on AI search, [our AEO framework covers the four layers](https://crawlytics.app/c/claude/resources/ai-search-optimization): technical accessibility, content structure, signal generation, and attribution recovery. ## Pre-flight checklist before you ship Before pushing `llms.txt` live, run through this list. The failure modes are all silent — your file will exist, agents will fetch it, and you won't know it's broken unless you check: 1. **File loads at `https://yoursite.com/llms.txt`.** Fetch it with curl. Confirm 200 status. No redirects, no auth challenge. 2. **Content-Type is `text/plain` or `text/markdown`.** Some hosts default to `application/octet-stream` for unknown extensions, which causes downloads instead of inline display. Set the MIME type explicitly. 3. **All URLs are absolute.** Relative URLs break for any agent that doesn't know your origin. 4. **No 404s in the link list.** Stale links cost agent fetches and degrade your citation quality. 5. **Total size under 100KB.** Above that, agents start truncating, and they truncate from the bottom — your less-important sections get cut first, but if your file is over 200KB, useful content gets dropped too. 6. **Descriptions are sentences, not keyword salad.** The description teaches the model what the page is about. Write it like a librarian, not a meta description. 7. **The file is in `robots.txt` as allowed.** If you have a blanket `Disallow: /`, allow `/llms.txt` explicitly so AI bots can still reach it. 8. **You have a re-generation plan.** Whether it's a cron, a build hook, or a hosted generator, the file needs to stay current. Stale beats nothing, but fresh beats stale. If you want a one-click check on all eight, the [free Agent-Ready Grader](https://crawlytics.app/c/claude/agent-ready) runs through them in 10 seconds and gives you a score plus the broken items. ## The bottom line llms.txt is a small file with a long tail of impact. It costs you 15 minutes to hand-write or a one-line snippet to automate. The downside is zero. The upside is being readable to every AI client that asks — which, on the current trajectory, is most of them by the end of 2026. Don't overthink the format. Ship it, point it at your best pages, and update it when you add new ones. The agents that matter are already looking for it. ## Related Written by Crawlytics Team. Crawlytics tracks AI bots, generates llms.txt, and powers WebMCP commerce, all from one snippet on any stack. [See how it works →](https://crawlytics.app/c/claude/) ## Frequently Asked Questions ### Does llms.txt help your SEO? Direct answer: not in the classic Google-ranking sense. Google has not confirmed that Googlebot reads llms.txt as a ranking signal, and AI Overviews still pull from the regular web index, not from llms.txt directly. What it does help is AI search — the layer of ChatGPT, Claude, Perplexity, Gemini, and the dozens of vertical assistants that fetch sites directly when answering questions. In that channel: --- title: "What Is WebMCP? AI Agent Actions Explained (2026)" type: [Organization, Article, BreadcrumbList, WebSite, FAQPage] author: Crawlytics Team publisher: Crawlytics datePublished: 2026-06-05 dateModified: 2026-06-05 canonical: https://crawlytics.app/c/claude/blog/webmcp-explained-ai-agent-actions category: blog wordCount: 2131 readingTime: 11 min crawledAt: 2026-06-21 16:40:13 lastVerified: 2026-08-25 13:09:19 site: https://crawlytics.app/c/claude/ --- # What Is WebMCP? AI Agent Actions Explained (2026) ## Summary WebMCP is the draft browser API letting sites expose tools (search, cart, booking) to in-browser AI agents. The spec, who invokes it today, and how to ship it. ## Key facts - For the past year, the AI-on-the-web playbook has been about being _readable_. - This is the section most WebMCP coverage skips. - The reason WebMCP can be shippable in a browser without a thousand abuse vectors is the consent model. - You can write WebMCP integrations from scratch using the raw `navigator. - When a WebMCP-aware agent takes an action on your site, you want to know which agent, which session, and whether the action converted. The read-only era of AI on the web is starting to give way to a read-and-do era. AI agents have spent the last year fetching pages, parsing content, and summarizing what they find — but stopping short of clicking, submitting, or buying anything on the user's behalf. WebMCP is the proposed browser API that lets a site offer those action surfaces to an agent that knows how to ask. The honest framing for mid-2026: WebMCP is real as a spec, prototyped in Chromium-based browsers and in agent-first browsers like Perplexity Comet, and actively used by a growing set of browser extensions and custom-built agents. It is not yet how ChatGPT and Claude's first-party apps operate — those still use citation rendering or screen-control. So adding a WebMCP snippet today is a forward investment: you become invocable by the WebMCP-aware agents that exist now, and you're ready when the larger consumer agents add support. This is the explainer. What the spec does, who invokes it today vs who doesn't, what an agent action looks like, the safety model, and how to add it to your site without rewriting anything. ## What WebMCP actually changes — the shift from "read" to "do" For the past year, the AI-on-the-web playbook has been about being _readable_. Ship `llms.txt`. Make sure your meta descriptions are clean. Render server-side so agents don't choke on JavaScript. Optimize for citation. WebMCP moves the goal post. It lets your site register _tools_ — JavaScript functions with structured inputs and outputs — that an in-browser AI agent can invoke. A tool can be anything: `searchProducts(query)`, `addToCart(sku, qty)`, `requestQuote(name, email, project)`, `bookAppointment(slot, contact)`. A WebMCP-aware agent reads the tool catalog, decides which one matches the user's intent, and calls it. The browser shows the user a confirmation. The action happens. `llms.txt` made you readable. WebMCP makes you _actionable_, for agents that know how to act. Different layer, different upside, different timeline on adoption. ## The spec in four sentences 1. **`navigator.modelContext`** is the entry point. A browser-provided object that exposes `registerTool()`, `unregisterTool()`, and a tool registry. 2. **A tool is a JSON Schema + a handler function.** The schema describes the inputs (and the expected output). The handler is your normal site code — it runs in your page's JavaScript context with your normal session, cookies, and APIs. 3. **A WebMCP-aware AI agent reads the registered tools and invokes them.** The agent has to be implemented against the API — not every browser-resident agent is. 4. **The browser renders a confirmation UI for the user before the tool runs.** The site does not write its own consent dialog — the browser owns it, which is what makes the trust model work. That's the whole API surface a developer needs to think about. The complexity is on the browser and agent side, where the integration, sandboxing, and consent UI live. ## Who actually invokes WebMCP today (and who doesn't) This is the section most WebMCP coverage skips. The honest reality for mid-2026: ### Agents that invoke WebMCP today (small but real) - **Agent-first browsers** experimenting with the API — Perplexity Comet is the most active; Brave Leo's agent mode and Arc Search's agent flows are evaluating it. - **Browser extensions** that ship their own in-page agent — open-source projects, vertical shopping agents, research assistants. - **Custom enterprise agents** built on the Anthropic or OpenAI SDKs that target specific WebMCP-enabled sites. ### Agents that don't invoke WebMCP today (most consumer flows) - **ChatGPT first-party apps and chat.openai.com** use a mix of citation rendering and, in agentic browse mode, OpenAI's Operator-style screen-control approach — not WebMCP API calls. - **Claude first-party apps and claude.ai** use Anthropic's Computer Use, which takes screenshots and clicks at the OS level — also not WebMCP. - **Most mobile AI chat apps** when they open a URL in-app render content; they don't invoke registered tools. ### Browser-side support Chromium-based browsers expose `navigator.modelContext` behind a flag or origin trial in current builds. Safari and Firefox have been evaluating but have not shipped. Stable, default-on, every-browser support is still ahead. ### Why ship the snippet anyway Three reasons, in order of immediate vs eventual return: 1. **Today (small but real):** the agents listed above can invoke your tools right now. If your customers use Comet, an agent extension, or a custom buying agent, you become actionable to them. 2. **Within 6-12 months (the realistic adoption window):** as the spec stabilizes and consumer agents add WebMCP support, sites that already registered tools start showing up as the actionable choice. Being early in directory listings of WebMCP-enabled sites matters. 3. **As the better path than screen-control:** agents that don't use WebMCP today rely on screen-control, which is messy, slow, and breaks when you change your CSS. If you make yourself easy to invoke via the API, agents shift to it because it's more reliable. You're shaping the path of least resistance. ## What an agent action looks like — three examples Specs are abstract. Concrete examples are not. Here's what three flows look like end-to-end, when invoked by a WebMCP-aware agent. (Important context: these scenarios run today on Comet, agent extensions, and custom agents — not yet on first-party ChatGPT or Claude.) ### Example 1 — Product search User in Comet: _"I'm looking for a wireless dog fence for a 2-acre yard, around $300."_ 1. Comet navigates to a pet-supply site that has registered a `searchProducts` tool. 2. Comet reads the tool schema: inputs are `query`, `category`, `maxPrice`, `features`; output is an array of products with name, URL, price, and short description. 3. Comet calls `searchProducts({ query: "wireless dog fence", maxPrice: 300, features: ["2-acre range"] })`. 4. The browser shows: "Comet wants to search products on petsupplies.com. Allow?" User clicks allow. 5. The tool runs your normal product search code, returns three results. 6. Comet renders the three results inline in its chat, with your product names and links. That's a faster, better path than "the agent scraped your category page and guessed." You control the ranking, the description, and what the agent shows. ### Example 2 — Booking an appointment User in a Chrome agent extension: _"Find me a roof inspection appointment in Dallas next Tuesday morning."_ 1. The agent lands on a roofing company's site with a registered `getAvailableSlots` and `bookAppointment` tool pair. 2. The agent calls `getAvailableSlots({ city: "Dallas", date: "2026-06-09", timeOfDay: "morning" })`. Browser confirms. Tool returns three slots. 3. Agent tells the user the three options. User picks one. 4. Agent calls `bookAppointment({ slot, name, phone, email })`. Browser confirms with the user, showing the details to be submitted. 5. The tool runs the actual booking transaction. Returns a confirmation number. For local-service businesses, this is the kind of flow that will eventually move from "the agent surfaces a phone number" to "the agent books the appointment." The infrastructure to do it cleanly exists; the consumer-agent invocation that drives volume is still catching up. ### Example 3 — Lead capture / quote request User in a custom enterprise buying agent: _"Get me a quote for a 1,500 sqft kitchen remodel in Plano."_ 1. Agent lands on a remodeler's site with a `requestQuote` tool registered. 2. Agent calls `requestQuote({ projectType: "kitchen remodel", squareFeet: 1500, location: "Plano, TX", name, email, phone })`. 3. Browser confirms with the user, who reviews the data being submitted. 4. Tool runs your existing lead-capture logic — writes to CRM, fires email, triggers Slack notification. The agent did the work the user would have done by filling out a form. The form code on the page didn't change. ## The safety model: who gets to do what The reason WebMCP can be shippable in a browser without a thousand abuse vectors is the consent model. Three things hold the line: ### 1\. In-browser confirmation, not in-page The page cannot draw its own consent dialog. The browser renders the confirmation — same trust surface as a permission prompt for location or camera. A malicious page can't fake an "allow" click. ### 2\. Per-invocation, not blanket consent Approval is per-tool-per-action by default. Users can mark a specific tool as "always allow on this site" but that's a deliberate setting, not a one-time-blanket-OK. The default is: every call shows a prompt. ### 3\. Payment and authentication are carved out The spec explicitly forbids tools that take credit-card or password fields. The browser refuses to invoke them. Payment integrations (Stripe, PayPal, Apple Pay, Shop Pay) work by handing the agent a "ready-to-pay" URL that the user has to click through. The agent assembles the cart; the human authorizes payment. That carve-out is what makes "agentic commerce" not terrifying. The agent shops, the user buys. ## Setting it up: one script tag, the tool packs You can write WebMCP integrations from scratch using the raw `navigator.modelContext.registerTool()` API. For most sites, that's not the right play — you'd be writing the same five or six tools (search, add-to-cart, checkout-handoff, book-appointment, request-quote, lookup-order) that every other site is writing. The cleaner pattern is a tool pack: a hosted snippet that ships a library of common tools, each backed by a config block where you wire it to your existing site code or platform API. One tag, the tools register, the WebMCP-aware agents that visit your page can invoke them. For ecommerce on Shopify, BigCommerce, or WooCommerce, the snippet auto-wires search, cart, and order-lookup to your platform APIs without any custom code. For service businesses on Calendly, Acuity, or HubSpot, booking and lead-capture wire to those. For custom apps, you point each tool at the function or endpoint it should call. Crawlytics' Commerce tier ($49.99/mo) ships exactly this pattern — a single tag, a config block, and a dashboard that shows which agents invoked which tools when. The math: integrating yourself is doable, but the spec is still evolving and rolling your own means tracking those changes manually. ## What this means for conversion attribution When a WebMCP-aware agent takes an action on your site, you want to know which agent, which session, and whether the action converted. Otherwise the channel is a black box. The emerging attribution pattern: - Tool invocations carry an agent identifier in their metadata (where the implementation supports it — Comet exposes it; many extensions do too) - You log the tool call, the agent, the user session, and the downstream conversion (purchase, booking, lead) - You attribute conversions back to the agent the same way you'd attribute to a paid channel — by source, by campaign equivalent This connects back to the broader AI-attribution problem. ChatGPT, Claude, and Perplexity in-app browsers already strip the Referer header on outbound clicks — most sites are losing AI referral attribution to "(direct)" in Google Analytics. We covered the fix in [our piece on ChatGPT direct traffic](https://crawlytics.app/c/claude/blog/chatgpt-direct-traffic-fix). WebMCP attribution is the same problem at a different layer. ## Does WebMCP replace llms.txt? No, and the order of investment matters: **ship `llms.txt` first**. The audience that benefits from `llms.txt` — every AI client that fetches your pages — is much larger today than the audience that invokes WebMCP. WebMCP is the next layer on top. `llms.txt` tells an agent _what_ is on your site — the catalog, the order, the descriptions. WebMCP tells the agent _what it can do_ on your site — the actions, the inputs, the outputs. A WebMCP-aware agent picking between two sites will use `llms.txt` to read them and WebMCP to act on the one that lets it complete the user's task. If you have `llms.txt` but not WebMCP, the agent reads you and refers the user back to manual action. If you have both, the agent reads you and (if it's a WebMCP-aware one) completes the task. [The llms.txt setup guide is here](https://crawlytics.app/c/claude/blog/what-is-llms-txt-guide) — ship that first, then come back for WebMCP. ## Where this leaves you WebMCP is the next layer of the AI-readiness stack — the layer where agents stop referring users and start completing tasks. But the realistic 2026 picture is that adoption is in early-prototype phase. Today's WebMCP invokers are a small set: Comet, agent extensions, custom buying agents. The major consumer agents (ChatGPT, Claude in their first-party apps) use other approaches today and may take 6-12 months or more to add WebMCP support. So the honest framing: adding the snippet is a forward investment, not a 2026 conversion engine. The integration is small. The upside compounds as adoption grows. The downside is zero — on browsers without WebMCP support, the registration call no-ops and the page renders normally. If you're already shipping `llms.txt`, WebMCP is the natural next layer. If you're not, ship that first. ## Related Written by Crawlytics Team. Crawlytics tracks AI bots, generates llms.txt, and powers WebMCP commerce, all from one snippet on any stack. [See how it works →](https://crawlytics.app/c/claude/) ## Frequently Asked Questions ### Does WebMCP replace llms.txt? No, and the order of investment matters: ship llms.txt first. The audience that benefits from llms.txt — every AI client that fetches your pages — is much larger today than the audience that invokes WebMCP. WebMCP is the next layer on top. llms.txt tells an agent what is on your site — the catalog, the order, the descriptions. WebMCP tells the agent what it can do on your site — the actions, the inputs, the outputs. A WebMCP-aware agent picking between two sites will use llms.txt to read them and WebMCP to act on the one that lets it complete the user's task. If you have llms.txt but not WebMCP, the agent reads you and refers the user back to manual action. If you have both, the agent reads you and (if it's a WebMCP-aware one) completes the task. The llms.txt setup guide is here — ship that first, then come back for WebMCP. --- title: "How to Track AI Citations (ChatGPT, Claude, Perplexity) 2026" type: [Organization, Article, BreadcrumbList, WebSite] author: Crawlytics Team publisher: Crawlytics datePublished: 2026-06-11 dateModified: 2026-06-11 canonical: https://crawlytics.app/c/claude/blog/how-to-track-ai-citations category: blog wordCount: 2204 readingTime: 11 min crawledAt: 2026-06-21 16:40:19 lastVerified: 2026-08-25 13:09:22 site: https://crawlytics.app/c/claude/ --- # How to Track AI Citations (ChatGPT, Claude, Perplexity) 2026 ## Summary Server logs show which AI bots fetched your pages; prompt-testing shows which answers cite you. Practical playbook with bot UAs, grep commands, and benchmarks. ## Key facts - First, the uncomfortable part. - Before any tooling, separate the two questions, because they have different answers and different fixes. - Pull the last 30 days of access logs. - Logs tell you what's being fetched. - Once you have the data, the next question is whether what you're seeing is good. The most common question I get from marketing teams in 2026 is some version of: "Is ChatGPT citing my site?" The answer is usually disappointing — not because the data isn't there, but because most teams are looking in the wrong place. Google Analytics will not tell you. Your CDN dashboard probably won't either. The data exists in your raw server logs and in the AI tools themselves, but you have to know what to look for. This is the practical playbook. Four detection steps, the User-Agent strings to grep for, a way to prompt-test your own brand, and the benchmarks that tell you whether what you're seeing is good, average, or a warning sign. ## The attribution gap: why your analytics can't see this First, the uncomfortable part. When someone discovers you through ChatGPT and visits your site later, the visit almost never carries an AI fingerprint. It lands in GA as **(direct) / (none)**, or as a branded Google search three days later when the buyer types your name to find pricing. The discovery happened inside an AI answer; the analytics record says otherwise. Zero-click answers are worse. A buyer asks Perplexity to compare five tools in your category, reads the synthesis, shortlists you, and never clicks anything. You influenced a deal and generated zero rows in any analytics table. Sales teams call this the dark funnel, and AI assistants are pumping more of the buying journey into it every quarter. So when teams try to track AI search visibility through referral reports and conclude nothing is happening, they're usually wrong. The influence is there. It's wearing a disguise — direct visits, branded searches, "a colleague mentioned you" — and the job is to find signals that don't depend on a referrer header. ### Four proxy signals (and the layer they miss) Casey Nifong made this case well in a June 2026 Search Engine Land piece on [tracking AI search visibility when attribution falls short](https://searchengineland.com/track-ai-search-visibility-attribution-falls-short-479510). Her argument: no single metric explains AI-driven influence, so you triangulate across four signals. Assisted conversions, branded search growth, direct traffic trends, and brand visibility inside the AI systems themselves. All four are sound, and the first three are exactly where AI's invisible influence leaks back into measurable data. (The fourth is usually measured by prompt sampling, which has real limits — more on that in our piece on [why AI share of voice is a made-up number](https://crawlytics.app/c/claude/blog/ai-share-of-voice).) What the article doesn't cover is the one dataset that records AI activity directly instead of by proxy: your server logs. Every `ChatGPT-User` or `Perplexity-User` hit is a timestamped, page-level record of an assistant pulling your content for a live answer. The four proxy signals tell you _something_ is happening. The logs tell you which assistant, which page, and when. That pairing is the practical move. Treat AI bot crawl spikes as your leading indicator and branded search lift as the lagging one. If Claude-User fetches of your comparison page triple in March and branded search impressions climb in April, you've connected crawl data to business impact without a referrer header in sight. Watch the two lines together for a quarter and the lag between them becomes your attribution model. ## The two questions: fetching vs citing Before any tooling, separate the two questions, because they have different answers and different fixes. **Question 1: Is AI fetching my pages?** This is a server-side question. AI assistants have crawlers that visit your URLs, parse the content, and return it (or a summary) to whoever asked. Your access logs show every fetch. If you don't see fetches from named AI bots, the agent doesn't have your content — full stop. **Question 2: Is AI citing me in answers?** This is a client-side question. Even if AI is fetching you, the model may or may not surface your URL when answering a user's question. Citation is a separate event from retrieval. You measure it by asking the AI a question and seeing if you show up. The two failure modes are different. If you're not being fetched, the fix is technical — robots.txt, llms.txt, agent-accessibility. If you're being fetched but not cited, the fix is editorial — your content isn't answering the question well enough, or competitors are answering it more clearly. ## Step 1 — Server log signals: what to grep for Pull the last 30 days of access logs. The format varies (Apache, nginx, Cloudflare, Vercel, Netlify) but every one of them records the User-Agent. Here's the bot taxonomy you should be grepping for: | UA pattern | Who | What it means | | --- | --- | --- | | GPTBot | OpenAI | Training crawler. Fetches pages to potentially include in model training. Not a real-time answer signal. | | ChatGPT-User | OpenAI | Live fetch. Triggered when a ChatGPT user asks a question and the model decides to browse your URL. | | OAI-SearchBot | OpenAI | ChatGPT Search index crawler. Real-time-ish — populates the in-product web index. | | ClaudeBot | Anthropic | Training crawler. | | Claude-User | Anthropic | Live fetch. Claude is browsing your URL on behalf of a user prompt. | | Claude-SearchBot | Anthropic | Claude Search index crawler. | | PerplexityBot | Perplexity | Index crawler. | | Perplexity-User | Perplexity | Live fetch on behalf of a user. | | Google-Extended | Google | Gemini training crawler. Separate UA from Googlebot so you can opt out of AI training without losing Search. | | Bytespider | ByteDance | Doubao / Chinese-market crawlers. Often confused for malicious traffic. | | Amazonbot | Amazon | Alexa+ / Rufus crawler. | A nginx one-liner to count the last 30 days of AI-bot hits by UA: ``` grep -E 'GPTBot|ChatGPT-User|OAI-SearchBot|ClaudeBot|Claude-User|PerplexityBot|Perplexity-User|Google-Extended|Bytespider|Amazonbot' /var/log/nginx/access.log* \ | awk -F'"' '{print $6}' \ | sed -E 's/.*(GPTBot|ChatGPT-User|OAI-SearchBot|ClaudeBot|Claude-User|PerplexityBot|Perplexity-User|Google-Extended|Bytespider|Amazonbot).*/\1/' \ | sort | uniq -c | sort -rn ``` What you want to see: a healthy mix, with the User-suffixed bots (ChatGPT-User, Claude-User, Perplexity-User) showing up at all. Those are the ones tied to real user prompts in real time. If you only see training crawlers (GPTBot, ClaudeBot) but never the User variants, you're indexed but not being browsed. ### The page-level rollup Counts by bot are useful, but the more actionable view is bot-by-page. Which of your pages are AI assistants actually fetching? Sort the request paths by bot fetch count and look at the top 20. Common patterns: - **Your docs/getting-started page** dominates if you publish a developer product. Agents fetch it as a primer. - **Comparison and "vs" posts** get heavy live-fetch traffic — agents lean on them for "should I use X or Y?" questions. - **Pricing pages** get fetched when users ask cost-related questions. - **Glossary or definition pages** get fetched for "what is X?" prompts. If the top pages don't match the pages you want surfaced in AI answers, that's a content gap, not a tracking gap. Write the page that the agent is asking for. ## Step 2 — Prompt-test your own brand Logs tell you what's being fetched. They don't tell you whether you're being cited. For that, you have to be the user. Open ChatGPT, Claude, and Perplexity (signed-out incognito sessions for each — your account history biases results). Run a battery of prompts a buyer in your category would actually type. Record which sources are cited in the response, and where you appear. A working test set has three tiers: 1. **Branded prompts.** "Tell me about ." If you don't appear here, you have a foundational problem — likely no `llms.txt`, robots.txt blocking AI bots, or fresh-domain trust issues. 2. **Category prompts.** "What's the best tool for ?" "How do I ?" This is the real visibility test — you against competitors. 3. **Long-tail / pain prompts.** "Why is my ?" These are the highest-converting prompts because the asker is mid-buying-decision. Score each prompt: **cited** (your URL appears in the source list), **mentioned** (your brand name appears in the answer but no URL), or **absent**. Track this monthly. Three months of data shows whether your AI visibility is trending up or down. You can do this by hand for 20-30 prompts in an afternoon. Past that, automate it — there are tools (Profound, Otterly, AI Brand Rank) that run scheduled prompts against each model and chart your appearance over time. ## Step 3 — Per-bot fetch frequency and what "normal" looks like Once you have the data, the next question is whether what you're seeing is good. Here are the benchmarks I see across Crawlytics customers in mid-2026, split by site size: | Site size (monthly human pageviews) | Healthy AI-bot fetches/month | Healthy bot-to-human ratio | | --- | --- | --- | | Under 10k | 200 - 1,000 | 1:30 to 1:50 | | 10k - 100k | 2,000 - 15,000 | 1:25 to 1:50 | | 100k - 1M | 15,000 - 200,000 | 1:10 to 1:30 | | 1M+ | 200,000+ | 1:5 to 1:20 | Two failure modes to watch for: - **Below 1:200 (way too few bot fetches).** Agents are skipping you. Check robots.txt for accidental blocks, confirm `llms.txt` exists, check whether your sitemap.xml is up to date and discoverable. - **Above 1:5 (way too many bot fetches).** You have a crawl-spend or bandwidth problem. If you're not on a flat-rate host, this can show up as a hosting bill. Worth rate-limiting the training crawlers (GPTBot, ClaudeBot, Google-Extended) while keeping the live User variants open. ## Step 4 — Catch the human follow-up Bots fetching your page isn't the end of the funnel. Some percentage of the users who saw your citation in ChatGPT will click through and visit your site. That visit is where revenue happens — and it's where most analytics stacks go blind. The reason: ChatGPT, Claude, and Perplexity all open citations in an in-app browser that strips the `Referer` header. Google Analytics sees the visit as **(direct) / (none)** and you have no idea it came from AI. We covered the full mechanics and the fix in our piece on [why ChatGPT traffic shows as direct in Google Analytics](https://crawlytics.app/c/claude/blog/chatgpt-direct-traffic-fix) — the short version is you have to inject UTM parameters into the URLs that AI assistants fetch, before they fetch them. Detection and attribution work together. Detection tells you whether you're showing up. Attribution tells you whether the visits convert. Without both, you're flying half-blind. ## What "good" looks like across three site sizes Three real-shape benchmarks from Crawlytics customer cohorts (anonymized): ### Small SaaS marketing site (40 pages, 8k monthly visits) - ~600 AI-bot fetches/month — split 50% ChatGPT, 25% Claude, 15% Perplexity, 10% other - Top-fetched page: the comparison post against the category leader - Brand prompt appearance: cited in 60% of branded ChatGPT prompts, 40% of branded Claude prompts - Category prompt appearance: cited in 8% — there's headroom ### Mid-size docs site (400 pages, 80k monthly visits) - ~9,000 AI-bot fetches/month - Top-fetched page: Getting Started, followed by the four most-popular concept docs - Brand prompt appearance: cited 95% of the time across all three engines - Category prompt appearance: cited in 25% — strong - llms-full.txt fetches: ~400/month (code agents pulling the whole reference) ### Enterprise ecommerce (15k SKU pages, 800k monthly visits) - ~120,000 AI-bot fetches/month, concentrated on category and "best X for Y" pages - Bot-to-human ratio sits at 1:6.5 — getting heavy. Training bots rate-limited; live-fetch bots untouched - Category prompt appearance: cited in 35% of "best \[product type\]" prompts - Detected at least one new AI client every quarter that wasn't on the radar last year None of these are "industry averages" — your mileage will vary. They're shapes to compare against. If yours is dramatically lower at a given site size, dig in. ## When grep-the-logs stops scaling The whole playbook above works without any paid tool for a site under ~10k monthly pageviews. Past that scale, three things break: 1. **Log rotation.** Default rotation is 7-14 days. To track 30-day or 90-day trends you need to archive logs somewhere queryable. That's a S3-plus-Athena project, or a hosted tool. 2. **Per-page rollups.** Counting fetches by bot is one grep. Counting fetches per bot per page per day, with time-series charts, is a database problem. 3. **Citation tracking.** Manual prompt-testing 20 prompts is doable. Running 200 prompts against three engines weekly and charting your share of voice is not. At that point you want a dashboard that does both halves — per-bot fetch counts AND scheduled prompt-tests with citation tracking — and ties them to the same per-page rollup. That's what Crawlytics does. [The Visibility tier ($29.99/mo)](https://crawlytics.app/c/claude/pricing) covers fetch detection plus llms.txt generation; citation tracking is on the roadmap for the next tier. If you're not at that scale yet, the grep-and-prompt-test loop is more than enough. Run it once a month. Track the trend. The day you can't keep up by hand is the day to graduate to a tool. ## Related Written by Crawlytics Team. Crawlytics tracks AI bots, generates llms.txt, and powers WebMCP commerce, all from one snippet on any stack. [See how it works →](https://crawlytics.app/c/claude/) --- title: "Crawlytics vs Profound: AI Brand Visibility Tools Compared (2026)" type: [Organization, Article, BreadcrumbList, WebSite, FAQPage] author: Crawlytics Team publisher: Crawlytics datePublished: 2026-06-05 dateModified: 2026-06-05 canonical: https://crawlytics.app/c/claude/blog/crawlytics-vs-profound category: blog wordCount: 2379 readingTime: 12 min crawledAt: 2026-06-21 16:40:20 lastVerified: 2026-08-25 13:09:23 site: https://crawlytics.app/c/claude/ --- # Crawlytics vs Profound: AI Brand Visibility Tools Compared (2026) ## Summary Profound vs. Crawlytics: Profound is the high-end share-of-voice tool for enterprise; Crawlytics is the technical stack for bot tracking, llms.txt, and WebMCP commerce from $49.99/mo. ## Key facts - Five things, all done well: - The overlap with Profound is exactly one feature: **tracking whether AI assistants cite your URLs. - Five places where Profound is straightforwardly better, and the gap isn't going to close from Crawlytics' end any time soon: - Five places where Crawlytics is straightforwardly better, and where Profound either doesn't compete or hasn't shipped: - Profound publishes pricing only on request. Quick answer Profound is the enterprise share-of-voice dashboard for AI search — it runs hundreds of prompts daily across every major model, charts your visibility against named competitors, and ships executive-ready reports. It starts in the four figures monthly and earns it for the brands it serves. Crawlytics is the technical stack underneath — per-bot server log analytics, llms.txt generation, per-LLM UTM attribution that recovers ChatGPT/Claude/Perplexity referrals from "(direct)" in GA, and WebMCP agentic commerce — at $29.99-$49.99/mo. They overlap in the citation-tracking question but optimize for different customers. **Pick Profound if you're a $100M+ brand that needs share-of-voice reporting. Pick Crawlytics if you're a sub-$50M business that needs bot tracking, llms.txt, attribution recovery, or WebMCP. Run both if you're large enough that both jobs apply — they overlap less than 20%.** This is the comparison post I get asked for most. Profound ([tryprofound.com](https://tryprofound.com/)) is the best-funded, best-marketed entrant in the AI brand visibility space, and they've done a clean job of defining the share-of-voice category. Crawlytics covers a different surface — the technical AI-readiness stack — at a price point that doesn't compare. People want to know which to pick. The honest answer is that for most companies it's not a choice between them. They solve different jobs at different price points for different customers. Below I'll walk through what Profound actually does well, where Crawlytics genuinely wins, where Profound clearly wins, and the decision tree for picking one (or both) by company size. ## What Profound actually does Five things, all done well: 1. **Daily prompt-set automation.** Profound runs hundreds (in enterprise plans, thousands) of prompts every day against ChatGPT, Claude, Perplexity, Gemini, and Copilot. The prompt sets are curated for your category — branded prompts, competitor prompts, category prompts, long-tail buyer-intent prompts. 2. **Share of Voice dashboard.** The headline metric. For each prompt set, you see what percentage of responses cite your brand vs each named competitor. Trended over time. Sliceable by model, by prompt cluster, by geography. 3. **Conversation analytics.** Profound captures the full LLM response text, not just whether you were cited. They surface the language models use to describe you, which adjectives recur, which competitors get co-mentioned. Real qualitative signal an enterprise team can act on. 4. **Source URL tracking.** When ChatGPT cites a URL, Profound logs which URL — so you can see which of your pages are driving citations, and which third-party pages (Reddit threads, news articles, review sites) are influencing your brand mentions. The competitive intel version of backlink analysis. 5. **Executive reporting.** Their reports look like the slides a CMO would put in front of a board. PDF exports, scheduled email digests, a polished "this is your AI visibility score" headline metric that translates into a meeting agenda. That's a real product, well-executed, and the customers I've talked to who use it at scale (large CPG, large fintech, large travel) consistently say it earns its price. The category Profound defined did not exist three years ago and they put a flag in it convincingly. ## What Crawlytics does in the overlap zone The overlap with Profound is exactly one feature: **tracking whether AI assistants cite your URLs.** Crawlytics approaches this from the server side: it tells you which AI bots (ChatGPT-User, Claude-User, Perplexity-User, OAI-SearchBot, Claude-SearchBot) fetched which of your pages, when, how often. Real fetch traffic, real URL-level granularity, real-time, from your access logs. Profound approaches it from the prompt side: it asks the models questions a buyer would ask and records whether your URL shows up in the cited sources. Synthetic prompts, full response text, share-of-voice math, daily cadence. Both signals are valid. They answer different versions of the same question. Profound's prompt-test catches the case where you're "in the index but not cited" — you're being read but not surfaced. Crawlytics' log signal catches the case where you're "cited but not being fetched" — your training-era reputation is carrying you and the models aren't refreshing. You want both signals; you usually don't need to pay for both at the same price tier. ## Where Profound clearly wins Five places where Profound is straightforwardly better, and the gap isn't going to close from Crawlytics' end any time soon: ### 1\. Share of Voice as a dashboard Profound's share-of-voice chart is the artifact a marketing team brings to a stand-up. It compresses "how visible is our brand in AI search" into a single number, trended over time, with named competitor bars next to yours. Crawlytics does not have this view. Per-bot fetch counts are a different question and a different audience. ### 2\. Prompt-set scale and curation Running 500-2,000 curated prompts daily against five models is a real infrastructure problem. Profound has solved it, runs it as a service, and curates prompts by vertical so you don't have to design your own test bench. Crawlytics' citation tracking on the roadmap is going to be smaller-scale and DIY-flavored — 50-200 prompts on a slower cadence. Different product. ### 3\. Enterprise reporting Branded PDF exports, scheduled executive digests, white-glove account management, custom dashboards for the CMO. Profound has built the enterprise-buying surface — and the price reflects it. If you need the report to look like McKinsey wrote it, Profound delivers; Crawlytics produces dashboards designed for the operator, not the C-suite. ### 4\. The deep conversation analytics Capturing the full response text and analyzing the language models use about your brand — "fast, reliable, expensive" vs "innovative, complex, premium" — is a qualitative layer that requires both the prompt automation and an NLP layer on top. Profound ships it. It's a real input to brand and positioning work that no log-analytics tool reproduces. ### 5\. Multi-model breadth Profound tracks ChatGPT, Claude, Perplexity, Gemini, and Copilot in one dashboard. Crawlytics tracks every named AI bot that hits your server (which is broader on the crawler-coverage side) but doesn't yet run synthetic prompts against all five models. For "are we cited in Gemini's answer to this buyer question," Profound is the answer today. ## Where Crawlytics clearly wins Five places where Crawlytics is straightforwardly better, and where Profound either doesn't compete or hasn't shipped: ### 1\. Per-bot server log analytics Profound asks the model what it knows about you. Crawlytics watches your server and reports who fetched what, when. That distinction matters: server-log signal catches every fetch from every bot the moment it happens — not just the prompts in your test set. If a new AI client launches tomorrow and starts crawling your site, Crawlytics shows the traffic on day one. Profound shows it the first time you add a prompt that surfaces the new model. Different latency, different completeness. ### 2\. llms.txt generation Profound does not ship an `llms.txt` file. They're a measurement layer, not a publishing layer. **[llms.txt](https://crawlytics.app/c/claude/features/llms-txt-generator)** — llms-full.txt from your sitemap, re-crawls daily, and serves the files at stable URLs every AI bot knows to look for. The publishing layer is where AI-readiness starts — measurement comes after. ### 3\. Per-LLM UTM attribution When a ChatGPT user clicks your citation, the in-app browser strips the Referer header and Google Analytics logs the visit as **(direct) / (none)**. Profound does not solve this — they measure citations, they don't fix the downstream attribution gap. Crawlytics injects per-LLM UTM tags into the AI-Optimized HTML it serves to bots, so citation clicks arrive at your site with `utm_source=chatgpt` instead of being invisible. [Full mechanics here.](https://crawlytics.app/c/claude/blog/chatgpt-direct-traffic-fix) Different surface entirely. ### 4\. WebMCP agentic commerce Profound does not ship a WebMCP layer. They don't expose tools to in-browser agents, they don't handle cart-assembly, they don't attribute conversions back to the agent that drove them. Crawlytics' Commerce tier ($49.99/mo) does all three. If your business model includes agents being able to _buy_ things on your site — not just cite them — Profound is the wrong tool. [WebMCP explainer here.](https://crawlytics.app/c/claude/blog/webmcp-explained-ai-agent-actions) ### 5\. Price Crawlytics tops out at $49.99/mo. Profound starts in the high three figures and goes up from there. That gap is the single biggest reason to pick Crawlytics for any business under about $50M revenue — at that scale you can't justify a four-figure monthly AI visibility line, but you absolutely can justify $29.99 for the bot tracking and llms.txt. ## Pricing reality check Profound publishes pricing only on request. Based on public mentions and customer disclosures in 2026, the entry tier sits somewhere between $1,000 and $1,500/mo, and enterprise plans for the brands they market to most heavily (large CPG, large tech, large retail) run several thousand per month. The pricing matches the customer — those plans include custom prompt curation, account management, and the executive reporting layer that justifies the cost for a CMO with a $5M+ annual marketing budget. Crawlytics has three published tiers: - **Free** — Agent-Ready Grader, basic bot detection on one site, llms.txt audit. - **Visibility ($29.99/mo)** — Full bot tracking, llms.txt generation, per-LLM UTM attribution, one site. - **Commerce ($49.99/mo)** — Everything in Visibility plus the WebMCP snippet, per-agent conversion attribution, and multi-site dashboards. For a sub-$50M business, the math is decisive. Crawlytics Commerce is roughly 2-5% the cost of Profound's entry tier and covers a strictly different — and in most cases more technically essential — set of jobs. For a $100M+ brand with a real CMO and a board-level AI visibility narrative, Profound's price is reasonable for what it delivers. It's a tool for marketing teams, not a tool for engineering teams. Different buyer, different budget line. ## The decision tree by company size The honest matrix: ### If your annual revenue is under $5M **Crawlytics, almost certainly.** You need bot tracking, llms.txt, and attribution recovery. You don't yet need a share-of-voice dashboard against named competitors — at your size you can hand-run 20 prompts against ChatGPT once a month and get most of the signal Profound provides. The $29.99/mo Visibility tier covers the technical ground. Skip Profound until you've outgrown DIY prompt-testing. ### If your annual revenue is $5M-$50M **Crawlytics, with an honest conversation about whether you need share-of-voice yet.** At this size you might be approaching the point where a Profound-style dashboard pays for itself — particularly if your category is competitive in AI search and your marketing team needs the benchmark to justify content investment. But the technical AI-readiness stack (bot tracking, llms.txt, WebMCP if you sell online) is non-negotiable and Crawlytics delivers it at a price the CFO won't blink at. Start there. Add Profound when the share-of-voice question becomes a quarterly discussion at the executive level. ### If your annual revenue is $50M-$500M **Probably both, sequenced.** Crawlytics for the technical stack — you need llms.txt, bot tracking, attribution, and (if you sell online) WebMCP. Profound for the share-of-voice dashboard your CMO will use in board updates. The combined monthly cost is still negligible against revenue at this size, and the two tools overlap less than 20% — you're getting two different jobs done for two different audiences inside your company. ### If your annual revenue is $500M+ or you're a category leader **Both, definitively.** Profound is the share-of-voice tool you brief the C-suite with. Crawlytics is the engineering-team tool that ships the publishing layer (llms.txt), recovers the attribution your data team is missing in GA, and enables agentic commerce if you have a transactional surface. They are complementary, not competitive, at this scale. ## "Use both" — when it makes sense The case for running both is stronger than people expect. Profound tells you whether you're cited; Crawlytics tells you whether you're being fetched. Profound tells you what share of voice you have against Competitor X in branded prompts; Crawlytics tells you which of your pages those citations are landing on and whether the traffic converts. Profound's qualitative conversation analytics shapes the language you use in your copy; Crawlytics' WebMCP layer captures the conversions that copy drives. The two products are sufficiently non-overlapping that for any brand above $50M revenue running them together costs less than 0.1% of marketing budget and produces strictly more signal than either alone. The mistake is treating them as alternatives when they're complements. ## Related Written by Crawlytics Team. Crawlytics tracks AI bots, generates llms.txt, and powers WebMCP commerce, all from one snippet on any stack. [See how it works →](https://crawlytics.app/c/claude/) ## Frequently Asked Questions ### Is Profound worth the price? For the brands they're built for — $100M+ revenue, CMO-led AI search strategy, board-level reporting requirements — yes. The share-of-voice dashboard, the prompt-set scale, and the executive reporting layer compound into a real input to brand strategy. For smaller brands, the price is hard to justify against what a $29.99/mo log analytics tool plus a monthly hand-run prompt audit can produce. ### Can Crawlytics replace Profound for an enterprise brand? No. Crawlytics does not run automated prompt sets at Profound's scale, does not produce share-of-voice dashboards against named competitors, and does not ship executive reporting in the same form. If those things are your buying criteria, Crawlytics is the wrong tool. Crawlytics covers a different (and at enterprise, complementary) set of jobs. ### Does Profound do bot tracking? Not in the per-server-log sense. Profound is a prompt-side measurement tool — it asks the models questions and records the answers. It does not analyze your access logs to tell you which AI bots fetched which pages, and it doesn't generate or audit your llms.txt. For that you need a server-side tool like Crawlytics or a custom log pipeline. ### Which has better ChatGPT citation accuracy? They measure different things. Profound has higher accuracy on "am I cited in synthetic ChatGPT prompts" because that's literally what they measure, repeatedly, with controlled prompt sets. Crawlytics has higher accuracy on "is ChatGPT-User fetching my pages" because that's literally what they measure, from real server logs. Neither is a stand-in for the other. ### Are there cheaper alternatives to both? For prompt-side measurement at the Profound scale: Otterly.ai, AI Brand Rank, and Peec.ai are positioned as more affordable share-of-voice tools, though none have matched Profound's coverage and reporting depth as of mid-2026. For the technical stack Crawlytics covers (bot tracking, llms.txt, attribution, WebMCP), the alternative is DIY — grep your access logs, hand-write your llms.txt, build your own UTM injection, write your own WebMCP integration. That's a real option for engineering-heavy teams and a multi-week project for most others. --- title: "What Schema Markup Still Matters in the AI Search Era" type: [Organization, Article, BreadcrumbList, WebSite, FAQPage] author: Crawlytics Team publisher: Crawlytics datePublished: 2026-06-05 dateModified: 2026-06-05 canonical: https://crawlytics.app/c/claude/blog/schema-markup-ai-search category: blog wordCount: 2347 readingTime: 12 min crawledAt: 2026-06-21 16:40:24 lastVerified: 2026-08-25 13:09:25 site: https://crawlytics.app/c/claude/ --- # What Schema Markup Still Matters in the AI Search Era ## Summary Most schema is noise to LLMs. Four types still earn their keep: Article, FAQPage, Organization, BreadcrumbList. What to ship, what to skip, and why. ## Key facts - Here's what's actually happening under the hood. - Article schema is the highest-leverage type to ship on every blog post and editorial page. - A short tour of the schema types you can safely skip if AI search is your priority. - You don't need a full schema audit tool to know what you're shipping. - A few practical traps that come up repeatedly: The most honest sentence I can write about schema markup in 2026 is this: most of what you've been told to ship doesn't matter for AI search. ChatGPT does not parse your `HowTo` schema. Claude does not care about your `Speakable` markup. The detailed `Product` blocks SEO consultants spent the late 2010s telling you to add are doing nothing for your citation rate in AI answers. That doesn't mean schema is dead. A short list still moves the needle — for both Google rich results and the small piece of AI retrieval that does pick it up. This guide separates the schema worth shipping from the schema you can quietly remove, and explains why the line falls where it does. ## The honest answer — most schema is noise to LLMs Here's what's actually happening under the hood. When ChatGPT, Claude, or Perplexity fetches your page, the retrieval pipeline does one of two things: it renders the page and extracts the visible text, or it fetches a clean markdown version (via `llms.txt`, content negotiation, or a built-in HTML-to-markdown converter). In neither path does the standard JSON-LD blob get parsed as Schema.org structured data. The retrieval scorers look at headings, paragraphs, lists, tables, and entity mentions. They don't query `@type: HowTo` to decide whether to cite you. That's a Google rich-result behavior, not an LLM behavior. When people insist schema is "critical for AI search," they're usually conflating Google AI Overviews (which still leans on the classic Google index) with the standalone AI assistants (which mostly don't). There's one nuance. Some schema content — names, descriptions, dates — does end up parsed because it's also present in visible page text or in meta tags the LLM reads. The _signal_ survives; the schema container does not. That's why the four schema types below still matter: not because LLMs parse the JSON-LD directly, but because the data they contain ends up where LLMs can read it, and because Google's rich-result coverage compounds the benefit. ## The four schema types that still matter ### Article (with dateModified) Article schema is the highest-leverage type to ship on every blog post and editorial page. Google uses it for the "Top Stories" carousel, article cards in AI Overviews, and date-stamping in search results. The single most important property is `dateModified` — it's the signal that tells both Google and downstream LLM retrieval that the content is fresh. A minimum-viable Article block: ``` ``` Update `dateModified` every time you retrofit the post. Google sometimes shows this date in SERPs, and fresher dates correlate with higher click-through. ChatGPT and Perplexity also pick up date freshness from the visible page header — but the schema acts as a fallback when the visible date isn't crawlable. ### FAQPage FAQPage schema is the second-highest-leverage type. Google still expands FAQ snippets in some verticals, and the Q/A structure happens to mirror exactly what AI retrieval scorers love — discrete, chunkable answers to discrete questions. Even though the LLM doesn't read the JSON-LD, the visible Q/A section it duplicates is the single most-cited part of most posts. The trick is to mirror your visible FAQ section exactly. Don't ship FAQPage schema with questions that aren't visible on the page (Google deprecated that pattern in 2023 and may flag it as deceptive). The schema's job is to reinforce what's already in the rendered DOM. If your post has a "Common questions" H2 with 3-5 H3 questions and paragraph answers, ship the matching FAQPage block. That's all the lift this type can give you. ### Organization (especially sameAs) Organization schema is the entity layer. It tells search engines and any LLM retrieval pipeline that scrapes it (Bing's certainly does, Google's does, OpenAI's index appears to ingest it inconsistently) who you are, what you're called, and which other web properties you own. The single highest-value property in Organization is `sameAs`, the array of canonical URLs for your brand on other platforms — your LinkedIn, X, GitHub, Crunchbase, Wikipedia (if applicable), and any other authoritative profile. This is the entity-disambiguation signal that helps an AI engine resolve "Crawlytics" to a specific company rather than a generic term. Ship Organization schema once, at the site root (typically in your homepage layout or a shared `BaseLayout`). Get `sameAs` right and you get cumulative benefit across every page that inherits the schema. ### BreadcrumbList BreadcrumbList is the least sexy of the four but it punches above its weight. It tells search engines (and any LLM that parses it) how a page fits in your site's hierarchy — useful for context, useful for showing breadcrumb trails in Google SERPs, useful for the rare AI engine that uses page position as a relevance signal. It also costs almost nothing to ship. If your site has any nested structure (blog/post-slug, features/feature-slug, resources/resource-slug), generate BreadcrumbList per page from the URL path. Ten lines of template code. One-time setup, perpetual benefit. ## Schema types that don't matter for AI (and why people still ship them) A short tour of the schema types you can safely skip if AI search is your priority. None of these will hurt you; most are just a waste of template-author time. - **Speakable.** Designed for voice-assistant readouts. Largely retired by Google and never picked up by Alexa, Siri, or Google Assistant in any meaningful way. LLMs do not use it. Remove it from your templates — no one is reading your articles via voice in 2026, and the schema is the cargo-cult artifact of a moment that didn't materialize. - **HowTo.** Google deprecated HowTo rich results in late 2023 for non-mobile, and tightened restrictions further in 2024. LLMs do not parse it. The structure was useful in 2018 when "how to" rich results were a real SERP feature; it has been functionally dead since 2024. Removing it tightens your JSON-LD blob and improves page weight by a few kilobytes per page. - **Recipe.** Useful if you run a recipe site that wants to appear in Google's recipe carousel — which is still alive and meaningful for food publishers. Useless for everyone else. ChatGPT will summarize a recipe page from the visible text just fine without it. - **Product.** Important for Google Shopping rich results, and worth shipping if Shopping is a meaningful channel for you. Less important for AI search — when a user asks ChatGPT about a product, the model leans on the visible page content (price, features, reviews) rather than the JSON-LD. Ship Product schema if you're commerce; don't expect it to lift AI citations. - **Event, Course, JobPosting.** Same pattern — useful for the specific Google rich result, irrelevant to LLM retrieval. Ship if you need the Google feature; skip if you don't. - **VideoObject.** Useful for YouTube and Google Video. Not parsed by LLMs. If you embed video heavily, ship it for Google. Otherwise, skip. The pattern across all of these: the schema type was designed for a specific Google rich-result feature. If you care about that feature, ship the schema. If you only care about AI citations, the schema is a no-op. ## The new thing LLMs DO use that isn't classic schema Here's the substitution. The signal that classic schema was supposed to provide — "this page is structured, here's what it contains" — is now provided in AI search by clean markdown delivery. `llms.txt` tells the AI client what your site contains and where to find it. `llms-full.txt` bundles the actual content. Content-negotiated markdown rendering (or a `.md` companion route) lets the AI fetch a clean version of any specific page in one request instead of scraping HTML. These are not schema in the Schema.org sense. They're a parallel content-delivery layer that gives AI clients the same context that schema was supposed to give Google. The [llms.txt setup guide](https://crawlytics.app/c/claude/blog/what-is-llms-txt-guide) covers the mechanics. The decision rule is the same as for the surviving schema types: ship it if you care about AI citations, skip it if you don't. If you're sequencing the work: ship `llms.txt` first (it's the bigger AI lift), then ship the four schema types listed above (small lift but easy, and you get Google rich results as a side benefit), then audit and remove the deprecated schema types from your templates. ## How to audit your existing schema in 5 minutes You don't need a full schema audit tool to know what you're shipping. The quickest path: 1. **View source on a top page.** Search for `application/ld+json`. Count the blocks. Each one is a schema object. 2. **For each block, check the `@type`.** Anything in the "keep" list (Article, FAQPage, Organization, BreadcrumbList) stays. Anything else gets a question mark. 3. **Validate the keepers with Google's Rich Results Test.** Paste the URL, see whether the schema parses without errors. Fix the errors — schema that throws warnings often throws away the benefit entirely. 4. **Confirm visible content mirrors the schema.** If your FAQPage block has a question that isn't in the visible Q/A section, either add the question to the page or remove it from the schema. Mismatches are penalty risks. 5. **For each deprecated type, decide.** Remove if you're not using the corresponding Google feature; keep if you are. Don't keep them "just in case" — they bloat the template and add maintenance surface. Five minutes per page. Most sites end up removing 1-2 schema types and tightening 1-2 keepers. The result is a smaller, more focused schema footprint that actually helps where it can. ## Implementation gotchas A few practical traps that come up repeatedly: - **JSON-LD beats Microdata.** If you have a choice, ship JSON-LD in a ` ``` Save the theme. Open your storefront in a browser with WebMCP support (Chromium-based browsers with the flag enabled, or an agent-first browser like Comet). You can verify the tools are live in the DevTools console: ``` navigator.modelContext.getRegisteredTools() // → [{ name: 'searchProducts', ... }, { name: 'addToCart', ... }, ...] ``` That's the entire install for a default Shopify store. The snippet auto-discovers your Storefront API schema, pulls your active product types, wires the search predicates to your tag taxonomy, and registers the tools. When a WebMCP-aware agent visits, your store is actionable to it. ## What the snippet auto-wires for Shopify The Shopify-flavored snippet does a few things you'd otherwise hand-write: - **Storefront API query construction.** When an agent calls `searchProducts({ query: "running shoes size 11", maxPrice: 150 })`, the snippet builds the right GraphQL query with `product_type`, `variants.price`, and `variants.option2` predicates against your shop. You don't write the GraphQL. - **Variant resolution.** Agents speak in human terms ("size 11"). The snippet maps that to the right variant ID using your store's option labels — works across `Size`, `Color`, `Material`, `Style` (the four standard Shopify option names), plus any custom ones. - **Cart persistence.** The snippet uses the same cart token Shopify themes use, so if an agent builds a cart and the user later visits the storefront manually, the cart is already there. No double-cart bug. - **Customer-scoped order lookup.** If the user is logged in (Shopify customer account session), `lookupOrder` uses their access token. If not, the tool returns a friendly "log in to check order status" error the agent can relay. - **Currency + market handling.** If you have Shopify Markets enabled and the user is in a region with a different currency, the snippet uses the buyer-context query so returned prices match what the customer would see at checkout. None of that requires custom Liquid. If your theme is stock Dawn, Sense, Studio, or any of the other free themes — or any of the major paid themes (Impulse, Prestige, Motion, Symmetry) — the default tool pack works without modification. ## When you do need to write a custom tool handler Three patterns where the default snippet isn't enough: ### Custom product configurators If you sell custom-printed apparel, configured furniture, or build-your-own subscription boxes, the agent needs a `buildConfiguration` tool that knows your option tree and returns a quote. You write that one tool handler, hook it into your existing configurator state, and register it alongside the defaults. ### Subscription products (Shopify Subscriptions / Recharge / Bold) Default `addToCart` doesn't know about selling plans. For subscription SKUs you either pass a `sellingPlanId` through `addToCart` (the snippet supports this if your products have selling plans defined) or register a separate `subscribeToProduct` tool that wraps your subscription app's API. ### B2B / wholesale catalogs If you run a B2B catalog with company-specific pricing and net-terms checkout, the default Storefront API query misses your customer pricing. You override `searchProducts` with a handler that queries the Shopify Plus B2B endpoints scoped to the logged-in company. In each case you're writing 30-60 lines of tool handler, not rebuilding the integration from scratch. The base snippet handles registration, schema, the agent confirmation flow, and the attribution beacon — you supply the function body for the custom tool. ## The Shopify Payments carve-out The single most-asked Shopify WebMCP question: "can the agent complete checkout without my customer involved?" The answer is no, and the reason is structural. The WebMCP spec explicitly forbids agents from entering credit card details or typing passwords. The browser enforces this at the consent layer — when a tool's schema includes a `cardNumber` or `password` field, the browser refuses to invoke it. So even a malicious site that tried to register `completeCheckout({ cardNumber, cvv })` would have the call blocked. The Shopify pattern works around the limitation cleanly: the agent assembles the cart, calls `checkoutHandoff()`, gets back a `checkoutUrl` like `https://your-store.myshopify.com/checkouts/cn/abc123`. The agent surfaces that URL to the user. The user clicks. Shopify Payments runs the real checkout — same Apple Pay button, same Shop Pay button, same address autofill — and the user authorizes payment. No PCI scope, no card data through the agent. For Shop Pay specifically, the flow is even tighter: customers who've previously authorized Shop Pay see a one-tap confirm instead of a full form. That's the agent-friendly checkout the spec is implicitly designed around. ## Conversion attribution — which agent drove which sale Shopify Analytics will tell you that a sale came from "Direct" or "Other" when an agent drove it. That's because the agent's session has no UTM, no Referer, and no Shopify Sales Channel ID — it's anonymous from Shopify's point of view. The fix is to capture the agent identifier at the tool layer, where you actually have it. WebMCP invocations carry agent identity in metadata where the implementation exposes it — Comet does; many extensions do too. You log that on the tool call, persist it through the cart token, and attribute the eventual order to the agent that started the journey. What that gets you, in practice (calibrated to the realistic 2026 volume — small but trackable): - Sales by agent — which WebMCP-aware clients are actually converting on your store, which are window-shopping. - Cart-add to purchase rate by agent — useful for spotting when an agent's checkout flow is breaking (e.g. an agent that adds 80% of the time but never reaches checkout suggests a handoff bug). - Top collections by agent traffic — which categories agents lean on you for vs your competitors. - A baseline you can compare against when consumer agents (ChatGPT, Claude) eventually add WebMCP — you'll already have the dashboard, the data shape, and a sense of normal. None of that is in Shopify's native reports. Any WebMCP-aware analytics layer can capture it; Crawlytics' Commerce tier ships it as a default dashboard. ## The Shopify Plus checkout extensibility note If you're on Shopify Plus and you've customized your checkout with Checkout Extensibility, double-check one thing after installing WebMCP: that your custom checkout extensions still fire on the agent-handed-off URL. They should — the `checkoutUrl` from `cart.checkoutUrl` uses the same checkout pipeline as a normal cart-to-checkout transition — but if you've got conditional logic gated on a specific session attribute or marketing source, an agent-driven cart may not match that attribute. Quick test: open your store in a WebMCP-capable browser, manually trigger an agent-style cart build via the DevTools console (`navigator.modelContext.invokeTool('addToCart', { ... })`), grab the checkoutUrl, walk through it, confirm your extensions render. If they do, you're done. If they don't, the fix is usually adding a fallback condition that catches the agent-source attribute the WebMCP layer sets. ## Related Written by Crawlytics Team. Crawlytics tracks AI bots, generates llms.txt, and powers WebMCP commerce, all from one snippet on any stack. [See how it works →](https://crawlytics.app/c/claude/) ## Frequently Asked Questions ### Does WebMCP work on Shopify Basic? Yes. All Shopify plans have access to the Storefront API, which is the only API the standard tool pack needs. No Shopify Plus required for the install. ### Will WebMCP slow down my Shopify storefront? The snippet is ~12KB gzipped and loads async. It does not block render and does not run any code until an agent actually invokes a tool. Real-user perf impact is negligible. ### Can AI agents complete checkout on Shopify without my customer's involvement? No, and this is by design. The WebMCP spec carves out payment and authentication — agents cannot enter card details. The Shopify pattern is agent assembles cart → agent hands off URL → human pays. Shop Pay one-tap is the closest thing to "agent buys it for you" and even that requires the customer's prior Shop Pay authorization. ### Does WebMCP work with Shopify Markets (multi-region)? Yes. The default snippet uses Shopify's buyer-context query, so prices, currencies, and product availability returned to the agent match what the customer would see based on their region. If you have Markets-specific catalogs (different SKUs per region), the snippet respects that automatically. ### Which agents will actually invoke my WebMCP tools today? Today, primarily: Perplexity Comet, browser extensions with built-in agents, and custom enterprise buying agents. ChatGPT and Claude's first-party apps don't currently invoke WebMCP — they use citation rendering or screen-control. The mainstream consumer agent rollout is the 6-12 month bet you're making by installing now. If you need conversion volume today, prioritize llms.txt and earning AI citations first; ship WebMCP as the next layer. ### How do I see which agent drove a sale in Shopify analytics? You don't, natively — Shopify Analytics aggregates agent traffic into Direct/Other. You need a layer above that captures the agent ID at the WebMCP tool call and persists it through the cart. The free DIY version: write a small app that listens for cart-create events, tags the cart with the agent metadata, and writes order-level notes you can filter in reports. The paid version: Crawlytics' Commerce dashboard does it as a default chart. --- title: "Default-Deny AI Crawlers: Why Reuters and Publishers Are Switching" type: [Organization, Article, BreadcrumbList, WebSite, FAQPage] author: Crawlytics Team publisher: Crawlytics datePublished: 2026-06-10 dateModified: 2026-06-10 canonical: https://crawlytics.app/c/claude/blog/default-deny-ai-crawlers category: blog wordCount: 2204 readingTime: 11 min crawledAt: 2026-06-21 16:40:20 lastVerified: 2026-08-25 13:09:23 site: https://crawlytics.app/c/claude/ --- # Default-Deny AI Crawlers: Why Reuters and Publishers Are Switching ## Summary Reuters, Time, and People Inc. are switching robots.txt from a blocklist to an allowlist. What default-deny means, why blocklists are failing, and what to do if you're not a publisher. ## Key facts - Allowlisting raises an obvious question: how do you decide who gets in? - Here's the catch nobody at the IAB event had to say out loud. - You don't need Reuters' clout to borrow Reuters' discipline. - Default-deny is the right instinct because the math of the open web changed. - Written by Crawlytics Team. At a late-May IAB Tech Lab event, Lindsay Van Kirk, SVP of Innovation at People Inc., gave a number that reframes the entire AI-crawler debate. When her team switched from blocking bots by name to allowing only a short approved list, the count of blocked user agents went from about 2,100 to more than 30,000. Nothing about the open web changed that week. The 28,000-bot gap had been crawling People's titles all along, and the old blocklist simply never knew their names. That gap is why Reuters, Time, People Inc., and a lengthening list of publishers are rewriting the most boring file on their servers. They're moving robots.txt from "block what you recognize" to "allow only what you approve." If you manage a site and you've been maintaining a list of bad bots to block, this is the shift you need to understand, because the list approach you're using is the one these publishers just abandoned. ## What "default-deny" actually means A blocklist robots.txt names the crawlers you want to keep out and lets everything else through. An allowlist robots.txt inverts that: it names the handful of crawlers you permit and refuses everyone else by default. Default-deny is the security term for that posture. You don't enumerate threats, you enumerate the exceptions and treat the rest of the world as untrusted until proven otherwise. For two decades, the blocklist model worked because the bots that mattered were a known, slow-moving set: Googlebot, Bingbot, a few SEO crawlers. You could name the bad actors because there weren't many. The AI boom broke that assumption. There are now dozens of training crawlers, live-fetch agents, and search indexers across providers, and a new one can show up under a brand-new user-agent string any week. A blocklist is only as good as your last update, and nobody updates robots.txt weekly. Reuters' live robots.txt is the cleanest example of the inverted model. It explicitly allows crawlers from Amazon, Google, Bing/Microsoft, Yahoo, and OpenAI, then disallows other bots across most of the site. Five names in, everyone else out. That file doesn't need to know that 30,000 other agents exist. It refuses them by structure, not by enumeration. ## The number that should scare you: 2,100 to 30,000+ People Inc.'s jump from roughly 2,100 to over 30,000 blocked agents isn't a story about new bots appearing overnight. It's a story about how much the blocklist was missing the whole time. The company didn't suddenly attract 28,000 new crawlers. Those crawlers were already fetching People.com, Travel + Leisure, Food & Wine, and the rest of the portfolio. Switching to an allowlist just made the invisible visible. This is the part that should land for any site owner. Your blocklist isn't a measure of crawler traffic. It's a measure of the crawlers you happened to hear about. The ones you don't name aren't absent, they're unmeasured, and unmeasured bot traffic is exactly the kind that scrapes content, drives up bandwidth bills, and gives nothing back. The People Inc. number is what the gap between the agents you block and the agents that actually visit looks like at scale. ## Why robots.txt was never built for this Robots.txt was published as a convention in 1994. It was designed for a cooperative web where a handful of search engines wanted to be polite about which directories they indexed. It has no authentication, no enforcement, and no way to verify that the bot reading it is the bot it claims to be. Compliance is entirely voluntary. That voluntary model is now the core problem. A Tollbit report found that 30% of total AI bot scrapes didn't comply with the explicit permissions in robots.txt. Nearly a third of AI crawler activity simply ignores the file. Some of that is bad actors spoofing user agents; some is crawlers that read robots.txt and fetch anyway. Either way, a robots.txt rule is a request, not a wall, and a meaningful share of AI traffic treats it as optional. Publishers know this, which is why robots.txt is becoming the policy layer rather than the enforcement layer. The enforcement happens at the CDN or WAF, where you can actually drop a request. The allowlist in robots.txt states the intent clearly enough that downstream tools, licensing negotiations, and legal positions have a documented baseline. The industry is also organizing around it: the publisher-backed SPUR Coalition grew to 36 organizations after adding 30 members in May, aiming to set shared standards for how content gets licensed and used. Regulators are moving too. A new UK conduct requirement forces Google to let sites opt out of AI search features, a sign that opt-out is becoming a right rather than a favor. ## Reuters' "fair value exchange" test Allowlisting raises an obvious question: how do you decide who gets in? Reuters built an explicit rubric. Josh London, head of Reuters Professional, told Digiday that a bot earns access only if it offers a "fair value exchange" across four dimensions: - **Licensing** — does the operator pay to use the content, or have a deal in place? - **Traffic** — does the bot send referral visitors back to the site? - **Uptime** — does its crawl behavior respect the site's stability instead of hammering it? - **Monetization** — does the relationship support the business, directly or indirectly? Run the major crawlers through that filter and Reuters' five-name allowlist makes sense. Google and Bing send search traffic and underpin discovery. Amazon and Yahoo fit existing commercial relationships. OpenAI has been signing licensing deals with publishers, which buys it a seat. A training crawler that pays nothing, sends nothing, and respects nothing fails all four tests, so it doesn't make the list. The framework turns an emotional "block the AI" reaction into a business decision you can defend line by line. ## What smaller sites can't copy from Reuters Here's the catch nobody at the IAB event had to say out loud. Reuters can demand a fair value exchange because Reuters has bargaining power. Its archive is worth licensing, so AI companies negotiate. When you run a 40-page SaaS site, a regional services business, or a personal blog, no one is lining up to license your content, and a hard default-deny can quietly cost you the visibility you actually want. The asymmetry is real. Anthropic's crawler documentation now warns publishers about the visibility trade-off of blocking its search bot: refuse the crawler that feeds AI answers and you opt out of being cited in those answers. For a publisher with a paywall and a licensing team, that trade can be worth it. For a business whose growth depends on being found, blocking the bots that surface you in ChatGPT or Claude is a way to make yourself invisible to the fastest-growing discovery channel on the web. Copying Reuters' robots.txt without Reuters' business model can backfire. The distinction that matters is the one between training crawlers and live-fetch or search crawlers. The training kind takes your content to improve a model and usually sends nothing back. The live-fetch and search kind pulls your page in response to a real user question and cites you, which sends traffic. A smart allowlist isn't "block AI." It's "permit what sends readers, scrutinize what only takes." We walk through that split crawler by crawler in the [GPTBot decision guide](https://crawlytics.app/c/claude/blog/block-gptbot-decision-guide), and the [AI bots list](https://crawlytics.app/c/claude/resources/ai-bots-list) maps every major user agent to what it actually does. ## A default-deny playbook for sites without a licensing team You don't need Reuters' clout to borrow Reuters' discipline. The order of operations is what matters, and most sites get it backwards by editing robots.txt first and measuring never. Flip that. **Measure before you block.** Pull your server or CDN logs and find out which bots actually hit your site, how often, and which pages they hammer. The People Inc. lesson is that the bots you don't track are the ones costing you the most. You can't make a value-exchange call on a crawler you didn't know was there. **Sort by what each bot gives back.** Group the crawlers you find into three buckets: send-me-traffic (search and live-fetch bots like Googlebot, Bingbot, ChatGPT-User, OAI-SearchBot, Claude-User), take-only (training crawlers and scrapers that never refer a visitor), and unknown. Allow the first bucket without hesitation. Scrutinize the second. Investigate the third before it grows. **Start with a soft allowlist, not a hard one.** You don't have to go full default-deny on day one. Begin by allowing your known-good search and AI-answer bots explicitly, then disallow the specific take-only crawlers you've identified. That captures most of the upside with far less risk of accidentally blocking a bot that was sending you readers. The [manage AI crawlers guide](https://crawlytics.app/c/claude/resources/manage-ai-crawlers) has the ready-to-paste robots.txt, Cloudflare, and nginx configs for each posture. **Enforce where it counts.** Remember the Tollbit 30%. Robots.txt states intent, but the bots that ignore it only stop at the CDN or WAF. If a specific scraper is costing you real bandwidth and ignoring the file, rate-limit or block it at Cloudflare or nginx by user agent, where the request can actually be dropped. **Re-measure on a schedule.** New crawlers launch constantly. The whole reason blocklists fail is that they go stale, and an allowlist goes stale the same way if you never check what's hitting the gate. A monthly look at your bot traffic is enough to catch a new entrant before it becomes a 28,000-agent surprise. ## The bottom line Default-deny is the right instinct because the math of the open web changed. When new bots outpace any blocklist and a third of them ignore the rules anyway, "allow only what you approve" is the only posture that scales. The publishers flipping their robots.txt aren't being paranoid, they're being realistic about a file that was never designed for this. The honest caveat for everyone who isn't Reuters: an allowlist is a tool, not a reflex. Block the wrong bots and you lock yourself out of AI search at the exact moment it's becoming how people find things. Start with measurement, allow the crawlers that bring readers, and refuse the ones that only take. That's the version of default-deny that works whether you have a licensing team or just a robots.txt file and a bandwidth bill. ## Related Written by Crawlytics Team. Crawlytics tracks AI bots, generates llms.txt, and powers WebMCP commerce, all from one snippet on any stack. [See how it works →](https://crawlytics.app/c/claude/) ## Frequently Asked Questions ### What is the difference between an allowlist and a blocklist for AI crawlers? A blocklist names the specific bots you want to keep out and allows everyone else by default. An allowlist (the default-deny model) names the few bots you permit and refuses everyone else by default. The practical difference is coverage: a blocklist only stops crawlers you've heard of, while an allowlist stops every bot you haven't explicitly approved. People Inc. found that switching from a blocklist to an allowlist raised its blocked-agent count from about 2,100 to more than 30,000, because the blocklist had been missing tens of thousands of crawlers it never knew to name. ### Does a default-deny robots.txt actually stop AI bots? Not on its own. Robots.txt is a voluntary convention with no enforcement, and a Tollbit report found that about 30% of AI bot scrapes ignore the permissions in the file entirely. A default-deny robots.txt clearly states your intent and gives well-behaved crawlers a rule to follow, but the bots that ignore it only stop at the CDN or WAF layer, where you can rate-limit or hard-block by user agent. Treat robots.txt as the policy layer and your CDN as the enforcement layer. ### Should a small website use a default-deny robots.txt? Usually not as an aggressive first step. Smaller sites rarely have the licensing clout that makes a hard allowlist pay off, and blocking the wrong bots can remove you from AI answers that drive discovery. A better approach is a soft allowlist: explicitly allow the search and live-fetch bots that send you traffic, then disallow the specific take-only crawlers you've identified in your logs. Measure your real bot traffic first, then tighten from there. ### Which AI crawlers should I allow if I switch to an allowlist? Allow the bots that send readers back to your site. That generally means search and live-fetch crawlers like Googlebot, Bingbot, ChatGPT-User, OAI-SearchBot, and Claude-User, which fetch a page in response to a real user query and cite you. Reuters allowlists Amazon, Google, Bing/Microsoft, Yahoo, and OpenAI based on a "fair value exchange" test of licensing, referral traffic, uptime, and monetization. Scrutinize training-only crawlers that take content without sending visitors, and investigate any user agent you don't recognize before allowing it. ### Will blocking AI crawlers hurt my search visibility? It can, depending on which crawlers you block. Blocking a training-only bot like GPTBot has no effect on traditional search rankings. But refusing live-fetch and AI-search crawlers removes you from the answers those assistants generate. Anthropic's own documentation now warns publishers about the visibility cost of blocking its search bot. If being found is part of your business model, allow the crawlers that cite you in AI answers and reserve blocking for the ones that only take content without referring traffic. --- title: "AI Agent Transactions: Chrome Auto-Browse Hits 200M+ Phones" type: [Organization, Article, BreadcrumbList, WebSite, FAQPage] author: Crawlytics Team publisher: Crawlytics datePublished: 2026-06-10 dateModified: 2026-06-10 canonical: https://crawlytics.app/c/claude/blog/ai-agent-transactions category: blog wordCount: 1848 readingTime: 9 min crawledAt: 2026-06-21 16:40:18 lastVerified: 2026-08-25 13:09:22 site: https://crawlytics.app/c/claude/ --- # AI Agent Transactions: Chrome Auto-Browse Hits 200M+ Phones ## Summary AI agent transactions arrive on 200M+ Android phones via Chrome auto-browse in late June 2026. What makes a site agent-transactable — not just agent-readable — and how to audit yours. ## Key facts - The old visibility question (does an AI assistant mention you in its answer? - App-level agents reach the people who chose to install them. - Most "AI-ready" work to date optimized for one bar: being _readable_. - Auto-browse uses Gemini 3's multimodal model to read a page, identify what is on it, fill forms, navigate the flow, and complete the transaction. - If you have read about [WebMCP](https://crawlytics. For two years, "AI visibility" meant one thing: does an LLM cite your site when someone asks? That question is about to get a more expensive sibling. Starting late June 2026, Google's Chrome auto-browse lands on Android at the operating-system level, default-on for everyone with a Pixel 10 or Galaxy S26, with Google's stated path reaching more than 200 million devices by the end of the year. When an agent shows up on a user's phone to book the appointment, the question is no longer just whether it found you. It is whether it can finish. ## From one question to two The old visibility question (does an AI assistant mention you in its answer?) measured whether you exist in the model's worldview. The new one measures whether you can be operated. An agent that cites you but can't complete your checkout sends the user a recommendation. An agent that can complete your checkout sends you a sale. That gap is the whole story. If your site can be read but not driven, you don't lose a citation, you lose the conversion that the citation used to lead to. The agent reads three roofing companies, picks the one whose booking form it can actually fill, and books it. The other two never find out they were in the running. ## Why OS-level distribution changes the stakes App-level agents reach the people who chose to install them. A ChatGPT app or a Perplexity Comet browser is opt-in, which kept agentic transactions in early-adopter territory through 2025 and early 2026. An agent baked into the operating system reaches everyone who bought the phone. There is nothing to download and nothing to enable. That is what shifts late June 2026 from a product launch into a distribution event. The Pixel 10 and Galaxy S26 are the first wave, and Google has said the same capability extends to watches, cars, glasses, and laptops across the rest of the year. The audience for "can an agent transact on you" jumps from a sliver of power users to a meaningful share of your actual mobile traffic, on a default setting, almost overnight. Two details say this is real capability rather than a demo. Google's underlying agent work, Project Mariner, scored 83.5% on the WebVoyager benchmark for completing real web tasks. And the feature is metered like something people use for high-value work: the AI Pro tier runs $19.99/month for 20 agent tasks a day, and AI Ultra runs $249.99/month for 200. People paying by the task tend to delegate the tasks that matter: the bookings, the orders, the reservations. ## Agent-readable is not agent-transactable Most "AI-ready" work to date optimized for one bar: being _readable_. Clean meta descriptions, server-rendered content, a tidy `llms.txt`, structured headings an assistant can quote. Readable means an agent can fetch your page and understand what it says. It is the bar that earns citations. Transactable is a higher bar. It means a non-human operator can complete a task on your page: submit the form, pick the slot, reach the confirmation screen. A site can be perfectly readable and completely un-transactable, and until this year that was fine, because nothing was trying to operate it. Chrome auto-browse is the thing that starts trying, at scale, by default. ## What actually makes you transactable Auto-browse uses Gemini 3's multimodal model to read a page, identify what is on it, fill forms, navigate the flow, and complete the transaction. Google has not published the exact pathway, but it combines vision with DOM access and accessibility-tree reads. The practical translation: the agent operates your real website the way a user does, faster and without anyone tapping. So readiness is not about adding a protocol. It is about whether a careful, non-human operator can drive the DOM you already have. There is a 30-second test for this, and it is worth running before you read another word of strategy. Open your booking or checkout flow in Chrome on a phone. Disable JavaScript in dev tools. Reload. Can you see the form, see the buttons, and finish the task with the keyboard alone? If yes, the agent can too. If the page goes blank or the flow breaks, you have work to do. The failure modes that stop an agent mid-task are mostly old accessibility sins wearing a new consequence: - **Client-side-only rendering** — if the page is blank without JavaScript, render server-side or hydrate before the form appears. - **Cookie and consent walls** — the agent has to find a real "Accept" button to get past them; a trap with no clear control stops it cold. - **Unlabeled form fields** — inputs need a real `