Reported Backfilled · Original publication August 21, 2026Crawlers & Content ControlsSearch & AI DiscoveryAgentic WebStandards & Protocols Cloudflare Blog ·

Cloudflare Launches Bot Preference Sync to Automatically Align robots.txt with Edge AI Policies

Open source ↗

Summary

Cloudflare has introduced Bot Preference Sync, a feature across all plan tiers that automatically prepends and maintains robots.txt directives corresponding to a site owner’s dashboard configurations for AI Search, Agent, and Training traffic. The system pairs edge-enforced blocks with matching robots.txt rules drawn from its tracked bot directory, and introduces an onboarding default for ad-supported publishers that disallows training while permitting search indexing. Additionally, Cloudflare established stricter verification criteria for mixed-use crawlers, requiring them to respect no-training signals, allow opt-outs from AI summaries, provide page-level training visibility, and prove search indexing parity.

Insight

Edge providers are bridging the gap between stated policy (robots.txt) and enforced policy (WAF/edge blocking) to prevent mixed-use crawlers from exploiting discrepancies between declared preferences and physical access rules.

Implication

Mixed-use crawler operators will face edge-level blocks unless they build granular opt-outs for training/summaries and demonstrate search neutrality, while site owners gain automated alignment across bot categories without manually editing static files.

Why it matters

Discrepancies between robots.txt files and edge firewall policies have historically provided AI crawlers technical or legal ambiguity to ignore site preferences; synchronizing them at the CDN layer standardizes machine-readable compliance enforcement at web scale.

Evidence