Weekly Dispatch · archived
Weekly Dispatch · Week 24 of 2026
Crawling & Publisher Controls
This week's discourse highlights the dual challenge for publishers: navigating the emerging AI content licensing market, which offers new revenue streams but raises concerns about power imbalances and data control, while simultaneously implementing granular controls over AI crawlers to manage traffic, prevent unauthorized scraping, and preserve content integrity amidst declining referral traffic.
- Publishers quietly cut 'six-figure' deals via Snowflake's AI licensing platform
Publishers are securing AI licensing deals through Snowflake, enabling monetized RAG access to their content while preventing scraping.
""Publishers are quietly cutting six-figure AI licensing deals on Snowflake, as the data giant positions itself as matchmaker-in-chief between locked-down news content and enterprises keen to plug reliable publisher content into their own internal AI tools via retrieval-augmented generation (RAG).""
- The emerging AI content licensing market puts news publishers in a “double bind,” a new report warns
A report warns that the AI content licensing market creates a "double bind" for publishers, with Big Tech controlling both content creation and monetization.
""The same big tech companies that are developing commercial AI products and stripping news publishers of site traffic are the ones dictating what alternative revenue will look like.""
- AI training data is becoming a seller's market. Here's what it's worth
The market for AI training data is evolving into a seller's market, with publishers securing significant licensing deals and shifting towards usage-based pricing models.
""The era of free AI training data is over. What is replacing it has a price list.""
- The internet's bot majority has arrived ahead of schedule
Bot traffic now dominates the internet, with AI crawlers altering the traditional web economics by extracting content without returning comparable referral traffic to publishers.
""AI breaks that loop because it can turn source material into an answer surface outside the source.""
- SEO was never born at Google
The historical tension over web crawlers and robots.txt continues with AI, as Google still frames robots.txt as a crawl-access tool, not a privacy or deindexing mechanism.
""That same tension remains visible in debates over AI crawlers, scraping, and publisher controls. Google's current documentation still frames robots.txt as a crawl-access tool rather than a true privacy or deindexing mechanism.""
- Microsoft MAI-Thinking-1: Clean Licensed Data Claims Clash With Common Crawl
Microsoft's claims of using "clean licensed data" for AI training are questioned due to its reliance on Common Crawl, which blurs the line between public and commercially licensed content.
""Common Crawl is a crawl, not a clearinghouse. It can preserve and index public pages, but the fact that a page was reachable by a crawler is not the same thing as the page being commercially licensed for model training.""
- Fake Reddit posts are becoming the new SEO for AI search
Reddit's data is now a valuable, licensed source for AI, but this creates an asymmetry where high-quality sources block access, potentially degrading AI answer quality.
""If high-quality sources restrict access while low-quality or manipulative sources remain open, AI systems may face a source diet problem.""
- AI agents are now buying things - and fraud looks identical
The rise of AI agents and scrapers makes it difficult to distinguish legitimate AI crawlers from fraudulent activity, as attackers spoof user-agent strings to bypass controls.
""Attackers spoof user-agent strings to exploit the trust organizations extend to recognized AI crawlers, bypassing robots.txt allowlists and rate-limit exemptions.""
- DuckDuckGo Makes AI-Free Search Easier To Set as Default
AI controls are shifting into search tools, with Google's AI rewriting headlines raising concerns about publisher control and the accuracy of information presented to users.
""Google Search is rewriting headlines with AI, raising concerns about accuracy, trust, and publisher control.""
- AI Crawler Access Control: The 2026 Decision Matrix
Publishers need a nuanced AI crawler access strategy, distinguishing between training and search bots, and potentially blocking at the WAF/IP level due to compliance inconsistencies and referral asymmetry.
""The defensible default is to disallow training crawlers (GPTBot, ClaudeBot, CCBot, Google-Extended, Applebot-Extended) while allowing search and retrieval crawlers (OAI-SearchBot, Claude-SearchBot, PerplexityBot) that send referral traffic.""
Agents
This week's reporting highlights the ongoing transformation of enterprise operations by agentic AI, emphasizing the critical need for robust data access protocols like MCP and careful governance to manage agent security and ensure effective deployment. Field reports also showcase practical implementations of agents for specific business functions.
- How Agentic AI is Transforming Enterprise Integration in 2026
Agentic AI is revolutionizing enterprise integration by enabling autonomous goal pursuit, multi-step actions, and tool use across systems, leading to structured intelligence and task completion.
"Agentic AI describes AI systems that pursue goals autonomously: planning multi-step actions, calling tools such as APIs, databases, and search systems, and adapting based on results without human-defined sequences."
- AI agents can't help if they can't see your marketing data
AI agents require live, current, and unmediated access to enterprise data via protocols like MCP to be effective, as raw API access lacks necessary guardrails and context.
"The problem is getting that data to them live, current, and without a human in the middle copying it across. It's the reason most PPC accounts in 2026 still run almost exactly the way they did before anyone started talking about agents."
- The Best Code I Ever Wrote Was the Code I Stopped Rewriting. Here Are the 3 Repos I Built Instead.
The author discusses building custom AI agents and an MCP Server to connect external AI clients to Salesforce, highlighting the shift towards standardized agent-to-tool communication in enterprise environments.
"With Salesforce now supporting MCP natively in Agentforce, the protocol is becoming the standard for agent-to-tool communication. SAAF's MCP Server works the other direction: it exposes your Salesforce org as a set of tools for any MCP-compatible AI client."
- DLP alerts in Security Copilot: are your controls keeping up?
DLP investigation agents require careful governance and context-driven approaches to ensure controls keep up with their capabilities, treating them as decision support rather than autonomous authorities.
"Security teams should treat DLP agents as decision support, not as autonomous authorities."
Copyright & Legal
This week saw significant developments in AI and copyright, including a major lawsuit against Meta by publishers and authors for alleged copyright infringement in training its AI models. Concurrently, the UK's competition watchdog mandated that Google allow publishers to opt out of AI content scraping for search summaries, while legal scholars analyzed AI's impact on tort litigation and the broader legal industry.
- Mark Zuckerberg 'personally authorized' Meta's copyright infringement, publishers allege
Five publishing houses and author Scott Turow sued Meta and CEO Mark Zuckerberg for allegedly using copyrighted works to train its AI system Llama.
"“Defendants reproduced and distributed millions of copyrighted works without permission, without providing any compensation to authors or publishers, and with full knowledge that their conduct violated copyright law,” the complaint reads in part."
- Hurwitz and Lu, 'An Initial Assessment of Standards in Technology Tort Litigation'
This article assesses how tort law treats standards in technology litigation, finding courts resist treating them as duty-defining baselines, with implications for AI.
"Because AI standards are emerging unusually early in the technological lifecycle – before stable engineering norms or accumulated accident experience exist – courts are unlikely to treat them as duty-defining baselines or safe harbors."
Web Ecosystem & AI Impact
This week's reporting highlights Google's efforts to address publisher concerns over AI Overviews' impact on traffic, while publishers simultaneously pursue content licensing deals as a new revenue stream and engage in lawsuits against AI companies for content usage. The broader consensus is that AI answer engines are significantly reducing referral traffic, necessitating new strategies for journalism and media outlets.
- Google looks to ease publisher concerns over the impact of AI Overviews on referral traffic
Google is introducing new tools and options to address publisher worries about traffic losses due to AI Overviews, despite past reports of significant declines.
"Google's looking to ease publisher concerns about traffic losses due to the rise of AI overviews in Search."
- AI Licensing Deals: A New Revenue Stream for Publishers
Publishers are securing incremental revenue through AI licensing deals via platforms like Snowflake, though this won't fully offset AI-driven traffic declines.
"AI licensing is an incremental revenue line, not a lifeline. It won't offset the referral traffic losses that have compounded across the industry for several years."
- How AI Is Changing Journalism and Media in 2026 — What Nobody Is Telling You
AI answer engines are significantly reducing publisher referral traffic, forcing newsrooms to adapt to a new reality where direct answers bypass their sites.
"AI answer engines — ChatGPT, Perplexity, Google's AI Overviews — are answering questions directly. No click required. No visit to the publisher's website. No ad revenue for the outlet that paid a journalist to do the original reporting."