#webspidering — Public Fediverse posts
Live and recent posts from across the Fediverse tagged #webspidering, aggregated by home.social.
-
Search Engine Journal: OpenAI Says Robots.txt May Not Apply To ChatGPT’s Fetch Bot. “ChatGPT’s page-fetching bot is disallowed by more sites than any other AI bot of its kind. It also reached disallowed pages on more sites than any other bot. OpenAI says robots.txt rules may not apply to it because a person asked for the page.”
https://rbfirehose.com/2026/08/15/search-engine-journal-openai-says-robots-txt-may-not-apply-to-chatgpts-fetch-bot/ -
Ars Technica: Inside the web infrastructure revolt over Google’s AI Overviews. “It could be a consequential act of quiet regulation. Cloudflare, a web infrastructure company, has updated millions of websites’ robots.txt files in an effort to force Google to change how it crawls them to fuel its AI products and initiatives. We spoke with Cloudflare CEO Matthew Prince about what exactly is […]