home.social

#webcrawling — Public Fediverse posts

Live and recent posts from across the Fediverse tagged #webcrawling, aggregated by home.social.

fetched live
  1. PatronView blocks Amazon's AI crawler after 117,000 daily page reads: Anthropic's crawler hit a 35,000 to 1 crawl ratio, and CAPTCHA solve rates measured just 0.24%. The findings show why small operators are locking down servers. ppc.land/patronview-blocks-ama #AI #WebCrawling #CyberSecurity #TechNews #DataPrivacy

  2. PatronView blocks Amazon's AI crawler after 117,000 daily page reads: Anthropic's crawler hit a 35,000 to 1 crawl ratio, and CAPTCHA solve rates measured just 0.24%. The findings show why small operators are locking down servers. ppc.land/patronview-blocks-ama #AI #WebCrawling #CyberSecurity #TechNews #DataPrivacy

  3. RE: mastodon.social/@koen_hufkens/

    This is absolute insanity.
    We could have decentralized index and no reason to crawl websites like madman.
    We have search engines that crawl the web regularly, including Google, and they can behave!

    But ofc people in silicon Valley building shit mentality prefer the hammer style. Brute forcing everything.

    See OpenAI and anthropic bellow. That is just evil, it should be considered a DDOS and punishable by law.

    What is the solution?
    We can't avoid what they are doing right now, the PandoraBox is open, but what can we do?

    I wonder if we could get traffic shapes to tell us if this is coming from harnesses or directly from different LLM vendors such as OpenAI, Anthropic, ZAI, etc.
    This website seems to be facing the later, but traffic from harnesses are still egregious already.

    If the traffic come from vendors we should sue them for DDOS, if it's from harnesses, I think many of them are open source we could contribute so they start using a distributed index or existing search engine results.

    Any other ideas?

    #search #distributedsystems #openai #anthropic
    #searchengine #webcrawler #webcrawling

  4. RE: mastodon.social/@koen_hufkens/

    This is absolute insanity.
    We could have decentralized index and no reason to crawl websites like madman.
    We have search engines that crawl the web regularly, including Google, and they can behave!

    But ofc people in silicon Valley building shit mentality prefer the hammer style. Brute forcing everything.

    See OpenAI and anthropic bellow. That is just evil, it should be considered a DDOS and punishable by law.

    What is the solution?
    We can't avoid what they are doing right now, the PandoraBox is open, but what can we do?

    I wonder if we could get traffic shapes to tell us if this is coming from harnesses or directly from different LLM vendors such as OpenAI, Anthropic, ZAI, etc.
    This website seems to be facing the later, but traffic from harnesses are still egregious already.

    If the traffic come from vendors we should sue them for DDOS, if it's from harnesses, I think many of them are open source we could contribute so they start using a distributed index or existing search engine results.

    Any other ideas?

    #search #distributedsystems #openai #anthropic
    #searchengine #webcrawler #webcrawling

  5. ICYMI: Google's crawler math turns against it as the open web pushes back: Cloudflare shows AI bots crawling sites 50,000 times per visitor as Munich strips Google's cover for Overviews and Google dismisses a rival crawler standard. ppc.land/googles-crawler-math- #Google #WebCrawling #AI #Cloudflare #OpenWeb

  6. ICYMI: Google's crawler math turns against it as the open web pushes back: Cloudflare shows AI bots crawling sites 50,000 times per visitor as Munich strips Google's cover for Overviews and Google dismisses a rival crawler standard. ppc.land/googles-crawler-math- #Google #WebCrawling #AI #Cloudflare #OpenWeb

  7. Google's crawler math turns against it as the open web pushes back: Cloudflare shows AI bots crawling sites 50,000 times per visitor as Munich strips Google's cover for Overviews and Google dismisses a rival crawler standard. ppc.land/googles-crawler-math- #Google #AI #WebCrawling #Cloudflare #SEO

  8. Google's crawler math turns against it as the open web pushes back: Cloudflare shows AI bots crawling sites 50,000 times per visitor as Munich strips Google's cover for Overviews and Google dismisses a rival crawler standard. ppc.land/googles-crawler-math- #Google #AI #WebCrawling #Cloudflare #SEO

  9. Engadget: Cloudflare will filter out web crawlers that serve AI companies . “Cloudflare has announced plans to automatically block mixed-use web crawlers that index websites for search engines and act as AI agents and trainers at the same time. The company previously offered its customers the optional ability to prevent crawlers from scraping their sites for AI chatbots, but now Cloudflare’s […]

    https://rbfirehose.com/2026/07/06/engadget-cloudflare-will-filter-out-web-crawlers-that-serve-ai-companies/
  10. Engadget: Cloudflare will filter out web crawlers that serve AI companies . “Cloudflare has announced plans to automatically block mixed-use web crawlers that index websites for search engines and act as AI agents and trainers at the same time. The company previously offered its customers the optional ability to prevent crawlers from scraping their sites for AI chatbots, but now Cloudflare’s […]

    https://rbfirehose.com/2026/07/06/engadget-cloudflare-will-filter-out-web-crawlers-that-serve-ai-companies/
  11. RT @Hobo_Web: Achtung... Ein Blick auf die frisch aktualisierten Daten der Google Search Console zeigt, dass Google erneut bestimmte Seiten/URLs aus dem Web entfernt hat. Dies könnte die Verzögerung der letzten Wochen erklären und ein Indikator für solche zukünftigen Google-Aktivitäten sein. In diesem Beispiel ist die Bereinigung im Bereich der nicht indizierten Seiten sichtbar. Beim letzten Mal wurden dabei auch indizierte Inhalte zu Waren/Produkten betroffen.

    mehr auf Arint.info

    #GoogleSearchConsole #GoogleUpdate #SEO #Suchmaschinenoptimierung #WebCrawling #Webmaster #arint_info

    https://x.com/Hobo_Web/status/2073171374554665217#m

  12. FYI: Cloudflare stops charging AI per crawl and starts paying per answer: Cloudflare ties AI payments to citations as half of bot crawls fetch unchanged pages, and a randomized study finds AI Overviews cut publisher clicks 39.8%. ppc.land/cloudflare-stops-char #Cloudflare #AI #ArtificialIntelligence #WebCrawling #DigitalMarketing

  13. FYI: Cloudflare stops charging AI per crawl and starts paying per answer: Cloudflare ties AI payments to citations as half of bot crawls fetch unchanged pages, and a randomized study finds AI Overviews cut publisher clicks 39.8%. ppc.land/cloudflare-stops-char #Cloudflare #AI #ArtificialIntelligence #WebCrawling #DigitalMarketing

  14. ICYMI: Cloudflare stops charging AI per crawl and starts paying per answer: Cloudflare ties AI payments to citations as half of bot crawls fetch unchanged pages, and a randomized study finds AI Overviews cut publisher clicks 39.8%. ppc.land/cloudflare-stops-char #Cloudflare #AI #DigitalMarketing #WebCrawling #SEO

  15. ICYMI: Cloudflare stops charging AI per crawl and starts paying per answer: Cloudflare ties AI payments to citations as half of bot crawls fetch unchanged pages, and a randomized study finds AI Overviews cut publisher clicks 39.8%. ppc.land/cloudflare-stops-char #Cloudflare #AI #DigitalMarketing #WebCrawling #SEO

  16. ICYMI: Cloudflare ties AI payouts to citations as 50% of crawls waste: Ceramic.ai and You.com will now pay publishers per query result, not per page fetch, since over half of good-bot crawls only re-fetch pages that never changed. ppc.land/cloudflare-ties-ai-pa #Cloudflare #AI #DigitalMarketing #SEO #WebCrawling

  17. ICYMI: Cloudflare ties AI payouts to citations as 50% of crawls waste: Ceramic.ai and You.com will now pay publishers per query result, not per page fetch, since over half of good-bot crawls only re-fetch pages that never changed. ppc.land/cloudflare-ties-ai-pa #Cloudflare #AI #DigitalMarketing #SEO #WebCrawling

  18. Cloudflare ties AI payouts to citations as 50% of crawls waste: Ceramic.ai and You.com will now pay publishers per query result, not per page fetch, since over half of good-bot crawls only re-fetch pages that never changed. ppc.land/cloudflare-ties-ai-pa #Cloudflare #AI #DigitalMarketing #SEO #WebCrawling

  19. Cloudflare ties AI payouts to citations as 50% of crawls waste: Ceramic.ai and You.com will now pay publishers per query result, not per page fetch, since over half of good-bot crawls only re-fetch pages that never changed. ppc.land/cloudflare-ties-ai-pa #Cloudflare #AI #DigitalMarketing #SEO #WebCrawling

  20. FYI: The user agent strings every SEO and site owner needs right now: A technical reference for GPTBot, OAI-SearchBot, ClaudeBot, PerplexityBot, and Bingbot - the exact user agent strings shaping how AI crawls the web in 2026. ppc.land/the-user-agent-string #SEO #UserAgent #AI #WebCrawling #DigitalMarketing

  21. FYI: The user agent strings every SEO and site owner needs right now: A technical reference for GPTBot, OAI-SearchBot, ClaudeBot, PerplexityBot, and Bingbot - the exact user agent strings shaping how AI crawls the web in 2026. ppc.land/the-user-agent-string #SEO #UserAgent #AI #WebCrawling #DigitalMarketing

  22. ICYMI: The user agent strings every SEO and site owner needs right now: A technical reference for GPTBot, OAI-SearchBot, ClaudeBot, PerplexityBot, and Bingbot - the exact user agent strings shaping how AI crawls the web in 2026. ppc.land/the-user-agent-string #SEO #WebCrawling #UserAgent #DigitalMarketing #AICrawlers

  23. ICYMI: The user agent strings every SEO and site owner needs right now: A technical reference for GPTBot, OAI-SearchBot, ClaudeBot, PerplexityBot, and Bingbot - the exact user agent strings shaping how AI crawls the web in 2026. ppc.land/the-user-agent-string #SEO #WebCrawling #UserAgent #DigitalMarketing #AICrawlers

  24. "The worst of these proposed standards would give websites far greater ability to automatically block legitimate, lawful scraping and crawling. For example, the AI Preferences working group is working on proposals to give publishers a way to express “preference signals” against crawling web data for AI-related purposes, including to train models, generate outputs, and help users search the web. These preference signals would be expressed through robots.txt and could potentially become legally binding in some jurisdictions.

    Another working group, called Web Bot Auth, is pursuing efforts to protect sites from overly-aggressive bots that strain website resources—a positive goal that could meaningfully improve the internet in the AI era. But Web Bot Auth is simultaneously pursuing a much more dangerous path as well: standards changes that would enable sites to cryptographically identify bots so that they can more easily block anyone they wish—not just “bad” actors, but competitors, dissidents, or anyone who hasn’t paid for the right to access sites using automated tools. If sites restrict crawling to a preapproved list of cryptographically authenticated bots, they could require licensing payments from those wishing to crawl their sites. This would close off the open web to researchers, archivists, and startups without the ability to pay for automated access.

    Websites may have legitimate reasons to worry about AI’s impacts on their traffic and advertising revenue, but those reasons must be weighed against the benefits of the open web."

    eff.org/deeplinks/2026/06/free

    #IETF #OpenWeb #WebCrawling #AI #Chatbots #LLMs

  25. "The worst of these proposed standards would give websites far greater ability to automatically block legitimate, lawful scraping and crawling. For example, the AI Preferences working group is working on proposals to give publishers a way to express “preference signals” against crawling web data for AI-related purposes, including to train models, generate outputs, and help users search the web. These preference signals would be expressed through robots.txt and could potentially become legally binding in some jurisdictions.

    Another working group, called Web Bot Auth, is pursuing efforts to protect sites from overly-aggressive bots that strain website resources—a positive goal that could meaningfully improve the internet in the AI era. But Web Bot Auth is simultaneously pursuing a much more dangerous path as well: standards changes that would enable sites to cryptographically identify bots so that they can more easily block anyone they wish—not just “bad” actors, but competitors, dissidents, or anyone who hasn’t paid for the right to access sites using automated tools. If sites restrict crawling to a preapproved list of cryptographically authenticated bots, they could require licensing payments from those wishing to crawl their sites. This would close off the open web to researchers, archivists, and startups without the ability to pay for automated access.

    Websites may have legitimate reasons to worry about AI’s impacts on their traffic and advertising revenue, but those reasons must be weighed against the benefits of the open web."

    eff.org/deeplinks/2026/06/free

    #IETF #OpenWeb #WebCrawling #AI #Chatbots #LLMs

  26. FYI: OpenAI tripled its web crawl after GPT-5 - but ChatGPT users may be declining: New log file analysis of 7 billion OpenAI bot events reveals a 3.5x surge in OAI-SearchBot activity after GPT-5, while ChatGPT user-driven events dropped 28%. ppc.land/openai-tripled-its-we #OpenAI #GPT5 #ChatGPT #AItrends #webcrawling

  27. FYI: Google-Agent joins the crawler list as AI browsing gets an official identity: Google on March 20 added Google-Agent to its user-triggered fetchers list, formalizing a new user agent for AI systems like Project Mariner that navigate the web on behalf of users. ppc.land/google-agent-joins-th #GoogleAgent #AIBrowsing #UserAgent #WebCrawling #ProjectMariner

  28. FYI: Googlebot is not a program - Google engineers finally explain what it really is: Google engineers reveal Googlebot is a misnomer for a central SaaS crawling platform serving dozens of products, with a 15 MB default file size limit and geo-crawling constraints. ppc.land/googlebot-is-not-a-pr #Googlebot #SEO #WebCrawling #DigitalMarketing #SaaS

  29. FYI: Googlebot is not a program - Google engineers finally explain what it really is: Google engineers reveal Googlebot is a misnomer for a central SaaS crawling platform serving dozens of products, with a 15 MB default file size limit and geo-crawling constraints. ppc.land/googlebot-is-not-a-pr #Googlebot #SEO #WebCrawling #DigitalMarketing #SaaS

  30. Googlebot is not a program - Google engineers finally explain what it really is: Google engineers reveal Googlebot is a misnomer for a central SaaS crawling platform serving dozens of products, with a 15 MB default file size limit and geo-crawling constraints. ppc.land/googlebot-is-not-a-pr #Googlebot #SEO #WebCrawling #SaaS #DigitalMarketing

  31. Googlebot is not a program - Google engineers finally explain what it really is: Google engineers reveal Googlebot is a misnomer for a central SaaS crawling platform serving dozens of products, with a 15 MB default file size limit and geo-crawling constraints. ppc.land/googlebot-is-not-a-pr #Googlebot #SEO #WebCrawling #SaaS #DigitalMarketing

  32. FYI: Google's secret crawl logic, finally explained in one page: Google published a new web crawling overview on March 3, 2026, detailing how Googlebot discovers, renders, and manages site access across 30+ years of web indexing. ppc.land/googles-secret-crawl- #Google #SEO #WebCrawling #Googlebot #DigitalMarketing

  33. FYI: Google's secret crawl logic, finally explained in one page: Google published a new web crawling overview on March 3, 2026, detailing how Googlebot discovers, renders, and manages site access across 30+ years of web indexing. ppc.land/googles-secret-crawl- #Google #SEO #WebCrawling #Googlebot #DigitalMarketing

  34. ICYMI: Google's secret crawl logic, finally explained in one page: Google published a new web crawling overview on March 3, 2026, detailing how Googlebot discovers, renders, and manages site access across 30+ years of web indexing. ppc.land/googles-secret-crawl- #Google #SEO #WebCrawling #Googlebot #DigitalMarketing

  35. ICYMI: Google's secret crawl logic, finally explained in one page: Google published a new web crawling overview on March 3, 2026, detailing how Googlebot discovers, renders, and manages site access across 30+ years of web indexing. ppc.land/googles-secret-crawl- #Google #SEO #WebCrawling #Googlebot #DigitalMarketing

  36. Google's secret crawl logic, finally explained in one page: Google published a new web crawling overview on March 3, 2026, detailing how Googlebot discovers, renders, and manages site access across 30+ years of web indexing. ppc.land/googles-secret-crawl- #Google #SEO #WebCrawling #Googlebot #DigitalMarketing

  37. Google's secret crawl logic, finally explained in one page: Google published a new web crawling overview on March 3, 2026, detailing how Googlebot discovers, renders, and manages site access across 30+ years of web indexing. ppc.land/googles-secret-crawl- #Google #SEO #WebCrawling #Googlebot #DigitalMarketing

  38. Smart TVs are now running Bright SDK to silently crawl the web for AI training, using residential proxies to bypass Google policy. The move sparks a compliance backlash and raises privacy concerns for home devices. How will regulators respond, and what does this mean for open‑source AI? Dive into the details. #SmartTV #BrightSDK #WebCrawling #DeviceCompliance

    🔗 aidailypost.com/news/smart-tvs

  39. Smart TVs are now running Bright SDK to silently crawl the web for AI training, using residential proxies to bypass Google policy. The move sparks a compliance backlash and raises privacy concerns for home devices. How will regulators respond, and what does this mean for open‑source AI? Dive into the details. #SmartTV #BrightSDK #WebCrawling #DeviceCompliance

    🔗 aidailypost.com/news/smart-tvs

  40. Ah, #wxpath, because using #XPath was just too easy before 🙄. Now with extra layers of #complexity, just in case you weren't already confused enough by web crawling! 🕸️😵‍💫
    github.com/rodricios/wxpath #webcrawling #technews #developerhumor #HackerNews #ngated

  41. Ah, #wxpath, because using #XPath was just too easy before 🙄. Now with extra layers of #complexity, just in case you weren't already confused enough by web crawling! 🕸️😵‍💫
    github.com/rodricios/wxpath #webcrawling #technews #developerhumor #HackerNews #ngated

  42. Cloudflare's 2025 data reveals Google's structural advantage in AI training: Googlebot crawled 11.6% of web pages vs OpenAI's 3.6%. Publishers face an impossible choice - they can't block Google's AI crawling without losing search visibility entirely, since the same bot handles both functions. #AI #WebCrawling #DigitalRights

    implicator.ai/googles-quiet-co

  43. Cloudflare's 2025 data reveals Google's structural advantage in AI training: Googlebot crawled 11.6% of web pages vs OpenAI's 3.6%. Publishers face an impossible choice - they can't block Google's AI crawling without losing search visibility entirely, since the same bot handles both functions. #AI #WebCrawling #DigitalRights

    implicator.ai/googles-quiet-co

  44. Search Engine Roundtable: OpenAI Scales Up Crawling & Bots For The Holidays. “OpenAI is reportedly scaling up its crawling infrastructure for the holiday shopping season. The folks at Merj noticed OpenAI adding a lot of new IP ranges for its bots and crawlers.”

    https://rbfirehose.com/2025/12/01/search-engine-roundtable-openai-scales-up-crawling-bots-for-the-holidays/

  45. Search Engine Roundtable: OpenAI Scales Up Crawling & Bots For The Holidays. “OpenAI is reportedly scaling up its crawling infrastructure for the holiday shopping season. The folks at Merj noticed OpenAI adding a lot of new IP ranges for its bots and crawlers.”

    https://rbfirehose.com/2025/12/01/search-engine-roundtable-openai-scales-up-crawling-bots-for-the-holidays/

  46. Released scrapy-contrib-bigexporter 1.0.0 (codeberg.org/ZuInnoTe/scrapy-c) - additional export formats for the webscraping framework Scrapy.

    Migrated parquet export from fastparquet to pyarrow as fastparquet is deprecated (docs.dask.org/en/stable/change)

    Migrated orc export from pyorc to pyarrow to reduce the number of dependencies

    #scrapy #crawling #python #parquet #orc #pyarrow #webcrawling #scraping

  47. Released scrapy-contrib-bigexporter 1.0.0 (codeberg.org/ZuInnoTe/scrapy-c) - additional export formats for the webscraping framework Scrapy.

    Migrated parquet export from fastparquet to pyarrow as fastparquet is deprecated (docs.dask.org/en/stable/change)

    Migrated orc export from pyorc to pyarrow to reduce the number of dependencies

    #scrapy #crawling #python #parquet #orc #pyarrow #webcrawling #scraping

  48. Remember when 'robots.txt' was supposed to solve all our crawling problems? Online media brands are trying a new protocol to deter 'unwanted' AI crawlers. Because clearly, we need more digital fences. What's your bet on how long it takes for a savvy AI to find a workaround?

    Read more: cnet.com/tech/services-and-sof

    #AI #TechNews #WebCrawling #DigitalRights #Privacy

  49. Search Engine Land: Google fixes reduced crawling issue impacting some websites. “Google has confirmed it fixed an issue with its crawlers impacting ‘some sites.’ The issue was ‘reduced / fluctuating crawling’ from Google’s end with Googlebot. It is now resolved and Google said the crawling should pick back up in the near future.”

    https://rbfirehose.com/2025/08/31/search-engine-land-google-fixes-reduced-crawling-issue-impacting-some-websites/