home.social

#crawlers — Public Fediverse posts

Live and recent posts from across the Fediverse tagged #crawlers, aggregated by home.social.

fetched live
  1. RE: social.edu.nl/@wlaatje/1169708

    "The web is full of independent archives, hobby databases, local news sites, forums, reference works. Decades of accumulated human effort, running on old code, maintained by small teams or single individuals, quietly holding up far more of our shared knowledge than anyone acknowledges."

    And proponents of so-called 'AI' are speedrunning their destruction.

    #noAI #scrapers #scraping #crawlers #AI #genAI

  2. RE: social.edu.nl/@wlaatje/1169708

    "The web is full of independent archives, hobby databases, local news sites, forums, reference works. Decades of accumulated human effort, running on old code, maintained by small teams or single individuals, quietly holding up far more of our shared knowledge than anyone acknowledges."

    And proponents of so-called 'AI' are speedrunning their destruction.

    #noAI #scrapers #scraping #crawlers #AI #genAI

  3. Thanks to content scraper from companies and protection solutions from companies like for delivering 56 kbps era response times in the age of fibre internet connections! How nice of them to constantly satisfy our nostalgia!

  4. Thanks to content scraper #crawlers from #AI companies and #bot protection solutions from companies like #Cloudflare for delivering 56 kbps #dialup era #website response times in the age of #gigabit fibre internet connections! How nice of them to constantly satisfy our nostalgia!

  5. ICYMI: AI crawlers hit sites 50,000 times per human visit, Cloudflare data shows: AI crawlers reach up to 50,000 visits per referred reader, Cloudflare data shows, as a Munich court strips Google of its liability shield over AI Overviews. ppc.land/ai-crawlers-hit-sites #AI #Crawlers #DataAnalysis #Cloudflare #Google

  6. ICYMI: AI crawlers hit sites 50,000 times per human visit, Cloudflare data shows: AI crawlers reach up to 50,000 visits per referred reader, Cloudflare data shows, as a Munich court strips Google of its liability shield over AI Overviews. ppc.land/ai-crawlers-hit-sites #AI #Crawlers #DataAnalysis #Cloudflare #Google

  7. #Cloudflare gives #AI #crawlers a September deadline: pay publishers or get blocked - thenextweb.com/news/cloudflare "From 15 September, Cloudflare will block crawlers that harvest content for AI training from any page carrying ads, unless the owner opts in, and pay publishers when their work shapes an AI answer. "

  8. #Cloudflare gives #AI #crawlers a September deadline: pay publishers or get blocked - thenextweb.com/news/cloudflare "From 15 September, Cloudflare will block crawlers that harvest content for AI training from any page carrying ads, unless the owner opts in, and pay publishers when their work shapes an AI answer. "

  9. ICYMI: Microsoft Clarity now flags robots.txt violations inside Bot Analytics: Microsoft Clarity now surfaces robots.txt violations in Bot Analytics, showing publishers which AI crawlers break access rules and what content they target. ppc.land/microsoft-clarity-now #MicrosoftClarity #BotAnalytics #SEO #WebAnalytics #Crawlers

  10. ICYMI: New York passes bill forcing AI crawlers to identify themselves to news sites: New York's Assembly passed A11292 on June 5, 2026, requiring AI crawlers to disclose identity and purpose to news publishers or face $15,000-per-day penalties. ppc.land/new-york-passes-bill- #NewYork #AI #Crawlers #News #Legislation

  11. ICYMI: New York passes bill forcing AI crawlers to identify themselves to news sites: New York's Assembly passed A11292 on June 5, 2026, requiring AI crawlers to disclose identity and purpose to news publishers or face $15,000-per-day penalties. ppc.land/new-york-passes-bill- #NewYork #AI #Crawlers #News #Legislation

  12. #Anti #AI #Software / #Website Protection / AI #poison

    iocaine.madhouse-project.org/

    iocaine features:

    > No sympathy towards #crawlers

    Stand in the way of AI but not your visitors

    > Lightweight = Very little #CPU and #RAM

    Tries its best to require little resources from your visitors too.

    If you can serve #static content, you'll be able to run #iocaine too.

  13. #Anti #AI #Software / #Website Protection / AI #poison

    iocaine.madhouse-project.org/

    iocaine features:

    > No sympathy towards #crawlers

    Stand in the way of AI but not your visitors

    > Lightweight = Very little #CPU and #RAM

    Tries its best to require little resources from your visitors too.

    If you can serve #static content, you'll be able to run #iocaine too.

  14. "Our new analysis shows that more than 340 local news sites across the United States are now limiting the #InternetArchive’s ability to access and preserve their stories.

    Researchers, historians, and citizens around the world rely on the web archives of #localnews sites to do their work.

    “Blocking the Internet Archive’s web #crawlers threatens one of the most effective ways that we capture and store news content for the long term.”"

    niemanlab.org/2026/05/more-tha

  15. "Our new analysis shows that more than 340 local news sites across the United States are now limiting the #InternetArchive’s ability to access and preserve their stories.

    Researchers, historians, and citizens around the world rely on the web archives of #localnews sites to do their work.

    “Blocking the Internet Archive’s web #crawlers threatens one of the most effective ways that we capture and store news content for the long term.”"

    niemanlab.org/2026/05/more-tha

  16. Fucking #crawlers of #meta #facebook #apple #google are easting bandwidth and creating nonsense. In the last 3 days meta developer crawlers alone ate up 850GB+ #bandwidth. Assholes.

  17. Fucking #crawlers of #meta #facebook #apple #google are easting bandwidth and creating nonsense. In the last 3 days meta developer crawlers alone ate up 850GB+ #bandwidth. Assholes.

  18. Blocking AI crawlers cost news publishers 7% of traffic, study finds: A Wharton and Rutgers study finds news publishers who blocked LLM crawlers lost 7% of weekly traffic in 6 weeks, with no measurable content protection gains. ppc.land/blocking-ai-crawlers- #AI #crawlers #NewsPublishers #TrafficLoss #ContentProtection

  19. Blocking AI crawlers cost news publishers 7% of traffic, study finds: A Wharton and Rutgers study finds news publishers who blocked LLM crawlers lost 7% of weekly traffic in 6 weeks, with no measurable content protection gains. ppc.land/blocking-ai-crawlers- #AI #crawlers #NewsPublishers #TrafficLoss #ContentProtection

  20. ICYMI: Google rewrites Googlebot's rulebook: 2MB limits, IP moves, and what crawlers really are: Google today published two blog posts revealing Googlebot's true architecture as a shared SaaS platform, a 2MB fetch limit, and a new IP ranges directory path. ppc.land/google-rewrites-googl #Google #Googlebot #SEO #Crawlers #WebMaster

  21. Fresh on my #blog: "There's a difference between 'scraping' and 'retrieving'". I have a dilemma; one is extractive, the other accessibility. What do I do?
    #ArtificialIntelligence #LLMs #crawlers #ethics #AIethics
    thomasrigby.com/posts/theres-a-difference-between-scraping-and-retrieving/

  22. 52467 requests in about 8h from #OpenAI #crawlers with no sign of stopping.