home.social

#aicrawlers — Public Fediverse posts

Live and recent posts from across the Fediverse tagged #aicrawlers, aggregated by home.social.

fetched live
  1. So, after #Level0, now even #PTNA got affected by AI crawlers... ( community.openstreetmap.org/t/ ) and there are still people that believe in the generally good usage of AI/LLM.
    A new tool is good or bad depending on the majority of the uses, and as long as the majority of the known usages is abuse and stealing of other people's work, then we can't talk about healthy AI/LLM uses.

    #OpenStreetMap #aicrawlers

  2. 🤖 Infrastructura Kernel.org este bombardată de milioane de cereri zilnice ale bot-ilor de Inteligență Artificială!

    Publicația Linuxiac relatează o problemă majoră de infrastructură cu care se confruntă administratorii Kernel.org (platforma oficială de găzduire și dezvoltare a nucleului Linux, administrată de Konstantin Ryabitsev): un val masiv de AI crawlers și scrapperi automatizați consumă resurse critice ale serverelor.

    ✨ Punctele cheie ale situației generate de bot-ii de IA:

    💥 Consum masiv de resurse procesor (CPU):
    • Serverele git.kernel.org primesc în jur de 6 milioane de cereri zilnice doar pentru paginile individuale ale commit-urilor Git.
    • Din cele 90 de nuclee CPU distribuite pe 5 locații geografice, aproximativ 14-16 nuclee sunt folosite constant doar pentru a genera dinamic pagini HTML pentru bot-ii care adună date pentru antrenarea modelelor de limbaj (LLM).

    📉 Traficul legitim a devenit o minoritate:
    • Tranzacțiile legitime (inclusiv descărcările și comanda git clone efectuate de dezvoltatori umani) reprezintă doar 2% din totalul cererilor, restul de 98% fiind generat de bot-i automatizați.

    ⚙️ Ineficiența scraperilor și fentarea restricțiilor:
    • În loc să cloneze depozitul Git o singură dată și să lucreze local, mulți crawleri interoghează individual fiecare URL de commit.
    • Deși s-au încercat blocări pe bază de IP/ASN și un sistem de verificare Proof-of-Work (Anubis), bot-ii au adaptat rapid atacul trecând prin rețele proxy de IP-uri rezidențiale și de mobil (fentând blocările) și alocând mai multă putere de calcul pentru a trece de verificări.

    🛡️ Măsurile drastice luate în considerare:
    • Echipa de infrastructură ia în calcul reducerea drastică a numărului de URL-uri indexabile/generabile dinamic și restricționarea acțiunilor costisitoare din punct de vedere CPU pentru utilizatorii anonimi, pentru a asigura stabilitatea platformei.

    📌 Concluzie:

    Această situație evidențiază un efect secundar major al cursei pentru antrenarea modelelor de IA: exploatarea excesivă a infrastructurii publice gratuite Open Source, punând o presiune financiară și tehnică uriașă pe proiectele comunitare! 🚀

    #KernelOrg #Linux #AICrawlers #Git #SysAdmin #OpenSource #Linuxiac #TechNews #DevOps

  3. Git repositories are made to be cloned, not scraped commit page by commit page.

    At scale, HTML crawling wastes CPU on rendered pages, creates higher costs for open source hosts, and can lead to proof-of-work checks or rate limits that affect real users first.

    Clone once. Fetch updates. Use Git as intended.

    x.com/GitGem/status/2094005395

    #Git #OpenSource #Linux #FOSS #AICrawlers #DevOps #GitGem

  4. FYI: Cloudflare blocks opaque AI crawlers from sites that disallow training: Bot Preference Sync rewrites robots.txt from dashboard settings on every plan, free to Enterprise. Ad-funded domains now default to disallowing AI training. ppc.land/cloudflare-blocks-opa #Cloudflare #AICrawlers #RobotsTxt #WebDevelopment #AIEthics

  5. ICYMI: Cloudflare blocks opaque AI crawlers from sites that disallow training: Bot Preference Sync rewrites robots.txt from dashboard settings on every plan, free to Enterprise. Ad-funded domains now default to disallowing AI training. ppc.land/cloudflare-blocks-opa #Cloudflare #AICrawlers #RobotsTXT #WebSecurity #SEO

  6. Cloudflare blocks opaque AI crawlers from sites that disallow training: Bot Preference Sync rewrites robots.txt from dashboard settings on every plan, free to Enterprise. Ad-funded domains now default to disallowing AI training. ppc.land/cloudflare-blocks-opa #Cloudflare #AICrawlers #WebScraping #DataPrivacy #SEO

  7. 🔍 Wow, who knew small businesses wanted to be 'visible' only to later be snubbed by AI's selective memory? 😲 Apparently, only a minuscule percentage even care to block the AI crawlers, while the rest are blissfully unaware they're not even invited to the AI party! 🎉
    website-auditor.io/ai-visibili #smallbusiness #AIvisibility #selectiveAI #awareness #AIcrawlers #HackerNews #ngated

  8. FYI: IAB Australia forces every crawler into one of four verdicts: Just 2.6% of AI crawler traffic serves live queries versus 52% for training, IAB Australia finds, ahead of Cloudflare's default block starting in September. ppc.land/iab-australia-forces- #IAB #Australia #AICrawlers #WebTraffic #DigitalMarketing

  9. IAB Australia forces every crawler into one of four verdicts: Just 2.6% of AI crawler traffic serves live queries versus 52% for training, IAB Australia finds, ahead of Cloudflare's default block starting in September. ppc.land/iab-australia-forces- #IABAustralia #AICrawlers #DigitalMarketing #Cloudflare #AdTech

  10. All About Berlin loses 75% of traffic as zero-click searches hit 68%: Affiliate income stayed down 30% despite renegotiated commissions, and Time now routes AI crawlers to a stripped-down copy of its site. Who survives the shift? ppc.land/all-about-berlin-lose #Berlin #TrafficLoss #ZeroClickSearches #AffiliateMarketing #AICrawlers

  11. Welp, now they came for my personal site.

    I already reluctantly added Cloudflare to botwiki.org. But I *really* don't want to have to do that to my site.

    This kind of blows. Thanks AI bros.

    #indieweb #PersonalWebsite #devops #bots #spam #AICrawlers #enshittification

  12. FYI: AI crawlers hit sites 50,000 times per human visit, Cloudflare data shows: AI crawlers reach up to 50,000 visits per referred reader, Cloudflare data shows, as a Munich court strips Google of its liability shield over AI Overviews. ppc.land/ai-crawlers-hit-sites #AICrawlers #Cloudflare #DigitalMarketing #WebTraffic #SEO

  13. FYI: Cloudflare exposes AI crawlers hitting sites 50000 times per visitor: Bot Management customers can now filter traffic by Training, Search or Agent use case and compare operators side by side. Will the data shift licensing talks? ppc.land/cloudflare-exposes-ai #Cloudflare #AICrawlers #BotManagement #WebTraffic #DataAnalytics

  14. ICYMI: Cloudflare exposes AI crawlers hitting sites 50000 times per visitor: Bot Management customers can now filter traffic by Training, Search or Agent use case and compare operators side by side. Will the data shift licensing talks? ppc.land/cloudflare-exposes-ai #Cloudflare #AICrawlers #BotManagement #Cybersecurity #TrafficFilter

  15. AI crawlers hit sites 50,000 times per human visit, Cloudflare data shows: AI crawlers reach up to 50,000 visits per referred reader, Cloudflare data shows, as a Munich court strips Google of its liability shield over AI Overviews. ppc.land/ai-crawlers-hit-sites #AICrawlers #Cloudflare #DataAnalytics #WebTraffic #SEO

  16. Cloudflare exposes AI crawlers hitting sites 50000 times per visitor: Bot Management customers can now filter traffic by Training, Search or Agent use case and compare operators side by side. Will the data shift licensing talks? ppc.land/cloudflare-exposes-ai #Cloudflare #AICrawlers #BotManagement #DigitalMarketing #WebTraffic

  17. #Cloudflare will block #AIcrawlers from ad-supported pages starting 15 September, unless site owners opt in. This move aims to address the issue of #AI companies using #webcontent for #training without compensating #publishers, potentially harming the open web. Cloudflare is also introducing a “#PayPerUse” system to compensate publishers when their content influences AI answers. thenextweb.com/news/cloudflare #tech #media #news

  18. ICYMI: The user agent strings every SEO and site owner needs right now: A technical reference for GPTBot, OAI-SearchBot, ClaudeBot, PerplexityBot, and Bingbot - the exact user agent strings shaping how AI crawls the web in 2026. ppc.land/the-user-agent-string #SEO #WebCrawling #UserAgent #DigitalMarketing #AICrawlers

  19. It looks like whoever is behind this has finally come for one of my sites. Manually blocking datacenter IPs has worked for a while, but it seems like the traffic is now coming either from residential IPs, or the IPs are just getting spoofed.

    Real conundrum. I didn't want to have to set up Cloudflare for the site, and I have to wonder how good they are at catching this anyway, but it might just come to that.

    Anyone else has been dealing with this?

    #indieweb #PersonalWebsite #devops #bots #spam #AICrawlers

  20. New York passes bill forcing AI crawlers to identify themselves to news sites: New York's Assembly passed A11292 on June 5, 2026, requiring AI crawlers to disclose identity and purpose to news publishers or face $15,000-per-day penalties. ppc.land/new-york-passes-bill- #NewYork #AICrawlers #DataPrivacy #TechNews #Legislation

  21. FYI: Kinsta adds free bot protection to all WordPress plans: Kinsta on June 9 launched Bot Protection for all plans, giving WordPress owners control over AI crawlers and automated traffic inside MyKinsta at no added cost. ppc.land/kinsta-adds-free-bot- #Kinsta #WordPress #BotProtection #AICrawlers #TechNews

  22. @kaffeeringe das hatte ich auch schon, ebenfalls der Hinweis, per htaccess einfach gewisse/die meisten IP Netze zu sperren.
    Was ich aber nicht verstehe: wenn so ein KI Bot eine Website "abgrast" sollte das doch mit einem Aufruf pro Seite erledigt sein. Wozu müssen die das 10.000e Male machen? Oder ist diese Technik des Einlesens einfach genauso schlecht und unausgereift, wie die Ausgaben?
    #KIBot #KICrawler #AIBot #AICrawler #AICrawlers #AI

  23. Attaque crawlers IA en cours sur le serveur gayfr.online : 98% des challenges #Anubis sont rejetés (donc sont des bots), plus de 800 000 requêtes sur la semaine.

    Du coup aucun impact sur les performances ni la disponibilité. Il reste juste un bot vietnamien malin qui passe sous les radars et interroge le service RSS avec plusieurs secondes ou minutes d'intervalle. Pas gênant mais je lui réglerai son compte.

    #gayfr #IA #AI #ArtificialIntelligence #IntelligenceArtificielle #AICrawlers #AIBots