home.social

#aiscrapers — Public Fediverse posts

Live and recent posts from across the Fediverse tagged #aiscrapers, aggregated by home.social.

  1. FYI: Fox buys Roku, Publicis and TTD end feud, UK publishers sue AI scrapers: Fox's $22bn Roku deal reshapes CTV, Publicis and The Trade Desk end dispute, UK publishers bill AI scrapers £500 per scraped article through county courts. ppc.land/fox-buys-roku-publici #FoxBuysRoku #CTV #Publicis #TheTradeDesk #AIscrapers

  2. FYI: Fox buys Roku, Publicis and TTD end feud, UK publishers sue AI scrapers: Fox's $22bn Roku deal reshapes CTV, Publicis and The Trade Desk end dispute, UK publishers bill AI scrapers £500 per scraped article through county courts. ppc.land/fox-buys-roku-publici #FoxBuysRoku #CTV #Publicis #TheTradeDesk #AIscrapers

  3. AI scrapers are driving up your hosting costs while real users are left waiting in the digital lobby 🤖

    It is time to take the pressure off your infrastructure by using robots.txt and cache normalization to manage those thirsty bots 💡

    We are sharing how to set sane application limits so your site stays fast for humans and does not turn into a villain story for your budget/.

    👉 developer.upsun.com/posts/insi

    #WebPerformance #DevOps #AIScrapers #TechInsights

  4. AI scrapers are driving up your hosting costs while real users are left waiting in the digital lobby 🤖

    It is time to take the pressure off your infrastructure by using robots.txt and cache normalization to manage those thirsty bots 💡

    We are sharing how to set sane application limits so your site stays fast for humans and does not turn into a villain story for your budget/.

    👉 developer.upsun.com/posts/insi

    #WebPerformance #DevOps #AIScrapers #TechInsights

  5. RE: mastodon.social/@gamingonlinux

    #RPCS3, open-source PlayStation 3 emulator and debugger, says on Twitter:

    > PSA: #Tencent is aggressively scraping the Internet to build yet another AI slop chatbot, DDoSing many websites in the process.
    >
    > We've found that, as of last week, their scraping bots can now solve Cloudflare challenges and behave like real users while ignoring robots.txt. In the last 24 hours alone, our website received more than 3 million successful requests from Tencent bot IP addresses, plus another 1 million that were blocked by Cloudflare challenges.
    >
    > These recurring DDoS attacks from Tencent have been going on for over a year, and we have been constantly adjusting our firewall rules to filter them while trying not to impact Tencent's real users. Because that is no longer possible, we're now fully blocking Tencent IP addresses, starting with ASN 132203. We recommend other sysadmins do the same.

    x.com/rpcs3/status/20619460007

    #DDoS #AI #AIChatBot #AIScraping #AIScraper #AIScrapers #PlayStation #PlayStation3 #PS3

  6. RE: mastodon.social/@gamingonlinux

    #RPCS3, open-source PlayStation 3 emulator and debugger, says on Twitter:

    > PSA: #Tencent is aggressively scraping the Internet to build yet another AI slop chatbot, DDoSing many websites in the process.
    >
    > We've found that, as of last week, their scraping bots can now solve Cloudflare challenges and behave like real users while ignoring robots.txt. In the last 24 hours alone, our website received more than 3 million successful requests from Tencent bot IP addresses, plus another 1 million that were blocked by Cloudflare challenges.
    >
    > These recurring DDoS attacks from Tencent have been going on for over a year, and we have been constantly adjusting our firewall rules to filter them while trying not to impact Tencent's real users. Because that is no longer possible, we're now fully blocking Tencent IP addresses, starting with ASN 132203. We recommend other sysadmins do the same.

    x.com/rpcs3/status/20619460007

    #DDoS #AI #AIChatBot #AIScraping #AIScraper #AIScrapers #PlayStation #PlayStation3 #PS3

  7. RE: mastodon.social/@gamingonlinux

    #RPCS3, open-source PlayStation 3 emulator and debugger, says on Twitter:

    > PSA: #Tencent is aggressively scraping the Internet to build yet another AI slop chatbot, DDoSing many websites in the process.
    >
    > We've found that, as of last week, their scraping bots can now solve Cloudflare challenges and behave like real users while ignoring robots.txt. In the last 24 hours alone, our website received more than 3 million successful requests from Tencent bot IP addresses, plus another 1 million that were blocked by Cloudflare challenges.
    >
    > These recurring DDoS attacks from Tencent have been going on for over a year, and we have been constantly adjusting our firewall rules to filter them while trying not to impact Tencent's real users. Because that is no longer possible, we're now fully blocking Tencent IP addresses, starting with ASN 132203. We recommend other sysadmins do the same.

    x.com/rpcs3/status/20619460007

    #DDoS #AI #AIChatBot #AIScraping #AIScraper #AIScrapers #PlayStation #PlayStation3 #PS3

  8. RE: mastodon.social/@gamingonlinux

    , open-source PlayStation 3 emulator and debugger, says on Twitter:

    > PSA: is aggressively scraping the Internet to build yet another AI slop chatbot, DDoSing many websites in the process.
    >
    > We've found that, as of last week, their scraping bots can now solve Cloudflare challenges and behave like real users while ignoring robots.txt. In the last 24 hours alone, our website received more than 3 million successful requests from Tencent bot IP addresses, plus another 1 million that were blocked by Cloudflare challenges.
    >
    > These recurring DDoS attacks from Tencent have been going on for over a year, and we have been constantly adjusting our firewall rules to filter them while trying not to impact Tencent's real users. Because that is no longer possible, we're now fully blocking Tencent IP addresses, starting with ASN 132203. We recommend other sysadmins do the same.

    x.com/rpcs3/status/20619460007

  9. RE: mastodon.social/@gamingonlinux

    #RPCS3, open-source PlayStation 3 emulator and debugger, says on Twitter:

    > PSA: #Tencent is aggressively scraping the Internet to build yet another AI slop chatbot, DDoSing many websites in the process.
    >
    > We've found that, as of last week, their scraping bots can now solve Cloudflare challenges and behave like real users while ignoring robots.txt. In the last 24 hours alone, our website received more than 3 million successful requests from Tencent bot IP addresses, plus another 1 million that were blocked by Cloudflare challenges.
    >
    > These recurring DDoS attacks from Tencent have been going on for over a year, and we have been constantly adjusting our firewall rules to filter them while trying not to impact Tencent's real users. Because that is no longer possible, we're now fully blocking Tencent IP addresses, starting with ASN 132203. We recommend other sysadmins do the same.

    x.com/rpcs3/status/20619460007

    #DDoS #AI #AIChatBot #AIScraping #AIScraper #AIScrapers #PlayStation #PlayStation3 #PS3

  10. Why is #twitter not properly identifying itself as a bot when trying to scrape my website? (69.12.56.0/21 is AS63179 is Twitter)

    Could it be cause they're a malicious party training an #aibot?

    (This is extremely low-intensity, but based on the combination of this specific UA and the pages they're trying to reach, I've seen them before, coming in from residential proxies.)

    The funny thing is that bots identifying as bots and observing robots.txt would actually be allowed to reach those particular pages.

    #aiscrapers #scrapers

  11. Why is #twitter not properly identifying itself as a bot when trying to scrape my website? (69.12.56.0/21 is AS63179 is Twitter)

    Could it be cause they're a malicious party training an #aibot?

    (This is extremely low-intensity, but based on the combination of this specific UA and the pages they're trying to reach, I've seen them before, coming in from residential proxies.)

    The funny thing is that bots identifying as bots and observing robots.txt would actually be allowed to reach those particular pages.

    #aiscrapers #scrapers

  12. Czech publishers get new robots.txt shield against AI scrapers: SPIR on March 19 updated its standard for Czech online publishers to opt out of AI text and data mining, adding real-time response crawlers to the scope of the robots.txt framework. ppc.land/czech-publishers-get- #CzechPublishing #AIScrapers #RobotsTxt #DataMining #OnlinePrivacy

  13. #AIBots may lead to the end of the internet as we know it

    In recent weeks, #OpenDemocracy’s website has been repeatedly brought down by an army of bots. We’re not the only ones

    Matthew Linares
    20 February 2026

    Excerpt: "Slater explained that 'the traffic often arrives through anonymous residential IPs', referring to residential proxy networks that route internet traffic through intermediary servers using IP addresses assigned by internet service providers to real homeowners. This, he said, makes it 'hard to distinguish ‘normal users’ from automated collection'. [That's not right and needs to be changed!!!]

    " 'We're being forced into permanent defence mode. #ResidentialProxyNetworks let #AIScrapers hide in plain sight, rotate identities, and extract data at scale. That shifts real costs onto projects that exist to serve people, not feed training pipelines."

    Read more:
    opendemocracy.net/en/ai-chatbo

    #AISucks #AI #DataMining #Internet #Websites #TechNews #AI #ArtificialIntelligence #BigTech #TechBros

  14. #AIBots may lead to the end of the internet as we know it

    In recent weeks, #OpenDemocracy’s website has been repeatedly brought down by an army of bots. We’re not the only ones

    Matthew Linares
    20 February 2026

    Excerpt: "Slater explained that 'the traffic often arrives through anonymous residential IPs', referring to residential proxy networks that route internet traffic through intermediary servers using IP addresses assigned by internet service providers to real homeowners. This, he said, makes it 'hard to distinguish ‘normal users’ from automated collection'. [That's not right and needs to be changed!!!]

    " 'We're being forced into permanent defence mode. #ResidentialProxyNetworks let #AIScrapers hide in plain sight, rotate identities, and extract data at scale. That shifts real costs onto projects that exist to serve people, not feed training pipelines."

    Read more:
    opendemocracy.net/en/ai-chatbo

    #AISucks #AI #DataMining #Internet #Websites #TechNews #AI #ArtificialIntelligence #BigTech #TechBros

  15. Webspace Invaders - Matthias Ott

    (…) In their hunger for data to train their large language models, companies from all over the world are systematically harvesting every word I’ve ever published, feeding it into their language models to keep them fresh – and the side effect, the collateral damage, is that Kevin in Montreal now can’t read my articles because my hosting provider decided the solution was to block Canada and half the rest of the world.
    I sat there staring at those logs for a while. The irony wasn’t lost on me. This is my little corner of the web. My writing. With my weird little style mixer up there in the top right. And now it is simultaneously being strip-mined by AI companies and effectively made inaccessible to actual humans around the world who might want to read it.
    This is where we are in 2026. (…) Yes, the AI companies need to do better. They actually should throttle their scraping to reasonable levels. They actually should respect the limited resources of small sites. They actually should develop industry standards that don’t externalize costs onto individuals who are just trying to share their work. (…) matthiasott.com

    I can't help but getting really really angry about all this and what it does to the web I used to love.

    #ai #aiScrapers #collateraldamage #exploitation #otemporaomores #Web

    https://webrocker.de/?p=29765

  16. Webspace Invaders - Matthias Ott

    (…) In their hunger for data to train their large language models, companies from all over the world are systematically harvesting every word I’ve ever published, feeding it into their language models to keep them fresh – and the side effect, the collateral damage, is that Kevin in Montreal now can’t read my articles because my hosting provider decided the solution was to block Canada and half the rest of the world.
    I sat there staring at those logs for a while. The irony wasn’t lost on me. This is my little corner of the web. My writing. With my weird little style mixer up there in the top right. And now it is simultaneously being strip-mined by AI companies and effectively made inaccessible to actual humans around the world who might want to read it.
    This is where we are in 2026. (…) Yes, the AI companies need to do better. They actually should throttle their scraping to reasonable levels. They actually should respect the limited resources of small sites. They actually should develop industry standards that don’t externalize costs onto individuals who are just trying to share their work. (…) matthiasott.com

    I can't help but getting really really angry about all this and what it does to the web I used to love.

    #ai #aiScrapers #collateraldamage #exploitation #otemporaomores #Web

    https://webrocker.de/?p=29765