home.social

#aiscraping — Public Fediverse posts

Live and recent posts from across the Fediverse tagged #aiscraping, aggregated by home.social.

fetched live
  1. RE: mastodon.online/@mwichary/1170

    Here's a thread about 2 experiments to stop AI from scraping website content by obfuscating the text. Unfortunately both lead to accessibility problems.

    #2 works OK with screen reader software but is unpleasant to read visually.

    #1 is not read by screen readers. The creators mitigate this problem by revealing the text through decryption in JavaScript. But, it's slow for screen reader users with JS enabled, and inaccessible with JS disabled.

    #Accessibility #ScreenReader #JavaScript #AIScraping

  2. “We've seen some creative innovations designed to protect #creativework online in the era of #AIscraping. Nightshade is a tool that adds a 'poison' to digital art files to pollute datasets when images are used for AI training. Now a type foundry and creative agency have done something comparable for #typography, creating a #typeface that anyone can use to shield their work.”

    Edited: Unfortunately, it's got some major accessibility problems -- caneandable.social/@WeirdWrite

    creativebloq.com/design/fonts-

  3. RE: mastodon.social/@gamingonlinux

    #RPCS3, open-source PlayStation 3 emulator and debugger, says on Twitter:

    > PSA: #Tencent is aggressively scraping the Internet to build yet another AI slop chatbot, DDoSing many websites in the process.
    >
    > We've found that, as of last week, their scraping bots can now solve Cloudflare challenges and behave like real users while ignoring robots.txt. In the last 24 hours alone, our website received more than 3 million successful requests from Tencent bot IP addresses, plus another 1 million that were blocked by Cloudflare challenges.
    >
    > These recurring DDoS attacks from Tencent have been going on for over a year, and we have been constantly adjusting our firewall rules to filter them while trying not to impact Tencent's real users. Because that is no longer possible, we're now fully blocking Tencent IP addresses, starting with ASN 132203. We recommend other sysadmins do the same.

    x.com/rpcs3/status/20619460007

    #DDoS #AI #AIChatBot #AIScraping #AIScraper #AIScrapers #PlayStation #PlayStation3 #PS3

  4. The Internet Archive just hit one trillion archived web pages—while major news sites block it over AI scraping fears. The irony? We’re losing history to protect it. arstechnica.com/tech-policy/20 #DigitalPreservation #InternetArchive #AIScraping #NyxIsAVirus

  5. @serigala_tropis ahh. AI scraping is also likely on flipboard.

    So I go through the trouble of avoiding ai scraping on my sites, flipboard requires full rss feeds, and suddenly you have a writing honeypot for ai scrapers.

    Sneaky.

    #flipboard #writing #aiscraping

  6. May have to put my rss feeds into flipboard.

    Edit: no. Full rss feeds are a requirement. It's an ai scraper honeypot.

    I will consider the admin overhead.

    I would rather be writing.

    techcrunch.com/2026/04/02/flip

    #flipboard #socialmedia #writing #aiscraping

  7. I heard back on this the other day. The #SJM aka #SJMN aka #mercurynews turned off their feeds due to #aiscraping , as they offered full articles in the feed.

    Ugh. Still a -4 on productivity.

  8. 🚨 Publishers Strike Back: EU Demands “Pay Up” & UK Says “Let Us Opt Out” of AI Search! 🤖💸

    The “wild west” of AI scraping just hit a massive roadblock. In a double-whammy update from Europe, lawmakers are finally drawing a line in the sand. If you own a website, create content, or work in SEO, the game is changing fast.

    Here is the breakdown of the two massive stories shaking up the tech world this week.

    nbloglinks.com/publishers-stri

    #AI #AIScraping #publishers #AIContent ##AIcontrol #UK #EU #technews #SEO

  9. No surprise here about #aiscraping. The question
    Is if it's efficient and produces value that redeems the cost.

    Yes, i laughed at the #typo in the title. Spellcheck alone should have caught it. 🤣

    #ai

    wired.com/story/ai-bots-are-no

  10. There is an ugly truth in this. I block ai scraping on my sites, but I am not blind to the fact that ai scraping can still happen.

    It's not an industry that pays more than lip service to social responsibility.

    They don't even need the data. That is how misguided it is.

    #ai #aiscraping

    axios.com/2026/02/02/iab-ai-ac

  11. Although the bland the "A.I." generated voice is detestable in the extreme the overall concept is mildly amusing if unoriginal, plus, there's a special treat for #DoctorWho fans in trying to fathom where the contents of the Laser Tracking Room were illicitly scraped from...

    youtube.com/watch?v=sZkB11pO9R8

    #cats #Caturday #AIart #TARDIS #copyright #AIscraping

  12. Scraping for AI training may or may not be legal. But the effort crawlers put into evading detection and blocking is a smoking gun, an admission this scraping is not fair.

    arstechnica.com/tech-policy/20

    #AIscraping #scrapers #ai

  13. Today, Meta's list of sites they've targeted for training their AI was leaked. We're on their list.

    I do everything possible to block AI bots. I use Cloudflare AI bot protection. I block what I can. I don't know if they actually get to read anything, but they want to read us.

    dropsitenews.com/p/meta-facebo

    #Meta #AIscraping

  14. Habe den aktualisierten AGBs von #vinted widersprochen, da sie von Datenauswertung via #Ki #AIscraping sprechen, also Nutzerdaten damit verarbeiten möchten. Da der Account sowieso schon deaktiviert war, habe ich zusätzlich um Löschung gebeten. Nun kriege ich die Antwort, dass vinted ein legitimes Interesse hätte und laut Article 17 of the EU General Data Protection Regulation (GDPR) so ziemlich alle meine Daten weiterhin auswerten darf. Ich könne ja juristische Schritte gehen. (1/2)

  15. #KI randaliert im Netz 🤖🪓 – #Admins halten dagegen 🦸

    Meine @campact -Kolumne aus Mai ist heute tagesaktuell dabei!

    > Herzlichen Dank an alle Admins, die unermüdlich dafür kämpfen, uns Nutzende und den Planeten vor der Gier von KI zu schützen. Ich hoffe, dieser Text ist ein Beitrag für mehr Verständnis zu diesem Thema.

    👉 blog.campact.de/2025/05/ki-ran

    #SysAdmins #SystemadminAppreciationDay #FediAdmins #AI #KIScraping
    #AIScraping #TDM #AdminLeiden #MastoAdmin #DataPoisoning #aitxt #GPT #GreenIT

  16. 🔍 / #software / #automation / #scraping

    You can build some pretty insane applications using just #LLMs, even if you don't really know what you're doing. But what separates a good AI app from a great AI app is one thing, and that's data.

    🐱🔗 laravista.altervista.org/CatLi

    #catlink #SoftwareAutomation #SoftwareAutomationScraping #Python #BrightData #AIScraping #AI

  17. News Summary: Cloudflare Launches Pay Per Crawl for AI Scraping; Amazon Hits One Million Robots

    You’ve heard, of course, of pay-per-view. And we are used to streaming revenue on a pay-per basis from the likes of Audible and Spotify. This week has seen the launch (admittedly at the moment in beta) of possibly the most transformative source of pay-per revenue…
    selfpublishingadvice.org/cloud

    #AIscraping #Amazonrobots #Cloudflare #generativeAI #PayPerCrawl
    @indieauthors

  18. A website appears to be scraping hashtags and creating AI articles, and then replying to the OG post

    It stole one of my posts (oldfriends.live/@paul/11477009) for its AI created article then spammed me from [email protected]

    It's doing it with #HashTagGames tags and other trending hashtags.

    Edit: making links dead as it appears to serve malware now: www.trend247daily.com/articles

    #MastoAdmin

    Article created from scraped post: www.trend247daily.com/article/mastering-the-art-of-the-productive-day-wake-up-look-busy-go-to-bed

    See this thread above, unless the AI content spammer deletes its reply and breaks the thread.

    I don't know where it is getting its content, from it's Mastodon Account ( [email protected] ) account, rss, or the API. If it has an application I would hope [email protected] and [email protected] would shut it down from scraping the API.

    #Spam #Fediblock #AIScraping

  19. The web-scraping is aggressive not just to hoard training data, but also to keep other AI bots from doing the same.

    They're not satisfied with stealing all your content, they also want exclusivity by any means necessary.

    nature.com/articles/d41586-025

    #ai #aiscraping #aichatbots