#datascraping — Public Fediverse posts
Live and recent posts from across the Fediverse tagged #datascraping, aggregated by home.social.
-
Yet another reason for me to be disgusted with Amazon. Why am I not surprised?
https://www.404media.co/we-tracked-a-shipment-of-rare-books-it-ended-at-an-amazon-ai-training-facility/
#books #authors #writers #Amazon #DataPrivacy #AI #SurveillanceEconomy #DataScraping #CopyrightInfringement #PreserveCulture #AntiAmazon #AmazonSucks -
Yet another reason for me to be disgusted with Amazon. Why am I not surprised?
https://www.404media.co/we-tracked-a-shipment-of-rare-books-it-ended-at-an-amazon-ai-training-facility/
#books #authors #writers #Amazon #DataPrivacy #AI #SurveillanceEconomy #DataScraping #CopyrightInfringement #PreserveCulture #AntiAmazon #AmazonSucks -
Yet another reason for me to be disgusted with Amazon. Why am I not surprised?
https://www.404media.co/we-tracked-a-shipment-of-rare-books-it-ended-at-an-amazon-ai-training-facility/
#books #authors #writers #Amazon #DataPrivacy #AI #SurveillanceEconomy #DataScraping #CopyrightInfringement #PreserveCulture #AntiAmazon #AmazonSucks -
A hacker has leaked a 14.5GB database containing 7.3 million scraped #Chess.com user records. Hackread’s analysis found 4.6 million email entries, profile details, and recent login data.
Listen/Read: https://hackread.com/hacker-leaks-7-million-scraped-chess-com-user-records/
-
A hacker has leaked a 14.5GB database containing 7.3 million scraped #Chess.com user records. Hackread’s analysis found 4.6 million email entries, profile details, and recent login data.
Listen/Read: https://hackread.com/hacker-leaks-7-million-scraped-chess-com-user-records/
-
A hacker has leaked a 14.5GB database containing 7.3 million scraped #Chess.com user records. Hackread’s analysis found 4.6 million email entries, profile details, and recent login data.
Listen/Read: https://hackread.com/hacker-leaks-7-million-scraped-chess-com-user-records/
-
https://winbuzzer.com/2026/07/19/hacked-code-maps-sunos-alleged-music-scraping-pipeline-xcxwbn/
Shared hacked code details how Suno collected millions of music clips and lyrics, adding disputed scraping evidence to ongoing copyright litigation.
#AI #AIMusic #Suno #ContentScraping #DataScraping #AITraining #AIAudio #AudioGeneration #Copyright #FairUse
-
https://winbuzzer.com/2026/07/19/hacked-code-maps-sunos-alleged-music-scraping-pipeline-xcxwbn/
Shared hacked code details how Suno collected millions of music clips and lyrics, adding disputed scraping evidence to ongoing copyright litigation.
#AI #AIMusic #Suno #ContentScraping #DataScraping #AITraining #AIAudio #AudioGeneration #Copyright #FairUse
-
https://winbuzzer.com/2026/07/19/hacked-code-maps-sunos-alleged-music-scraping-pipeline-xcxwbn/
Shared hacked code details how Suno collected millions of music clips and lyrics, adding disputed scraping evidence to ongoing copyright litigation.
#AI #AIMusic #Suno #ContentScraping #DataScraping #AITraining #AIAudio #AudioGeneration #Copyright #FairUse
-
https://winbuzzer.com/2026/07/19/patreon-says-it-has-begun-blocking-ai-training-crawlers-xcxwbn/
Patreon has started blocking AI training crawlers while preserving search indexing that can direct visitors back to creators.
#AI #Patreon #AITraining #ContentScraping #DataScraping #CloudFlare
-
https://winbuzzer.com/2026/07/19/patreon-says-it-has-begun-blocking-ai-training-crawlers-xcxwbn/
Patreon has started blocking AI training crawlers while preserving search indexing that can direct visitors back to creators.
#AI #Patreon #AITraining #ContentScraping #DataScraping #CloudFlare
-
https://winbuzzer.com/2026/07/19/patreon-says-it-has-begun-blocking-ai-training-crawlers-xcxwbn/
Patreon has started blocking AI training crawlers while preserving search indexing that can direct visitors back to creators.
#AI #Patreon #AITraining #ContentScraping #DataScraping #CloudFlare
-
FYI: EDPB blocks AI firms from using consent as an excuse to scrape: Adopted 07 July 2026, the guidance sets a 3-condition legitimate interest test for firms scraping data, and public consultation runs until 30 October 2026. https://ppc.land/edpb-blocks-ai-firms-from-using-consent-as-an-excuse-to-scrape/ #EDPB #AIethics #DataPrivacy #Consent #DataScraping
-
FYI: EDPB blocks AI firms from using consent as an excuse to scrape: Adopted 07 July 2026, the guidance sets a 3-condition legitimate interest test for firms scraping data, and public consultation runs until 30 October 2026. https://ppc.land/edpb-blocks-ai-firms-from-using-consent-as-an-excuse-to-scrape/ #EDPB #AIethics #DataPrivacy #Consent #DataScraping
-
FYI: EDPB blocks AI firms from using consent as an excuse to scrape: Adopted 07 July 2026, the guidance sets a 3-condition legitimate interest test for firms scraping data, and public consultation runs until 30 October 2026. https://ppc.land/edpb-blocks-ai-firms-from-using-consent-as-an-excuse-to-scrape/ #EDPB #AIethics #DataPrivacy #Consent #DataScraping
-
News from the walled gardens:
See also:
Apparently now GitHub is trying to pull a Meta/Linkedin a la login walls.
Their argument is exploitation and excess AI load.
While both might be true, this might not be the entire argument and probably defies the purpose of the open access GitHub.1/2
#ai #datascraping #git #github #opensource #enshittification #hackernews #linkedIn #instagram #meta
-
News from the walled gardens:
See also:
Apparently now GitHub is trying to pull a Meta/Linkedin a la login walls.
Their argument is exploitation and excess AI load.
While both might be true, this might not be the entire argument and probably defies the purpose of the open access GitHub.1/2
#ai #datascraping #git #github #opensource #enshittification #hackernews #linkedIn #instagram #meta
-
News from the walled gardens:
See also:
Apparently now GitHub is trying to pull a Meta/Linkedin a la login walls.
Their argument is exploitation and excess AI load.
While both might be true, this might not be the entire argument and probably defies the purpose of the open access GitHub.1/2
#ai #datascraping #git #github #opensource #enshittification #hackernews #linkedIn #instagram #meta
-
It's turned on by default, and you won't be notified when your likeness is used. Private accounts are locked down and safe.
How to opt out of this #DataScraping: Settings ➡️ Sharing and reuse ➡️ Toggle off the AI reuse settings for Posts and Reels -
Learn everything you need to know about Data Scraping via these 70 free HackerNoon blog posts. https://hackernoon.com/70-blog-posts-to-learn-about-data-scraping #datascraping
-
Learn everything you need to know about Data Scraping via these 70 free HackerNoon blog posts. https://hackernoon.com/70-blog-posts-to-learn-about-data-scraping #datascraping
-
Learn everything you need to know about Data Scraping via these 70 free HackerNoon blog posts. https://hackernoon.com/70-blog-posts-to-learn-about-data-scraping #datascraping
-
https://winbuzzer.com/2026/04/09/youtubers-sue-apple-for-scraping-videos-to-train-ai-xcxwbn/
YouTubers Sue Apple for Scraping Videos to Train AI Models
#AI #Apple #YouTube #GenAI #AITraining #AIModels #DataScraping #Copyright #Lawsuits #FairUse #BigTech #AIEthics
-
https://winbuzzer.com/2026/04/09/youtubers-sue-apple-for-scraping-videos-to-train-ai-xcxwbn/
YouTubers Sue Apple for Scraping Videos to Train AI Models
#AI #Apple #YouTube #GenAI #AITraining #AIModels #DataScraping #Copyright #Lawsuits #FairUse #BigTech #AIEthics
-
https://winbuzzer.com/2026/04/09/youtubers-sue-apple-for-scraping-videos-to-train-ai-xcxwbn/
YouTubers Sue Apple for Scraping Videos to Train AI Models
#AI #Apple #YouTube #GenAI #AITraining #AIModels #DataScraping #Copyright #Lawsuits #FairUse #BigTech #AIEthics
-
I'm slightly creeped out but not surprised. I was editing a music score on my laptop recently and I added an instruction to play the piece "robotic". The next time I logged into Indeed, the first job recommendation to come up is for Robotics Operator. Is Indeed scraping data from my recent documents for keywords?
Always check your firewall.
-
I'm slightly creeped out but not surprised. I was editing a music score on my laptop recently and I added an instruction to play the piece "robotic". The next time I logged into Indeed, the first job recommendation to come up is for Robotics Operator. Is Indeed scraping data from my recent documents for keywords?
Always check your firewall.
-
I'm slightly creeped out but not surprised. I was editing a music score on my laptop recently and I added an instruction to play the piece "robotic". The next time I logged into Indeed, the first job recommendation to come up is for Robotics Operator. Is Indeed scraping data from my recent documents for keywords?
Always check your firewall.
-
🚀 Want to dive into data scraping? Check out MediaCrawler by NanmiCoder! This powerful tool lets you harvest comments and content from popular platforms like Xiaohongshu, Douyin, Kuaishou, Bilibili, Weibo, Baidu Tieba, and Zhihu. Perfect for learning and research—just remember to use it responsibly! 📊💻
Explore more here: https://github.com/NanmiCoder/MediaCrawler
-
The “17.5 million Instagram user data leak” making rounds in 2026? Old news
The data from 2022 was already leaked in 2023.
We broke down all 3 dumps - same records
Don’t fall for clickbait reports!
Read: https://hackread.com/instagram-user-data-leak-scraped-records-2022/
-
The “17.5 million Instagram user data leak” making rounds in 2026? Old news
The data from 2022 was already leaked in 2023.
We broke down all 3 dumps - same records
Don’t fall for clickbait reports!
Read: https://hackread.com/instagram-user-data-leak-scraped-records-2022/
-
The “17.5 million Instagram user data leak” making rounds in 2026? Old news
The data from 2022 was already leaked in 2023.
We broke down all 3 dumps - same records
Don’t fall for clickbait reports!
Read: https://hackread.com/instagram-user-data-leak-scraped-records-2022/
-
LinkedIn's 2025 Data Crisis: 4.3 Billion Records Leaked, Risks Rise https://www.webpronews.com/linkedins-2025-data-crisis-4-3-billion-records-leaked-risks-rise/ #cybersecurity #LinkedIn #DataTheft #scams #spam #DataScraping
-
LinkedIn's 2025 Data Crisis: 4.3 Billion Records Leaked, Risks Rise https://www.webpronews.com/linkedins-2025-data-crisis-4-3-billion-records-leaked-risks-rise/ #cybersecurity #LinkedIn #DataTheft #scams #spam #DataScraping
-
New York Times Sues Perplexity AI for Copyright Infringement and ‘Trademark Tarnishment’
#AI #Copyright #PerplexityAI #NYT #GenAI #SearchEngines #RAG #MediaLaw #IntellectualProperty #Hallucinations #Journalism #DataScraping
-
New York Times Sues Perplexity AI for Copyright Infringement and ‘Trademark Tarnishment’
#AI #Copyright #PerplexityAI #NYT #GenAI #SearchEngines #RAG #MediaLaw #IntellectualProperty #Hallucinations #Journalism #DataScraping
-
New York Times Sues Perplexity AI for Copyright Infringement and ‘Trademark Tarnishment’
#AI #Copyright #PerplexityAI #NYT #GenAI #SearchEngines #RAG #MediaLaw #IntellectualProperty #Hallucinations #Journalism #DataScraping
-
Reddit Sues Perplexity and Data Scrapers for 'Industrial-Scale' AI Content Theft
#AI #Reddit #Perplexity #Lawsuit #DataScraping #Copyright #TechLaw #DMCA #AIEthics #BigTech #IntellectualProperty #SerpApi #Oxylabs #Litigation
-
Reddit Sues Perplexity and Data Scrapers for 'Industrial-Scale' AI Content Theft
#AI #Reddit #Perplexity #Lawsuit #DataScraping #Copyright #TechLaw #DMCA #AIEthics #BigTech #IntellectualProperty #SerpApi #Oxylabs #Litigation
-
Reddit Sues Perplexity and Data Scrapers for 'Industrial-Scale' AI Content Theft
#AI #Reddit #Perplexity #Lawsuit #DataScraping #Copyright #TechLaw #DMCA #AIEthics #BigTech #IntellectualProperty #SerpApi #Oxylabs #Litigation
-
Reddit sues data scrapers and Perplexity over unauthorized content access: Reddit filed a lawsuit on October 22, 2025, against SerpApi, Oxylabs, AWMProxy, and Perplexity AI for circumventing security measures to scrape platform data. https://ppc.land/reddit-sues-data-scrapers-and-perplexity-over-unauthorized-content-access/ #Reddit #Lawsuit #DataScraping #Privacy #Cybersecurity
-
Stable?
I think I’m at a place where I can write about this now.
If you’re a faithfully follower of my blog, you may have noticed a degradation in performance over the past few weeks. Writing and posting to the blog has also become maddening during this time, as I would get frequent “You are offline” messages from WordPress, images wouldn’t update reliably, and at times I couldn’t connect to the site at all.
Our domains were living on a shared server via a hosting provider in Canada; there were many websites hosted on our meager shared virtual machine (the ‘cloud’, if you will).
When I asked about the performance issues I was seeing, the hosting provider let me know there were other sites on the server getting hammered by bots and the like, and they attributed it to that.
Over this past weekend everything to do with anything jpnearl.com, the websites, the email, all of it, either came to a screeching halt or disappeared from the Internet completely. I raised another ticket and the hosting company promptly responded.
Our domain was being overwhelmed by bots and AI systems scraping my blog for training data. It was to the point that no one could even get into the server to try to do anything.
This is when I put up the generic “Hello, world” message that was there for a couple of days.
After a big ding in our household budget, our domains were moved over to a standalone server. The standalone server is much more robust than the shared VM we called our virtual home. And all seemed well for a couple of hours.
The bots and other AI scraping devices found us and started scraping any and all data it could find in full force. Things started crashing again.
When it comes to hosting my own domain, email is my primary concern, with the blogs coming in second. I took down the blogs again to get email working. The hosting company’s support team jumped onto the server and made numerous adjustments to the configuration to help mitigate some of the automated attacks that were occurring. I also went ahead and put the entire domain behind CloudFlare, which is designed to keep this sort of thing at bay.
Don’t be surprised if you get asked if you’re a human once in a while.
I also cleaned up a lot of outdated WordPress plugins I had installed over the year. In addition, I cleaned out a lot of cruft in the underlying file system; this domain has been around for over 20 years and there’s some files I’ve thrown on the server that I haven’t thought about in a long time, but the likes of ChatGPT found them very interesting.
I believe our migration is complete and the security around the server is stronger than it has ever been before. I was thinking I would completely turn off integration with the Fediverse, but I determined that wasn’t an issue and have turned it back on. I know several folks that follow along via Mastodon and the like. I don’t want to lose my connection with them.
The Internet of 2025 is nothing as it was intended to be and it’s primarily become an infestation of bots talking to bots and A.I. Large Language Models raping as much data as it can from sources all over the world all in the name of “training”. When people talk about the Internet being dead, I completely agree. It’s a shame, because back when President Clinton was talking about the “Information Superhighway”, I thought connecting computers together would enrich, enlighten, and teach us so many new things.
Never once did I think I would have to reboot the cat’s litter box because it is connected to the Internet.
Since the rebuilding of the support mechanisms around my blog has been a fairly pricey endeavor, it has prompted me to double down on what many consider to be an outdated mode of communication: long form writing on a personal blog.
I am focused more than ever on keeping this (repolished) nook on the Internet alive and well. At least until the next hosting bill arrives in a year or so.
-
Cloudflare Overhauls Web’s AI Rulebook with New Robots.txt ‘Content Signals’
#AI #Cloudflare #RobotsTxt #DataScraping #Publishing #GenerativeAI
-
Cloudflare Overhauls Web’s AI Rulebook with New Robots.txt ‘Content Signals’
#AI #Cloudflare #RobotsTxt #DataScraping #Publishing #GenerativeAI
-
Cloudflare Overhauls Web’s AI Rulebook with New Robots.txt ‘Content Signals’
#AI #Cloudflare #RobotsTxt #DataScraping #Publishing #GenerativeAI
-
LinkedIn, the social media titan known for its riveting inspirational #quotes and unsolicited connections, is now channeling its inner #superhero, battling the dastardly villains of data scraping. 🦸♂️💼 Apparently, charging $15k for harvested data is a crime—unless you're #LinkedIn, of course. 🤑🔍
https://therecord.media/linkedin-sues-data-scraping-company #DataScraping #SocialMedia #Crime #HackerNews #ngated -
LinkedIn, the social media titan known for its riveting inspirational #quotes and unsolicited connections, is now channeling its inner #superhero, battling the dastardly villains of data scraping. 🦸♂️💼 Apparently, charging $15k for harvested data is a crime—unless you're #LinkedIn, of course. 🤑🔍
https://therecord.media/linkedin-sues-data-scraping-company #DataScraping #SocialMedia #Crime #HackerNews #ngated -
Cloudflare launches Content Signals Policy to fight AI crawlers and scrapers
https://web.brid.gy/r/https://nerds.xyz/2025/09/cloudflare-content-signals-policy-ai-crawlers/
-
Cloudflare launches Content Signals Policy to fight AI crawlers and scrapers
https://web.brid.gy/r/https://nerds.xyz/2025/09/cloudflare-content-signals-policy-ai-crawlers/
-
Perplexity Fires Back at Cloudflare, Denying ‘Stealth Crawler’ Accusations
#AI #Cloudflare #Perplexity #WebCrawling #AIethics #DataScraping #SearchEngines #Web #AISearch
-
Perplexity Fires Back at Cloudflare, Denying ‘Stealth Crawler’ Accusations
#AI #Cloudflare #Perplexity #WebCrawling #AIethics #DataScraping #SearchEngines #Web #AISearch
-
Perplexity Fires Back at Cloudflare, Denying ‘Stealth Crawler’ Accusations
#AI #Cloudflare #Perplexity #WebCrawling #AIethics #DataScraping #SearchEngines #Web #AISearch
-
Cloudflare Accuses Perplexity of Using ‘Stealth Crawlers’ to Evade Web Standards
#AI #PerplexityAI #Cloudflare #DataScraping #AIEthics #WebSecurity
-
Cloudflare Accuses Perplexity of Using ‘Stealth Crawlers’ to Evade Web Standards
#AI #PerplexityAI #Cloudflare #DataScraping #AIEthics #WebSecurity
-
Cloudflare Accuses Perplexity of Using ‘Stealth Crawlers’ to Evade Web Standards
#AI #PerplexityAI #Cloudflare #DataScraping #AIEthics #WebSecurity
-
BBC Threatens Lawsuit against Perplexity AI over Verbatim Copying of Content
#AI #Copyright #PerplexityAI #BBC #TechLaw #Media #GenAI #DataScraping #FairUse #IntellectualProperty #AIethics
-
BBC Threatens Lawsuit against Perplexity AI over Verbatim Copying of Content
#AI #Copyright #PerplexityAI #BBC #TechLaw #Media #GenAI #DataScraping #FairUse #IntellectualProperty #AIethics
-
Anthropic Sued by Reddit for Unauthorized Use of AI Training Data
#AI #Reddit #Anthropic #AILawsuit #DataScraping #AIethics #TechLaw #DataRights #Copyright
-
Anthropic Sued by Reddit for Unauthorized Use of AI Training Data
#AI #Reddit #Anthropic #AILawsuit #DataScraping #AIethics #TechLaw #DataRights #Copyright
-
AI Crawlers Overwhelm Open-Source Projects, Forcing Developers to Block Entire Countries
#AI #Web #Robotstxt #AIScraping #OpenSource #Cybersecurity #DataScraping #Scraping #WebScraping