#scrapers — Public Fediverse posts
Live and recent posts from across the Fediverse tagged #scrapers, aggregated by home.social.
-
75 000+ IP au tapis ! ⛔️🛡️
Et encore, ce chiffre ne montre que ce qui réussit à passer la première ligne de défense.En amont, mon fichier de blocage #Nginx personnalisé fait déjà un gros tri sélectif :
🛑 #Scrapers #IA, faux Chrome et outils automatisés reçoivent un #403 immédiat.
🟢 Les instances du #Fediverse ( #Mastodon, #PeerTube... ), elles, entrent sans problème.Nginx filtre les indésirables, #Fail2Ban verrouille le reste. Résultat : un serveur qui respire et des ressources préservées ! 😎🧹
-
75 000+ IP au tapis ! ⛔️🛡️
Et encore, ce chiffre ne montre que ce qui réussit à passer la première ligne de défense.En amont, mon fichier de blocage #Nginx personnalisé fait déjà un gros tri sélectif :
🛑 #Scrapers #IA, faux Chrome et outils automatisés reçoivent un #403 immédiat.
🟢 Les instances du #Fediverse ( #Mastodon, #PeerTube... ), elles, entrent sans problème.Nginx filtre les indésirables, #Fail2Ban verrouille le reste. Résultat : un serveur qui respire et des ressources préservées ! 😎🧹
-
75 000+ IP au tapis ! ⛔️🛡️
Et encore, ce chiffre ne montre que ce qui réussit à passer la première ligne de défense.En amont, mon fichier de blocage #Nginx personnalisé fait déjà un gros tri sélectif :
🛑 #Scrapers #IA, faux Chrome et outils automatisés reçoivent un #403 immédiat.
🟢 Les instances du #Fediverse ( #Mastodon, #PeerTube... ), elles, entrent sans problème.Nginx filtre les indésirables, #Fail2Ban verrouille le reste. Résultat : un serveur qui respire et des ressources préservées ! 😎🧹
-
75 000+ IP au tapis ! ⛔️🛡️
Et encore, ce chiffre ne montre que ce qui réussit à passer la première ligne de défense.En amont, mon fichier de blocage #Nginx personnalisé fait déjà un gros tri sélectif :
🛑 #Scrapers #IA, faux Chrome et outils automatisés reçoivent un #403 immédiat.
🟢 Les instances du #Fediverse ( #Mastodon, #PeerTube... ), elles, entrent sans problème.Nginx filtre les indésirables, #Fail2Ban verrouille le reste. Résultat : un serveur qui respire et des ressources préservées ! 😎🧹
-
75 000+ IP au tapis ! ⛔️🛡️
Et encore, ce chiffre ne montre que ce qui réussit à passer la première ligne de défense.En amont, mon fichier de blocage #Nginx personnalisé fait déjà un gros tri sélectif :
🛑 #Scrapers #IA, faux Chrome et outils automatisés reçoivent un #403 immédiat.
🟢 Les instances du #Fediverse ( #Mastodon, #PeerTube... ), elles, entrent sans problème.Nginx filtre les indésirables, #Fail2Ban verrouille le reste. Résultat : un serveur qui respire et des ressources préservées ! 😎🧹
-
#AI #scrapers are the PFAS of what was once the fertile ground of an open and honest internet.
In the race of AI to.. (yeah to what?), websites with decades of knowledge are forced to spend enormous amounts of compute and bandwidth defending themselves against an opaque ecosystem of scrapers.
An article about Residential Proxy Networks, “ethically sourced” IP addresses, Iocaine, NetNut-infected apps, Anubis, and the invisible infrastructure s(cr/h)aping the web.
-
#AI #scrapers are the PFAS of what was once the fertile ground of an open and honest internet.
In the race of AI to.. (yeah to what?), websites with decades of knowledge are forced to spend enormous amounts of compute and bandwidth defending themselves against an opaque ecosystem of scrapers.
An article about Residential Proxy Networks, “ethically sourced” IP addresses, Iocaine, NetNut-infected apps, Anubis, and the invisible infrastructure s(cr/h)aping the web.
-
@weeklyOSM Please don’t be alarmed if you’ve seen your browser being tested – even if only briefl
Our website was under enormous strain from #AI #scrapers
To defend against these scrapers, we’ve installed Anubis
https://en.wikipedia.org/wiki/Anubis_(software) -
@weeklyOSM Please don’t be alarmed if you’ve seen your browser being tested – even if only briefl
Our website was under enormous strain from #AI #scrapers
To defend against these scrapers, we’ve installed Anubis
https://en.wikipedia.org/wiki/Anubis_(software) -
CW: Adult content
Reminder & FYI to ppl new to Bluesky & my page👇 reposting “leaked” (stolen & posted without personal) porn is really harmful to sw. It means you support that practice. I can’t and won’t allow that in my space. You shouldn’t either. #scrapers #sw #sexworkers #porn #swsolidarity #spreadtheword
-
I appreciate the wise minds that have jumped into my thread over the past 24 hours.
It's amazing to see the variety of tools and attitudes towards the problem. The more I talk it through and hear from other folks, the more I feel that shutting the door is ultimately a self-own. If I want humans to see my website, I need to focus on making that as simple and unfettered as possible.
I have lots of things I could be doing better for my fellow humans. The trackers from Google and FB alone are my major concern for me. There is also a trade off between being searchable at all, versus blocking the scrapers and removing the trackers.
Any tech "solution" is ultimately imperfect. I wish this was a simpler equation. It does not feel like a "nuance" situation, rather a "you get screwed either way" situation.
#Scrapers #AIbots #WebHosting -
I appreciate the wise minds that have jumped into my thread over the past 24 hours.
It's amazing to see the variety of tools and attitudes towards the problem. The more I talk it through and hear from other folks, the more I feel that shutting the door is ultimately a self-own. If I want humans to see my website, I need to focus on making that as simple and unfettered as possible.
I have lots of things I could be doing better for my fellow humans. The trackers from Google and FB alone are my major concern for me. There is also a trade off between being searchable at all, versus blocking the scrapers and removing the trackers.
Any tech "solution" is ultimately imperfect. I wish this was a simpler equation. It does not feel like a "nuance" situation, rather a "you get screwed either way" situation.
#Scrapers #AIbots #WebHosting -
Strava declara la lucha a los scrapers ayer de su salida a bolsa #antes #BOLSA #declara #guerra #IPO #los #raspado_de_datos #salida #scrapers #Strava #ButterWord #Spanish_News Comenta tu opinión 👇
https://butterword.com/strava-declara-la-lucha-a-los-scrapers-ayer-de-su-salida-a-bolsa/?feed_id=83426&_unique_id=6a1d89e384e07 -
Why is #twitter not properly identifying itself as a bot when trying to scrape my website? (69.12.56.0/21 is AS63179 is Twitter)
Could it be cause they're a malicious party training an #aibot?
(This is extremely low-intensity, but based on the combination of this specific UA and the pages they're trying to reach, I've seen them before, coming in from residential proxies.)
The funny thing is that bots identifying as bots and observing robots.txt would actually be allowed to reach those particular pages.
-
Why is #twitter not properly identifying itself as a bot when trying to scrape my website? (69.12.56.0/21 is AS63179 is Twitter)
Could it be cause they're a malicious party training an #aibot?
(This is extremely low-intensity, but based on the combination of this specific UA and the pages they're trying to reach, I've seen them before, coming in from residential proxies.)
The funny thing is that bots identifying as bots and observing robots.txt would actually be allowed to reach those particular pages.
-
Iocaine and my custom solution aren't good enough. :blobcatbigsob: I'm considering to add to login to my website rewrite as protection against bots.
I would always offer an anonymous session after completing a proof of work (which is also available without JS).
Do you think this is okay? Please don't hesitate to reply!
#website #personalBlog #PersonalSites #indieweb #spam #spamprotection #scrapers #selfhosting #iocaine
-
Iocaine and my custom solution aren't good enough. :blobcatbigsob: I'm considering to add to login to my website rewrite as protection against bots.
I would always offer an anonymous session after completing a proof of work (which is also available without JS).
Do you think this is okay? Please don't hesitate to reply!
#website #personalBlog #PersonalSites #indieweb #spam #spamprotection #scrapers #selfhosting #iocaine
-
Desde afuera todavia se nota cierta latencia, a veces, posiblemente porque no han cesado los ataques de scraping. En la red interna vuela, y en las metricas los servidores no estan bajo carga o demanda altos, estan normales. El problema en ese caso sería que todos esos ataques que el firewall esta bloqueando exitosamente, lo hace recien dentro de la red, por lo que ese trafico ocupa lugar en la conexión dejando menos ancho de banda neto para el tráfico legítimo... veremos si la cosa mejora en los próximos dias #undernet #ataque #bots #scrapers #iabot #peertube
-
Desde afuera todavia se nota cierta latencia, a veces, posiblemente porque no han cesado los ataques de scraping. En la red interna vuela, y en las metricas los servidores no estan bajo carga o demanda altos, estan normales. El problema en ese caso sería que todos esos ataques que el firewall esta bloqueando exitosamente, lo hace recien dentro de la red, por lo que ese trafico ocupa lugar en la conexión dejando menos ancho de banda neto para el tráfico legítimo... veremos si la cosa mejora en los próximos dias #undernet #ataque #bots #scrapers #iabot #peertube
-
🎩🤖 Oh, look, another #GitHub hero has blessed us with a "groundbreaking" #tool to trap #AI #web #scrapers in a "poison pit." Because clearly, what we all need is a #digital Venus flytrap for code 😏. Meanwhile, GitHub's feature salad just keeps growing, because who doesn't love a good menu with more options than a diner? 🍔💻
https://github.com/austin-weeks/miasma #innovation #featureupdate #codinghumor #HackerNews #ngated -
🎩🤖 Oh, look, another #GitHub hero has blessed us with a "groundbreaking" #tool to trap #AI #web #scrapers in a "poison pit." Because clearly, what we all need is a #digital Venus flytrap for code 😏. Meanwhile, GitHub's feature salad just keeps growing, because who doesn't love a good menu with more options than a diner? 🍔💻
https://github.com/austin-weeks/miasma #innovation #featureupdate #codinghumor #HackerNews #ngated -
Miasma: A tool to trap AI web scrapers in an endless poison pit
https://github.com/austin-weeks/miasma
#HackerNews #Miasma #AI #web #scrapers #Endless #pit #Tech #innovation #Open #source
-
Miasma: A tool to trap AI web scrapers in an endless poison pit
https://github.com/austin-weeks/miasma
#HackerNews #Miasma #AI #web #scrapers #Endless #pit #Tech #innovation #Open #source
-
No outages in the latest Apache logs. However, there is plenty of suspicious activity.
The log has 16,033 lines.
Of these, 1,559 lines feature the "RecentChanges" function for my wikis. Which is something regular users _might_ call up from time to time, but I suspect that #scrapers are the more likely culprits.
The vast majority of these requests come from a random assortment of IP addresses, and they usually end with something on the lines of:
"Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/123.0.0.0 Safari/537.36"
So yeah, "anonymous bot nets scraping the Interwebs for nefarious purposes" would be by first guess.
-
No outages in the latest Apache logs. However, there is plenty of suspicious activity.
The log has 16,033 lines.
Of these, 1,559 lines feature the "RecentChanges" function for my wikis. Which is something regular users _might_ call up from time to time, but I suspect that #scrapers are the more likely culprits.
The vast majority of these requests come from a random assortment of IP addresses, and they usually end with something on the lines of:
"Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/123.0.0.0 Safari/537.36"
So yeah, "anonymous bot nets scraping the Interwebs for nefarious purposes" would be by first guess.
-
Army of Bots
For some months now I have a simple detection against "bad" bots in place. Bots that scrape *everything* they find and very likely are vacuuming all the contents they get to feed the data grinders that train the LLMs of the world. Bots that not only ignore the "robots.txt" protocol, but actively see entries in the robots.txt file as an invitation to visit the contents that are listed there as "disallowed".
I always had a hunch that stating addresses in a publicy reachable text file and flagging those as "please stay out of there" wasn't the best idea, but well, it was the only thing we've got back in the days where the only bots out there were the crawlers of the search engines.
(…) There are two important considerations when using /robots.txt:
robots can ignore your /robots.txt. Especially malware robots that scan the web for security vulnerabilities, and email address harvesters used by spammers will pay no attention.
the /robots.txt file is a publicly available file. Anyone can see what sections of your server you don't want robots to use.
robotstxt.orgNow with all the content-sucking and scraping that the "AI" corporations let lose on the web, it is not unusual to haver a massive spike in bot-related visits even in the personal-website-space. And those scrapers are ruthless, they hammer the servers in high frequency and repeatetly, and are killing the web as we know it along the way.
(…) Many of these scrapers are so sophisticated that it is hard, or impossible, to detect them in action. They often ignore the websites’ programmatic pleas not to be scraped, and are known to hit the more fragile parts of a website repeatedly. opendemocracy.net
I created a directory with a random name in the top-level of my website.
I then added this directory in the robots.txt file with a disallow. This directory is not linked anywhere. Its name is so random and cryptical that it is highly unlikely that a "name guessing" bot will find it (like those exploit-searching idiot scripts that hammer on "wp-admin" or "typo3" URLs even on sites that don't use WordPress or TYPO3…). Inside the directory is a index script that
a) sends me an email,
b) logs the visit with user-agent-string and IP address and
c) saves the data in a nosql db.In front of my website I have a script that will check the current visitor's IP address against the nosql and if the IP matches, a HTTP 403 status is served.
Here's a best-of user agent strings that recently "visited" my hidden dir.
That last one is superb, considering that this one alone is several times in my log, of course with a different IP each time:PetalBot
Googlebot/2.1
Claude-SearchBot/1.0
Thinkbot/0 +In_the_test_phase,_if_the_Thinkbot_brings_you_trouble,_please_block_its_IP_address._Thank_you.Plus, there's a load more that pretend to be "normal" web browsers, of course. 🙄
It is a crude, a symbolical fist shaking yelling at clouds kind-of thing, especially compared to the things that Matthias Ott shared in his post, but it is better than nothing.
-
Army of Bots
For some months now I have a simple detection against "bad" bots in place. Bots that scrape *everything* they find and very likely are vacuuming all the contents they get to feed the data grinders that train the LLMs of the world. Bots that not only ignore the "robots.txt" protocol, but actively see entries in the robots.txt file as an invitation to visit the contents that are listed there as "disallowed".
I always had a hunch that stating addresses in a publicy reachable text file and flagging those as "please stay out of there" wasn't the best idea, but well, it was the only thing we've got back in the days where the only bots out there were the crawlers of the search engines.
(…) There are two important considerations when using /robots.txt:
robots can ignore your /robots.txt. Especially malware robots that scan the web for security vulnerabilities, and email address harvesters used by spammers will pay no attention.
the /robots.txt file is a publicly available file. Anyone can see what sections of your server you don't want robots to use.
robotstxt.orgNow with all the content-sucking and scraping that the "AI" corporations let lose on the web, it is not unusual to haver a massive spike in bot-related visits even in the personal-website-space. And those scrapers are ruthless, they hammer the servers in high frequency and repeatetly, and are killing the web as we know it along the way.
(…) Many of these scrapers are so sophisticated that it is hard, or impossible, to detect them in action. They often ignore the websites’ programmatic pleas not to be scraped, and are known to hit the more fragile parts of a website repeatedly. opendemocracy.net
I created a directory with a random name in the top-level of my website.
I then added this directory in the robots.txt file with a disallow. This directory is not linked anywhere. Its name is so random and cryptical that it is highly unlikely that a "name guessing" bot will find it (like those exploit-searching idiot scripts that hammer on "wp-admin" or "typo3" URLs even on sites that don't use WordPress or TYPO3…). Inside the directory is a index script that
a) sends me an email,
b) logs the visit with user-agent-string and IP address and
c) saves the data in a nosql db.In front of my website I have a script that will check the current visitor's IP address against the nosql and if the IP matches, a HTTP 403 status is served.
Here's a best-of user agent strings that recently "visited" my hidden dir.
That last one is superb, considering that this one alone is several times in my log, of course with a different IP each time:PetalBot
Googlebot/2.1
Claude-SearchBot/1.0
Thinkbot/0 +In_the_test_phase,_if_the_Thinkbot_brings_you_trouble,_please_block_its_IP_address._Thank_you.Plus, there's a load more that pretend to be "normal" web browsers, of course. 🙄
It is a crude, a symbolical fist shaking yelling at clouds kind-of thing, especially compared to the things that Matthias Ott shared in his post, but it is better than nothing.
-
https://OpenStreetMap.org has been disrupted today. We're working to keep the site online while facing extreme load from anonymous scrapers spread across 100,000+ IP addresses. Please be patient while we mitigate and protect the service. #OpenStreetMap #DDoS #Scrapers #AI
-
https://OpenStreetMap.org has been disrupted today. We're working to keep the site online while facing extreme load from anonymous scrapers spread across 100,000+ IP addresses. Please be patient while we mitigate and protect the service. #OpenStreetMap #DDoS #Scrapers #AI
-
Looks like those nasty AI scraper cannot follow 30x redirects
#webmaster #scrapers #website -
Any solution to get more SERP results from Google? Any hack/tricks? #BuildInPublic #scraping #scrapers #python
-
Any solution to get more SERP results from Google? Any hack/tricks? #BuildInPublic #scraping #scrapers #python
-
Posted some new blog-articles during the weekend .. Now I saw a quite substancial spike in traffic, that doesn't really look like normal human interaction ...
Grafana/Loki with some LogQL did quickly reveal it. It's scrapers of the AI-Slop generators hitting the webserver in bursts 🤦♂️
Time for come countermeasures 🙂
-
Posted some new blog-articles during the weekend .. Now I saw a quite substancial spike in traffic, that doesn't really look like normal human interaction ...
Grafana/Loki with some LogQL did quickly reveal it. It's scrapers of the AI-Slop generators hitting the webserver in bursts 🤦♂️
Time for come countermeasures 🙂
-
Really excited by becoming a collaborator on #stegodon - It's such an exciting piece of software! I did however spend most of my morning countering #scrapers on lemmy.zip - the fun of hosting a site on the #fediverse eh :) -
Really excited by becoming a collaborator on #stegodon - It's such an exciting piece of software! I did however spend most of my morning countering #scrapers on lemmy.zip - the fun of hosting a site on the #fediverse eh :) -
We can't have nice things because of AI scrapers
https://blog.metabrainz.org/2025/12/11/we-cant-have-nice-things-because-of-ai-scrapers/
#HackerNews #AI #Scrapers #Technology #Ethics #Online #Community #Digital #Rights
-
We can't have nice things because of AI scrapers
https://blog.metabrainz.org/2025/12/11/we-cant-have-nice-things-because-of-ai-scrapers/
#HackerNews #AI #Scrapers #Technology #Ethics #Online #Community #Digital #Rights
-
#fediauthors si je pars du postulat que les LLM scrappent le web et copient/aspirent le contenu que je crée sur mon blog. Comment faire pour m'en prémunir, continuer à partager mes idées et protéger mon contenu de ce vol et de l'utilisation de mes récits pour alimenter ces I.A. ?
Quelles solutions, quels outils ?
Dois-je simplement cesser de créer ? Les capsules #Gemini sont elles scrappées elles aussi ? Si non, combien de temps avant qu'elles ne le soient ?
Merci
#scrapers #LLM #voleDeDonnees #commentFaire #blog -
#fediauthors si je pars du postulat que les LLM scrappent le web et copient/aspirent le contenu que je crée sur mon blog. Comment faire pour m'en prémunir, continuer à partager mes idées et protéger mon contenu de ce vol et de l'utilisation de mes récits pour alimenter ces I.A. ?
Quelles solutions, quels outils ?
Dois-je simplement cesser de créer ? Les capsules #Gemini sont elles scrappées elles aussi ? Si non, combien de temps avant qu'elles ne le soient ?
Merci
#scrapers #LLM #voleDeDonnees #commentFaire #blog -
🤖🔒 A fox-led #crusade to bamboozle #AI #scrapers from a "Git forge" that sounds as mythical as it does unnecessary. 29 minutes of your life wasted on a convoluted game of hide-and-seek with bots, because apparently, that's the hill we're choosing to die on. 🦊💻
https://vulpinecitrus.info/blog/guarding-git-forge-ai-scrapers/ #Fox #GitForge #HideAndSeek #TechHumor #HackerNews #ngated -
🤖🔒 A fox-led #crusade to bamboozle #AI #scrapers from a "Git forge" that sounds as mythical as it does unnecessary. 29 minutes of your life wasted on a convoluted game of hide-and-seek with bots, because apparently, that's the hill we're choosing to die on. 🦊💻
https://vulpinecitrus.info/blog/guarding-git-forge-ai-scrapers/ #Fox #GitForge #HideAndSeek #TechHumor #HackerNews #ngated -
🤖🔒 A fox-led #crusade to bamboozle #AI #scrapers from a "Git forge" that sounds as mythical as it does unnecessary. 29 minutes of your life wasted on a convoluted game of hide-and-seek with bots, because apparently, that's the hill we're choosing to die on. 🦊💻
https://vulpinecitrus.info/blog/guarding-git-forge-ai-scrapers/ #Fox #GitForge #HideAndSeek #TechHumor #HackerNews #ngated -
🤖🔒 A fox-led #crusade to bamboozle #AI #scrapers from a "Git forge" that sounds as mythical as it does unnecessary. 29 minutes of your life wasted on a convoluted game of hide-and-seek with bots, because apparently, that's the hill we're choosing to die on. 🦊💻
https://vulpinecitrus.info/blog/guarding-git-forge-ai-scrapers/ #Fox #GitForge #HideAndSeek #TechHumor #HackerNews #ngated -
#Trump Media is partnering with #Perplexity to bring #AIsearch to #TruthSocial. Perplexity, which has a history of using #scrapers to evade websites and #plagiarising content. Despite Trump Media’s mission to end #BigTech’s influence, the partnership with Perplexity, whose investors include #JeffBezos, highlights the #complexities of Big Tech’s influence. https://www.404media.co/trump-is-launching-an-ai-search-engine-powered-by-perplexity/?eicker.news #tech #media #news
-
#Trump Media is partnering with #Perplexity to bring #AIsearch to #TruthSocial. Perplexity, which has a history of using #scrapers to evade websites and #plagiarising content. Despite Trump Media’s mission to end #BigTech’s influence, the partnership with Perplexity, whose investors include #JeffBezos, highlights the #complexities of Big Tech’s influence. https://www.404media.co/trump-is-launching-an-ai-search-engine-powered-by-perplexity/?eicker.news #tech #media #news