#scrapers — Public Fediverse posts
Live and recent posts from across the Fediverse tagged #scrapers, aggregated by home.social.
-
75 000+ IP au tapis ! ⛔️🛡️
Et encore, ce chiffre ne montre que ce qui réussit à passer la première ligne de défense.En amont, mon fichier de blocage #Nginx personnalisé fait déjà un gros tri sélectif :
🛑 #Scrapers #IA, faux Chrome et outils automatisés reçoivent un #403 immédiat.
🟢 Les instances du #Fediverse ( #Mastodon, #PeerTube... ), elles, entrent sans problème.Nginx filtre les indésirables, #Fail2Ban verrouille le reste. Résultat : un serveur qui respire et des ressources préservées ! 😎🧹
-
75 000+ IP au tapis ! ⛔️🛡️
Et encore, ce chiffre ne montre que ce qui réussit à passer la première ligne de défense.En amont, mon fichier de blocage #Nginx personnalisé fait déjà un gros tri sélectif :
🛑 #Scrapers #IA, faux Chrome et outils automatisés reçoivent un #403 immédiat.
🟢 Les instances du #Fediverse ( #Mastodon, #PeerTube... ), elles, entrent sans problème.Nginx filtre les indésirables, #Fail2Ban verrouille le reste. Résultat : un serveur qui respire et des ressources préservées ! 😎🧹
-
75 000+ IP au tapis ! ⛔️🛡️
Et encore, ce chiffre ne montre que ce qui réussit à passer la première ligne de défense.En amont, mon fichier de blocage #Nginx personnalisé fait déjà un gros tri sélectif :
🛑 #Scrapers #IA, faux Chrome et outils automatisés reçoivent un #403 immédiat.
🟢 Les instances du #Fediverse ( #Mastodon, #PeerTube... ), elles, entrent sans problème.Nginx filtre les indésirables, #Fail2Ban verrouille le reste. Résultat : un serveur qui respire et des ressources préservées ! 😎🧹
-
75 000+ IP au tapis ! ⛔️🛡️
Et encore, ce chiffre ne montre que ce qui réussit à passer la première ligne de défense.En amont, mon fichier de blocage #Nginx personnalisé fait déjà un gros tri sélectif :
🛑 #Scrapers #IA, faux Chrome et outils automatisés reçoivent un #403 immédiat.
🟢 Les instances du #Fediverse ( #Mastodon, #PeerTube... ), elles, entrent sans problème.Nginx filtre les indésirables, #Fail2Ban verrouille le reste. Résultat : un serveur qui respire et des ressources préservées ! 😎🧹
-
75 000+ IP au tapis ! ⛔️🛡️
Et encore, ce chiffre ne montre que ce qui réussit à passer la première ligne de défense.En amont, mon fichier de blocage #Nginx personnalisé fait déjà un gros tri sélectif :
🛑 #Scrapers #IA, faux Chrome et outils automatisés reçoivent un #403 immédiat.
🟢 Les instances du #Fediverse ( #Mastodon, #PeerTube... ), elles, entrent sans problème.Nginx filtre les indésirables, #Fail2Ban verrouille le reste. Résultat : un serveur qui respire et des ressources préservées ! 😎🧹
-
Oh brother—it appears that there is a lot more data scraping going on than we thought:
-
Oh brother—it appears that there is a lot more data scraping going on than we thought:
-
Oh brother—it appears that there is a lot more data scraping going on than we thought:
-
#AI #scrapers are the PFAS of what was once the fertile ground of an open and honest internet.
In the race of AI to.. (yeah to what?), websites with decades of knowledge are forced to spend enormous amounts of compute and bandwidth defending themselves against an opaque ecosystem of scrapers.
An article about Residential Proxy Networks, “ethically sourced” IP addresses, Iocaine, NetNut-infected apps, Anubis, and the invisible infrastructure s(cr/h)aping the web.
-
#AI #scrapers are the PFAS of what was once the fertile ground of an open and honest internet.
In the race of AI to.. (yeah to what?), websites with decades of knowledge are forced to spend enormous amounts of compute and bandwidth defending themselves against an opaque ecosystem of scrapers.
An article about Residential Proxy Networks, “ethically sourced” IP addresses, Iocaine, NetNut-infected apps, Anubis, and the invisible infrastructure s(cr/h)aping the web.
-
#AI #scrapers are the PFAS of what was once the fertile ground of an open and honest internet.
In the race of AI to.. (yeah to what?), websites with decades of knowledge are forced to spend enormous amounts of compute and bandwidth defending themselves against an opaque ecosystem of scrapers.
An article about Residential Proxy Networks, “ethically sourced” IP addresses, Iocaine, NetNut-infected apps, Anubis, and the invisible infrastructure s(cr/h)aping the web.
-
#AI #scrapers are the PFAS of what was once the fertile ground of an open and honest internet.
In the race of AI to.. (yeah to what?), websites with decades of knowledge are forced to spend enormous amounts of compute and bandwidth defending themselves against an opaque ecosystem of scrapers.
An article about Residential Proxy Networks, “ethically sourced” IP addresses, Iocaine, NetNut-infected apps, Anubis, and the invisible infrastructure s(cr/h)aping the web.
-
#AI #scrapers are the PFAS of what was once the fertile ground of an open and honest internet.
In the race of AI to.. (yeah to what?), websites with decades of knowledge are forced to spend enormous amounts of compute and bandwidth defending themselves against an opaque ecosystem of scrapers.
An article about Residential Proxy Networks, “ethically sourced” IP addresses, Iocaine, NetNut-infected apps, Anubis, and the invisible infrastructure s(cr/h)aping the web.
-
@weeklyOSM Please don’t be alarmed if you’ve seen your browser being tested – even if only briefl
Our website was under enormous strain from #AI #scrapers
To defend against these scrapers, we’ve installed Anubis
https://en.wikipedia.org/wiki/Anubis_(software) -
@weeklyOSM Please don’t be alarmed if you’ve seen your browser being tested – even if only briefl
Our website was under enormous strain from #AI #scrapers
To defend against these scrapers, we’ve installed Anubis
https://en.wikipedia.org/wiki/Anubis_(software) -
@weeklyOSM Please don’t be alarmed if you’ve seen your browser being tested – even if only briefl
Our website was under enormous strain from #AI #scrapers
To defend against these scrapers, we’ve installed Anubis
https://en.wikipedia.org/wiki/Anubis_(software) -
@weeklyOSM Please don’t be alarmed if you’ve seen your browser being tested – even if only briefl
Our website was under enormous strain from #AI #scrapers
To defend against these scrapers, we’ve installed Anubis
https://en.wikipedia.org/wiki/Anubis_(software) -
@weeklyOSM Please don’t be alarmed if you’ve seen your browser being tested – even if only briefl
Our website was under enormous strain from #AI #scrapers
To defend against these scrapers, we’ve installed Anubis
https://en.wikipedia.org/wiki/Anubis_(software) -
CW: Adult content
Reminder & FYI to ppl new to Bluesky & my page👇 reposting “leaked” (stolen & posted without personal) porn is really harmful to sw. It means you support that practice. I can’t and won’t allow that in my space. You shouldn’t either. #scrapers #sw #sexworkers #porn #swsolidarity #spreadtheword
-
CW: Adult content
Reminder & FYI to ppl new to Bluesky & my page👇 reposting “leaked” (stolen & posted without personal) porn is really harmful to sw. It means you support that practice. I can’t and won’t allow that in my space. You shouldn’t either. #scrapers #sw #sexworkers #porn #swsolidarity #spreadtheword
-
CW: Adult content
Reminder & FYI to ppl new to Bluesky & my page👇 reposting “leaked” (stolen & posted without personal) porn is really harmful to sw. It means you support that practice. I can’t and won’t allow that in my space. You shouldn’t either. #scrapers #sw #sexworkers #porn #swsolidarity #spreadtheword
-
I appreciate the wise minds that have jumped into my thread over the past 24 hours.
It's amazing to see the variety of tools and attitudes towards the problem. The more I talk it through and hear from other folks, the more I feel that shutting the door is ultimately a self-own. If I want humans to see my website, I need to focus on making that as simple and unfettered as possible.
I have lots of things I could be doing better for my fellow humans. The trackers from Google and FB alone are my major concern for me. There is also a trade off between being searchable at all, versus blocking the scrapers and removing the trackers.
Any tech "solution" is ultimately imperfect. I wish this was a simpler equation. It does not feel like a "nuance" situation, rather a "you get screwed either way" situation.
#Scrapers #AIbots #WebHosting -
I appreciate the wise minds that have jumped into my thread over the past 24 hours.
It's amazing to see the variety of tools and attitudes towards the problem. The more I talk it through and hear from other folks, the more I feel that shutting the door is ultimately a self-own. If I want humans to see my website, I need to focus on making that as simple and unfettered as possible.
I have lots of things I could be doing better for my fellow humans. The trackers from Google and FB alone are my major concern for me. There is also a trade off between being searchable at all, versus blocking the scrapers and removing the trackers.
Any tech "solution" is ultimately imperfect. I wish this was a simpler equation. It does not feel like a "nuance" situation, rather a "you get screwed either way" situation.
#Scrapers #AIbots #WebHosting -
I appreciate the wise minds that have jumped into my thread over the past 24 hours.
It's amazing to see the variety of tools and attitudes towards the problem. The more I talk it through and hear from other folks, the more I feel that shutting the door is ultimately a self-own. If I want humans to see my website, I need to focus on making that as simple and unfettered as possible.
I have lots of things I could be doing better for my fellow humans. The trackers from Google and FB alone are my major concern for me. There is also a trade off between being searchable at all, versus blocking the scrapers and removing the trackers.
Any tech "solution" is ultimately imperfect. I wish this was a simpler equation. It does not feel like a "nuance" situation, rather a "you get screwed either way" situation.
#Scrapers #AIbots #WebHosting -
I appreciate the wise minds that have jumped into my thread over the past 24 hours.
It's amazing to see the variety of tools and attitudes towards the problem. The more I talk it through and hear from other folks, the more I feel that shutting the door is ultimately a self-own. If I want humans to see my website, I need to focus on making that as simple and unfettered as possible.
I have lots of things I could be doing better for my fellow humans. The trackers from Google and FB alone are my major concern for me. There is also a trade off between being searchable at all, versus blocking the scrapers and removing the trackers.
Any tech "solution" is ultimately imperfect. I wish this was a simpler equation. It does not feel like a "nuance" situation, rather a "you get screwed either way" situation.
#Scrapers #AIbots #WebHosting -
I appreciate the wise minds that have jumped into my thread over the past 24 hours.
It's amazing to see the variety of tools and attitudes towards the problem. The more I talk it through and hear from other folks, the more I feel that shutting the door is ultimately a self-own. If I want humans to see my website, I need to focus on making that as simple and unfettered as possible.
I have lots of things I could be doing better for my fellow humans. The trackers from Google and FB alone are my major concern for me. There is also a trade off between being searchable at all, versus blocking the scrapers and removing the trackers.
Any tech "solution" is ultimately imperfect. I wish this was a simpler equation. It does not feel like a "nuance" situation, rather a "you get screwed either way" situation.
#Scrapers #AIbots #WebHosting -
Strava declara la lucha a los scrapers ayer de su salida a bolsa #antes #BOLSA #declara #guerra #IPO #los #raspado_de_datos #salida #scrapers #Strava #ButterWord #Spanish_News Comenta tu opinión 👇
https://butterword.com/strava-declara-la-lucha-a-los-scrapers-ayer-de-su-salida-a-bolsa/?feed_id=83426&_unique_id=6a1d89e384e07 -
Why is #twitter not properly identifying itself as a bot when trying to scrape my website? (69.12.56.0/21 is AS63179 is Twitter)
Could it be cause they're a malicious party training an #aibot?
(This is extremely low-intensity, but based on the combination of this specific UA and the pages they're trying to reach, I've seen them before, coming in from residential proxies.)
The funny thing is that bots identifying as bots and observing robots.txt would actually be allowed to reach those particular pages.
-
Why is #twitter not properly identifying itself as a bot when trying to scrape my website? (69.12.56.0/21 is AS63179 is Twitter)
Could it be cause they're a malicious party training an #aibot?
(This is extremely low-intensity, but based on the combination of this specific UA and the pages they're trying to reach, I've seen them before, coming in from residential proxies.)
The funny thing is that bots identifying as bots and observing robots.txt would actually be allowed to reach those particular pages.
-
Why is #twitter not properly identifying itself as a bot when trying to scrape my website? (69.12.56.0/21 is AS63179 is Twitter)
Could it be cause they're a malicious party training an #aibot?
(This is extremely low-intensity, but based on the combination of this specific UA and the pages they're trying to reach, I've seen them before, coming in from residential proxies.)
The funny thing is that bots identifying as bots and observing robots.txt would actually be allowed to reach those particular pages.
-
Why is #twitter not properly identifying itself as a bot when trying to scrape my website? (69.12.56.0/21 is AS63179 is Twitter)
Could it be cause they're a malicious party training an #aibot?
(This is extremely low-intensity, but based on the combination of this specific UA and the pages they're trying to reach, I've seen them before, coming in from residential proxies.)
The funny thing is that bots identifying as bots and observing robots.txt would actually be allowed to reach those particular pages.
-
Why is #twitter not properly identifying itself as a bot when trying to scrape my website? (69.12.56.0/21 is AS63179 is Twitter)
Could it be cause they're a malicious party training an #aibot?
(This is extremely low-intensity, but based on the combination of this specific UA and the pages they're trying to reach, I've seen them before, coming in from residential proxies.)
The funny thing is that bots identifying as bots and observing robots.txt would actually be allowed to reach those particular pages.
-
Iocaine and my custom solution aren't good enough. :blobcatbigsob: I'm considering to add to login to my website rewrite as protection against bots.
I would always offer an anonymous session after completing a proof of work (which is also available without JS).
Do you think this is okay? Please don't hesitate to reply!
#website #personalBlog #PersonalSites #indieweb #spam #spamprotection #scrapers #selfhosting #iocaine
-
Iocaine and my custom solution aren't good enough. :blobcatbigsob: I'm considering to add to login to my website rewrite as protection against bots.
I would always offer an anonymous session after completing a proof of work (which is also available without JS).
Do you think this is okay? Please don't hesitate to reply!
#website #personalBlog #PersonalSites #indieweb #spam #spamprotection #scrapers #selfhosting #iocaine
-
Iocaine and my custom solution aren't good enough. :blobcatbigsob: I'm considering to add to login to my website rewrite as protection against bots.
I would always offer an anonymous session after completing a proof of work (which is also available without JS).
Do you think this is okay? Please don't hesitate to reply!
#website #personalBlog #PersonalSites #indieweb #spam #spamprotection #scrapers #selfhosting #iocaine
-
Iocaine and my custom solution aren't good enough. :blobcatbigsob: I'm considering to add to login to my website rewrite as protection against bots.
I would always offer an anonymous session after completing a proof of work (which is also available without JS).
Do you think this is okay? Please don't hesitate to reply!
#website #personalBlog #PersonalSites #indieweb #spam #spamprotection #scrapers #selfhosting #iocaine
-
Iocaine and my custom solution aren't good enough. :blobcatbigsob: I'm considering to add to login to my website rewrite as protection against bots.
I would always offer an anonymous session after completing a proof of work (which is also available without JS).
Do you think this is okay? Please don't hesitate to reply!
#website #personalBlog #PersonalSites #indieweb #spam #spamprotection #scrapers #selfhosting #iocaine
-
Desde afuera todavia se nota cierta latencia, a veces, posiblemente porque no han cesado los ataques de scraping. En la red interna vuela, y en las metricas los servidores no estan bajo carga o demanda altos, estan normales. El problema en ese caso sería que todos esos ataques que el firewall esta bloqueando exitosamente, lo hace recien dentro de la red, por lo que ese trafico ocupa lugar en la conexión dejando menos ancho de banda neto para el tráfico legítimo... veremos si la cosa mejora en los próximos dias #undernet #ataque #bots #scrapers #iabot #peertube
-
Desde afuera todavia se nota cierta latencia, a veces, posiblemente porque no han cesado los ataques de scraping. En la red interna vuela, y en las metricas los servidores no estan bajo carga o demanda altos, estan normales. El problema en ese caso sería que todos esos ataques que el firewall esta bloqueando exitosamente, lo hace recien dentro de la red, por lo que ese trafico ocupa lugar en la conexión dejando menos ancho de banda neto para el tráfico legítimo... veremos si la cosa mejora en los próximos dias #undernet #ataque #bots #scrapers #iabot #peertube
-
Desde afuera todavia se nota cierta latencia, a veces, posiblemente porque no han cesado los ataques de scraping. En la red interna vuela, y en las metricas los servidores no estan bajo carga o demanda altos, estan normales. El problema en ese caso sería que todos esos ataques que el firewall esta bloqueando exitosamente, lo hace recien dentro de la red, por lo que ese trafico ocupa lugar en la conexión dejando menos ancho de banda neto para el tráfico legítimo... veremos si la cosa mejora en los próximos dias #undernet #ataque #bots #scrapers #iabot #peertube
-
Desde afuera todavia se nota cierta latencia, a veces, posiblemente porque no han cesado los ataques de scraping. En la red interna vuela, y en las metricas los servidores no estan bajo carga o demanda altos, estan normales. El problema en ese caso sería que todos esos ataques que el firewall esta bloqueando exitosamente, lo hace recien dentro de la red, por lo que ese trafico ocupa lugar en la conexión dejando menos ancho de banda neto para el tráfico legítimo... veremos si la cosa mejora en los próximos dias #undernet #ataque #bots #scrapers #iabot #peertube