#crawlers — Public Fediverse posts
Live and recent posts from across the Fediverse tagged #crawlers, aggregated by home.social.
-
I tried to extract what's not personal: https://got.thinkberg.com/?action=summary&path=gotwebd-guard.git
It may be useful as a pattern for web crawlers in general.
-
The #crawlers are pretty annoying, especially when looking at irrelevant stuff digging deeper than necessary. Fortunately, in #OpenBSD using the #fail2ban pattern can be applied as well. My #GoT web server got hammered and first I just banned all the found crawler names. However, now, every action they do on gotwebd is remembered with an effort number and if that adds up to 100, the IP is banned. Additonally, connection storms are also banned if they follow certain patterns. #pf, #perl and I am done using only on-board tools.
Why? I didn't want to install #anubis. Not because I don't like it, it is just because I like to do the minimum.
-
The #crawlers are pretty annoying, especially when looking at irrelevant stuff digging deeper than necessary. Fortunately, in #OpenBSD using the #fail2ban pattern can be applied as well. My #GoT web server got hammered and first I just banned all the found crawler names. However, now, every action they do on gotwebd is remembered with an effort number and if that adds up to 100, the IP is banned. Additonally, connection storms are also banned if they follow certain patterns. #pf, #perl and I am done using only on-board tools.
Why? I didn't want to install #anubis. Not because I don't like it, it is just because I like to do the minimum.
-
The #crawlers are pretty annoying, especially when looking at irrelevant stuff digging deeper than necessary. Fortunately, in #OpenBSD using the #fail2ban pattern can be applied as well. My #GoT web server got hammered and first I just banned all the found crawler names. However, now, every action they do on gotwebd is remembered with an effort number and if that adds up to 100, the IP is banned. Additonally, connection storms are also banned if they follow certain patterns. #pf, #perl and I am done using only on-board tools.
Why? I didn't want to install #anubis. Not because I don't like it, it is just because I like to do the minimum.
-
The #crawlers are pretty annoying, especially when looking at irrelevant stuff digging deeper than necessary. Fortunately, in #OpenBSD using the #fail2ban pattern can be applied as well. My #GoT web server got hammered and first I just banned all the found crawler names. However, now, every action they do on gotwebd is remembered with an effort number and if that adds up to 100, the IP is banned. Additonally, connection storms are also banned if they follow certain patterns. #pf, #perl and I am done using only on-board tools.
Why? I didn't want to install #anubis. Not because I don't like it, it is just because I like to do the minimum.
-
The #crawlers are pretty annoying, especially when looking at irrelevant stuff digging deeper than necessary. Fortunately, in #OpenBSD using the fail2ban pattern can be applied as well. My #GoT web server got hammered and first I just banned all the found crawler names. However, now, every action they do on gotwebd is remembered with an effort number and if that adds up to 100, the IP is banned. Additonally, connection storms are also banned if they follow certain patterns. #pf, #perl and I am done using only on-board tools.
Why? I didn't want to install #anubis. Not because I don't like it, it is just because I like to do the minimum.
-
ICYMI: Microsoft Clarity now flags robots.txt violations inside Bot Analytics: Microsoft Clarity now surfaces robots.txt violations in Bot Analytics, showing publishers which AI crawlers break access rules and what content they target. https://ppc.land/microsoft-clarity-now-flags-robots-txt-violations-inside-bot-analytics/ #MicrosoftClarity #BotAnalytics #SEO #WebAnalytics #Crawlers
-
Spider-Horror Pic ‘Crawlers’ Acquired By Roadside Attractions & Saban Films
#Acquisitions #News #Crawlers #RoadsideAttractions #SabanFilmshttps://deadline.com/2026/04/crawlers-roadside-attractions-1236877140/
-
@iagondiscord The problem is that it's not targeting Codeberg. It's the #AIgoldrush. The web was completely crawled, just not by everyone yet. So startups start their #crawlers, carelessly and explicitly ignoring robots.txt to get the #biggestdata.
It does not matter if the web can no longer serve humanity due to this. Training the #AI is the only thing that matters.
Maybe a bit like a sacrifice for faith.
~f #goldrush