home.social

#crawlers — Public Fediverse posts

Live and recent posts from across the Fediverse tagged #crawlers, aggregated by home.social.

  1. I tried to extract what's not personal: got.thinkberg.com/?action=summ

    It may be useful as a pattern for web crawlers in general.

    #gotwebd #crawlers #guard

  2. I will test it and put it on my got server later. #crawlers #OpenBSD

  3. The #crawlers are pretty annoying, especially when looking at irrelevant stuff digging deeper than necessary. Fortunately, in #OpenBSD using the #fail2ban pattern can be applied as well. My #GoT web server got hammered and first I just banned all the found crawler names. However, now, every action they do on gotwebd is remembered with an effort number and if that adds up to 100, the IP is banned. Additonally, connection storms are also banned if they follow certain patterns. #pf, #perl and I am done using only on-board tools.

    Why? I didn't want to install #anubis. Not because I don't like it, it is just because I like to do the minimum.

  4. The #crawlers are pretty annoying, especially when looking at irrelevant stuff digging deeper than necessary. Fortunately, in #OpenBSD using the #fail2ban pattern can be applied as well. My #GoT web server got hammered and first I just banned all the found crawler names. However, now, every action they do on gotwebd is remembered with an effort number and if that adds up to 100, the IP is banned. Additonally, connection storms are also banned if they follow certain patterns. #pf, #perl and I am done using only on-board tools.

    Why? I didn't want to install #anubis. Not because I don't like it, it is just because I like to do the minimum.

  5. The #crawlers are pretty annoying, especially when looking at irrelevant stuff digging deeper than necessary. Fortunately, in #OpenBSD using the #fail2ban pattern can be applied as well. My #GoT web server got hammered and first I just banned all the found crawler names. However, now, every action they do on gotwebd is remembered with an effort number and if that adds up to 100, the IP is banned. Additonally, connection storms are also banned if they follow certain patterns. #pf, #perl and I am done using only on-board tools.

    Why? I didn't want to install #anubis. Not because I don't like it, it is just because I like to do the minimum.

  6. The #crawlers are pretty annoying, especially when looking at irrelevant stuff digging deeper than necessary. Fortunately, in #OpenBSD using the #fail2ban pattern can be applied as well. My #GoT web server got hammered and first I just banned all the found crawler names. However, now, every action they do on gotwebd is remembered with an effort number and if that adds up to 100, the IP is banned. Additonally, connection storms are also banned if they follow certain patterns. #pf, #perl and I am done using only on-board tools.

    Why? I didn't want to install #anubis. Not because I don't like it, it is just because I like to do the minimum.

  7. The #crawlers are pretty annoying, especially when looking at irrelevant stuff digging deeper than necessary. Fortunately, in #OpenBSD using the fail2ban pattern can be applied as well. My #GoT web server got hammered and first I just banned all the found crawler names. However, now, every action they do on gotwebd is remembered with an effort number and if that adds up to 100, the IP is banned. Additonally, connection storms are also banned if they follow certain patterns. #pf, #perl and I am done using only on-board tools.

    Why? I didn't want to install #anubis. Not because I don't like it, it is just because I like to do the minimum.

  8. Yesterday I started blocking all non-European IP addresses on our Gitlab instance because of the huge flood of AI scrapers. Sad that we have to do these kind of things

  9. #Anti #AI #Software / #Website Protection / AI #poison

    iocaine.madhouse-project.org/

    iocaine features:

    > No sympathy towards #crawlers

    Stand in the way of AI but not your visitors

    > Lightweight = Very little #CPU and #RAM

    Tries its best to require little resources from your visitors too.

    If you can serve #static content, you'll be able to run #iocaine too.

  10. Blocking AI crawlers cost news publishers 7% of traffic, study finds: A Wharton and Rutgers study finds news publishers who blocked LLM crawlers lost 7% of weekly traffic in 6 weeks, with no measurable content protection gains. ppc.land/blocking-ai-crawlers- #AI #crawlers #NewsPublishers #TrafficLoss #ContentProtection

  11. New Episodes

    Hulu released three new episodes of The Handmaid’s Tale. I’m a few minutes into episode two. The best description of the first few seasons that I heard was “misery porn.” That is accurate. Painfully, crushingly accurate. The current season is not expected to tow that line. Instead it’s looking like it might be “revenge porn.” Given the state of the real world… here’s hoping.

    The Handmaid’s Tale this morning, Daredevil tonight, Doctor Who on Saturday, and The Last of Us on Sunday. Hey television world, why not spread out the riches a little. Share that wealth with the rest of the calendar.

    Something is up with my Flickr account. Yesterday my hit count was about 10 times what I usually see. This morning before 7:00am it’s already more than yesterday. I think some bot somewhere is crawling me. It might be time for a new backup. Just in case. I hope the bots enjoy all of the cat pictures.

    Okay, I have to post this now. I can’t keep typing up brain droppings while hanging on every word of one of the best television shows Earth has ever produced. Seriously. A million years from now some alien species is going to find our remains and turn our planet into an archeological site and watch this show and think… those beings were seriously fucked up but I cannot stop watching.

    #bots #crawlers #Flickr #hulu #miseryPorn #revengePorn #tech #Television #theHandmaidSTale