home.social

#robots_txt — Public Fediverse posts

Live and recent posts from across the Fediverse tagged #robots_txt, aggregated by home.social.

fetched live
  1. ICYMI: Cloudflare drops AI opt-outs from 118 sites' robots.txt, HasData finds: Of 592 sites declaring a GPTBot ban, 234 still served it a live page, while publishers turn away AI bots at 56.4%. Has enforcement moved from files to CDNs? ppc.land/cloudflare-drops-ai-o #Cloudflare #AI #digitalmarketing #robots_txt #GPTBot

  2. Oh wow, #OpenAI is #scraping #CT #logs like a kid in a candy store 🍬. Apparently, they're on a mission to hunt down... robots.txt files? 🤖🗂️ Because who doesn't love a treasure trove of 404 errors and TLS certificates? 💾🔍
    benjojo.co.uk/u/benjojo/h/Gxy2 #robots_txt #404_errors #TLS_certificates #tech_news #HackerNews #ngated

  3. Thinking about your robots.txt file? It might seem counterintuitive, but disallowing RSS feeds and certain pagination paths can be a smart SEO move.

    This technique helps search engines focus crawl budgets on your most important pages to avoid potential duplicate content.

    This post on WebHeads United looks at the technical reasons behind this strategy and whether it's right for your site.

    Read the SEO deep dive: webheadsunited.com/why-disallo

    #SEO #TechnicalSEO #CrawlBudget #WebDev #robots_txt

  4. I’ve made a little something, so I thought I'd share.

    Gort is a robots.txt parser and evaluator. It implements RFC 9309.

    More details in the ReadMe: github.com/pointlessone/gort

    #Ruby #rubygem #release #robotstxt #robots_txt #rfc9309

  5. My local government just launched a site redesign, changing CMSes and permalink structures.

    They didn't set up redirects for old URLs.

    Half the site is still blocked in robots.txt.

    I'm professionally flabbergasted.

    #webdev #redirects #robots_txt

Share on Mastodon

Enter the server where you have an account.