#robotsdottxt — Public Fediverse posts
Live and recent posts from across the Fediverse tagged #robotsdottxt, aggregated by home.social.
-
@lfa isn't that just another CC signal? Similar to a robots.txt that blocks all crawlers?
Of course our problem includes AI scrapers that ignore #RobotsDotTxt ...
-
@lfa isn't that just another CC signal? Similar to a robots.txt that blocks all crawlers?
Of course our problem includes AI scrapers that ignore #RobotsDotTxt ...
-
@compsci_discussions So what DOES one actually put in the robots.txt file? #webcrawler #robotsdottxt
-
@compsci_discussions So what DOES one actually put in the robots.txt file? #webcrawler #robotsdottxt
-
The Web has been voluntarily opt-out since the early to mid- 1990s.
Before that, there was no opt-ing out.
The act of putting something onto the (open) Web was (and still is) understood as consent for others to use your data.
This is the norm of the Web — and has been for decades.
I think it is unlikely this norm is going to change any time soon.
#Fediverse #OptOut #RobotsDotTXT #RobotsNotWantedDotTXT #WorldWideWeb
-
The Web has been voluntarily opt-out since the early to mid- 1990s.
Before that, there was no opt-ing out.
The act of putting something onto the (open) Web was (and still is) understood as consent for others to use your data.
This is the norm of the Web — and has been for decades.
I think it is unlikely this norm is going to change any time soon.
#Fediverse #OptOut #RobotsDotTXT #RobotsNotWantedDotTXT #WorldWideWeb
-
Question to the #FediFetcher and #MastoAdmin communities:
Should FediFetcher follow the robots.txt?
(Edit: I should probably rephrase this slightly: I'm mostly interested in whether you think FediFetcher should honour blanket bans through robots.txt.)
Arguments for are quite obvious imo. But there are a few arguments against as well: For starters I wouldn't consider FediFetcher a bot, as it doesn't really go out and crawl things, nor does it store or process anything.
What do you think?
I would really appreciate if you could leave a comment explaining your decision as well (either here, or on the github issue).
https://github.com/nanos/FediFetcher/issues/84
#SingleUserInstance #FediFetcher #MastoAdmin #robots #robotsdottxt #RobotsTxt
-
Question to the #FediFetcher and #MastoAdmin communities:
Should FediFetcher follow the robots.txt?
(Edit: I should probably rephrase this slightly: I'm mostly interested in whether you think FediFetcher should honour blanket bans through robots.txt.)
Arguments for are quite obvious imo. But there are a few arguments against as well: For starters I wouldn't consider FediFetcher a bot, as it doesn't really go out and crawl things, nor does it store or process anything.
What do you think?
I would really appreciate if you could leave a comment explaining your decision as well (either here, or on the github issue).
https://github.com/nanos/FediFetcher/issues/84
#SingleUserInstance #FediFetcher #MastoAdmin #robots #robotsdottxt #RobotsTxt
-
If you run a website and you think that adding anything to robots.txt "blocks" #ChatGPT or any other #WebCrawler from accessing your site, you may want to think that process through a bit more.
#RobotsDotTxt is *asking nicely* for the robot not to look places. If you don't trust the ethics of the scanning party, I'm not sure why you would think that asking them nicely not to mine your free content for their profit would even slow them down.
-
If you run a website and you think that adding anything to robots.txt "blocks" #ChatGPT or any other #WebCrawler from accessing your site, you may want to think that process through a bit more.
#RobotsDotTxt is *asking nicely* for the robot not to look places. If you don't trust the ethics of the scanning party, I'm not sure why you would think that asking them nicely not to mine your free content for their profit would even slow them down.