home.social

#archiveteam — Public Fediverse posts

Live and recent posts from across the Fediverse tagged #archiveteam, aggregated by home.social.

fetched live
  1. I donʼt know whoʼs running the #ArchiveTeam #ArchiveBot, but you just donʼt load someoneʼs server like that. Loads multiples greater than Googlebot and Baidu. Not the first dodgy scraper found on Github, and I canʼt trust that people arenʼt using them as theftbots.

  2. @jonny Nice. The torrents are seeding smoothly. sciop.net/tags/smithsonian

    However I'm not getting any downloads from the webseed, at least in Trasmission 4.0.6 (38c164933e). It's looking for an URL like smithsonian-open-access.s3.ama . Do they really have all files in a single directory?

    #digipres #ArchiveTeam

  3. @raffaele just a heads up: if you wanna help save what’s there you can use a #ArchiveTeam warrior warrior.archiveteam.org and set it to work on the goo.gl project

  4. @slashdot good, one less stupid link shortening service that breaks the web.

    Luckily those good folks at #ArchiveTeam have been trying to archive them as much as possible.

    You can help out save these short-sighted services by running your own VM that archives websites and uploads them to the @internetarchive

    Check out ArchiveTeam and get the software here: wiki.archiveteam.org/

  5. I still have a bunch of computers running ArchiveTeam Warrior and here are the totals for I've downloaded so far...

    💿 Telegram 107 GiB
    💿 Voice of America 94 GiB
    💿 US Government 44 GiB
    💿 Goo-gl 450 MiB
    💿 Twitch 135 MiB

    #ArchiveTeam #archive #data #web

  6. @eloquence
    Having posts or other indexed/indexable content refer to URL shorteners is dangerous for referrals/archiving/…:

    #ArchiveTeam, the people behind e.g. the effort to archive US government websites in a hurry—before they were deleted/changed in an even greater hurry by the current administration, write about #URLShorteners:

    "Such services are a ticking timebomb. If they go away, get hacked, or sell out, millions of links will be lost (see Wikipedia: Link Rot)."
    wiki.archiveteam.org/index.php

  7. The ArchiveTeam Warrior has been running intermittently on my laptop for ten days now.

    It downloads stuff and puts it into the Internet Archive.

    Everything's fine. It only runs while I use the laptop. I don't notice it. When the laptop goes into standby and wakes up again that doesn't seem to have any adverse effects.

    I've downloaded and uploaded gigabytes so far. The top of the leader board for this project is half a petabyte.

    Now I'm considering a installation where it could run around the clock. I don't want to increase our household's standby energy consumption too much, so I will see how that goes.

    @internetarchive

    #archiving
    #internetArchive
    #DataRescue
    #dataPreservation
    #digtitalPreservation
    #archiveTeamWarrior
    #archiveTeam

  8. Installed and started the ArchiveTeam Warrior. Very smooth experience.

    It downloads stuff and puts it into the Internet Archive.

    I took the "ArchiveTeam’s Choice" project and it chose public telegram channels. It's not taking a lot of bandwidth or memory or space or computing, as far as I can tell. It might take too much of my time and focus if I continue staring at the dashboard to try and figure out what all that stuff is.

    warrior.archiveteam.org/

    @internetarchive

    #archiving #internetArchive #DataRescue #dataPreservation #digtitalPreservation #archiveTeamWarrior #archiveTeam

  9. There was a lack of a decent web based leaderboard so I wrote one in python and got up to speed on publishing to pypi properly. It's a nice minimal example of publishing a single python file to an installable command.

    github.com/westonal/archive-wa

    #archiveteam #archiveteamwarrior

  10. Anyone else running an #ArchiveTeam instance on the usgov project? Mine was humming all week but not it's not getting any items. Is the archive... done?

  11. I managed to help archive ~35GB of US Government web content with my #ArchiveTeam Warrior instance. At the moment there are no more available to-do items, but I'm keeping my warrior alive if any items re-enter the queue. Glad I was able to donate some CPU and bandwidth to the cause :)

    #InternetArchive #USPol #Government

  12. I love that there are so many great efforts out there preserving Geocities, and the internet as a whole. So much wholesome, personal stuff, imortalised.

    To the #ArchiveTeam nerds, thank you for being awesome.

  13. I only just now took a good look at the Archive Team Warrior logo 😹 10/10 no notes

    #ArchiveTeam #ArchiveTeamWarrior

  14. Started helping #archiveteam backup all of the US government websites last night. Up to 7Gb today! Simple docker setup in #Unraid for me but If you have bandwidth and a computer doing nothing you can help too: wiki.archiveteam.org/index.php

  15. Really simple instructions on how to help archive the US Government's websites:

    https://wiki.archiveteam.org/index.php/ArchiveTeam_Warrior#Installing_and_running_with_Docker

    Although #ArchiveTeam are not affiliated with with the #InternetArchive, that is where the archived sites are stored.

    #Archive #Backups #USPol

  16. 🚨 Technologists, archivists, and internet historians—this is your call to action! 🚨

    As the authoritarian, censorship-happy Trump/Musk regime takes hold, federal government websites are getting wiped, altered, or disappeared entirely. Critical public records, scientific research, and historical data are at risk. We can’t let that happen.

    🔹 Take action now: Use #ArchiveTeam’s tool to help preserve government websites before they vanish.

    💾 How you can help:
    1️⃣ Set up the Archive Team’s docker-compose tool.
    2️⃣ Start archiving key sites you care about.
    3️⃣ Share this message with other technologists.

    The internet remembers only if we make it remember. Let’s keep public data public. 🛡️📜 #ArchivingMatters #InternetPreservation #SaveGovData

    github.com/goudarziha/archive-