#webarchiving — Public Fediverse posts
Live and recent posts from across the Fediverse tagged #webarchiving, aggregated by home.social.
-
The Case for Crowdfunding at Scale: How Donors are Fundraising With Us and Creating a Path to Lasting Preservation
-
RE: https://cloudisland.nz/@stuartyeates/116972405510951605
useful data for #webarchiving provenance
-
📰 asciimoo/omnom
A web content preservation service
Archives web pages as snapshots, aggregates feeds, and supports ActivityPub streams with a self-hosted multiuser interface
⭐ Stars: 651
📅 Last Update: Jul 19, 2026https://github.com/asciimoo/omnom
#selfhosted #homelab #selfhost #selfhosting #opensource #webarchiving #rss
-
Anyone got thoughts/recommendations on the feasibility of setting up 120k URL redirects (301 Moved Permanently) as part of an institutional repository migration?
-
🌐 Ray-D-Song/web-archive
Free web archiving and sharing service.
Archives webpages as single HTML files via a browser plugin, stores them on a server, and allows searching and sharing through a web client
⭐ Stars: 928
📅 Last Update: Jun 12, 2026https://github.com/Ray-D-Song/web-archive
#selfhosted #homelab #selfhost #selfhosting #opensource #webarchiving #cloudflare
-
Library of Congress: Remembering the Past, Preserving the Present: The America 250 Semiquincentennial Web Archive. “The America 250 Semiquincentennial Web Archive documents how Americans are commemorating and reflecting on the nation’s 250th anniversary. In this interview, Malea Walker discusses how the collection evolved from a project focused on government websites into a broader effort to […]
https://rbfirehose.com/2026/07/06/remembering-the-past-preserving-the-present-the-america-250-semiquincentennial-web-archive-library-of-congress/ -
Shortened links? Expand them and save the URLs
by @beet_keeperShortened links are a digital preservation and web archiving nightmare. You can imagine how they need to work:
Create a unique short code for a given (target) URL (like a hash, but far far shorter)
Pair the short code with your URL in a database.
Create a redirect rule on the URL shortening server from the new source URL to the target URL.
Send the shortened link to the caller, e.g. shortURL.com/123badf00d
In-perpetuity: continue to pay for your domain; maintain the database; look after redirect rules during server migrations; ensure duplicate short-codes are not created.
A URL-shortening business in five easy steps.
But what does that mean for digital preservation?
#Code #Coding #cURL #digipres #DigitalArchiving #DigitalPreservation #httpreserve #httpreserveLinkstat #JC #linkstat #Paradata #WebArchives #webArchiving -
We got educational Claude accounts at $work, so I succumbed and spent the better part of a day working on this #webarchiving thing I've had in the back of my mind for a while:
https://github.com/edsu/rustyweb
It is a Rust service that will index the contents of WACZs and make them available via the web like Wayback. Embedded search with tantivy, can index PDFs as well as HTML, and can index/replay remotely accessible WACZ files too.
-
Preserving the First Draft of History: Reflections from the National Summit on Local News Preservation
https://web.brid.gy/r/https://blog.archive.org/2026/06/29/newssummit/
-
Preserving the First Draft of History: Reflections from the National Summit on Local News Preservation
https://web.brid.gy/r/https://blog.archive.org/2026/06/29/newsummit/
-
🧵2/2
🕹️ Hosted by our friends at Vancouver Public Library, the kiosk transforms web archives into something you can browse, zoom through, and explore for yourself.
📚 It's also just steps away from Internet Archive Canada's live book-scanning station, where visitors can watch books being digitized in real time.
@internetarchivecanada #WebArchiving #DigitalPreservation #CanadianHistory
-
Why we do this work…
by @beet_keeperThere aren’t many rewards in a discipline that is about taking the long term view but occasionally something comes up that you can take some pride in.
Last month, Ed Summers put out a call on Mastodon: digipres.club where he was wrestling with a CD-R format that was difficult to recognize. The disks likely held precious data belonging to his late brother.
Much of the search area had already been examined and narrowed down by folks in the community, including Misty de Meo, Roxi Ruuska, Ethan Gates, and Johan van der Knijff who all contributed suggestions and analysis..
Ed was able to share a copy of one of his disk images, and I had some time that I could dedicate to taking a look as well.
Long-story short, we were able to identify the disks, and Ed has written up the background here: https://inkdroid.org/2026/06/12/tascam/
The situation might be familiar to others: a digital file that isn’t recognized by the major file format identification tools, and yet, because of its context, you know it is something that might be important.
I have different experiences with these types of files, sometimes they are valuable (and you want to look after them), sometimes they are not (and it can still benefit you to get rid of them). The process of finding this out often follows a similar path.
In this instance the files turned out to be incredibly valuable and I wanted to elaborate on the path of discovery. Even though it really isn’t very sophisticated, I hope it will be helpful to those with unidentified digital records who might find the task of identifying them quite daunting.
#DAW #digipres #DigitalForensics #digitalHeritage #DigitalPreservation #internetArchive #MusicProduction #personalDigitalArchiving #TASCAM #TEAC #WebArchives #webArchiving -
Digitale Belege sichern: Internet Archive nutzen und selbst anlegen – Mein Beitrag bei der Netzwerk Recherche 2026. Hier sind die Folien:
https://katharinabrunner.de/2026/06/digitale-belege-sichern-internet-archive-nutzen-und-selbst-anlegen-mein-beitrag-bei-der-netzwerk-recherche-2026/ -
Digitale Belege sichern: Internet Archive nutzen und selbst anlegen – Mein Beitrag bei der Netzwerk Recherche 2026
Wie sollen wir festhalten, was wir im Internet finden? Mit einem Screenshot? Ein Screenshot fast so schnell gefälscht wie er gemacht wurde. Entwickler-Tools öffnen, im HTML den Text ändern, fertig. Schon ist aus einem möglich auf der Startseite von Netzwerk Recherche ein unmöglich geworden.
echter Screenshot
gefälschter Screenshot
Das war der Einstieg meines Vortrags bei der Netzwerk […]
https://katharinabrunner.de/2026/06/digitale-belege-sichern-internet-archive-nutzen-und-selbst-anlegen-mein-beitrag-bei-der-netzwerk-recherche-2026/ #browsertrix #dataJournalism #DigitalArchive #InternetArchive #Journalismus #webarchiving #webrecorder -
seen on HN: https://kage.tamnd.com/
kage renders every page in headless Chrome, snapshots the final DOM, removes every script and event handler, and downloads and rewrites the CSS, images, and fonts.saves in ZIM Format, in the comments the author says it will support WARC too https://news.ycombinator.com/item?id=48529990
-
Addio Digilander: il 9 giugno si spegne un pezzo di storia digitale dei primi internauti italiani
La storica piattaforma Libero Community, incluso Digilander, si prepara alla chiusura definitiva. Gli utenti dovranno salvare i contenuti prima della disattivazione.
https://www.libero.it/tecnologia/addio-digilander-libero-community-chiude-salvare-contenuti-116560 #webarchiving -
I just wanted to share a (not so late) night rant with you.
In three days, the Italian web portal Libero.it is going to shut down thousands of early blogs that were originally created through the platform ItaliaOnline and later rebranded as Digilander.
I think this is a paradigmatic case of what we are going to experience more and more often in the near future. 1/#digitaloblivion #webarchiving #Digilander #Italianwebhistory
#lostinternet #earlyblogs -
Können wir das Alter von Webseiten abschätzen, wenn uns nur ein Crawl zur Verfügung steht?
Diese Frage hat Ira Kokoshko und Robert Jäschke beschäftigt – und sie haben dazu den Ancient GeoCities Datensatz auf der Web Science 2026 in Braunschweig vorgestellt. Wie gut ein LLM bei der Schätzung des Alters von Webseiten performt, könnt ihr im Paper nachlesen:
-
I was the first person to archive a webpage from Internet Archive Europe on the Internet Archive’s Wayback Machine.
LoL
#InternetArchive #InternetArchiveEurope #WaybackMachine #archive #archiving #WebArchiving #WebPreservation #inception
-
“People aren’t sure what’s true, and what libraries are here for is to help with that.”
Brewster Kahle, digital librarian of the Internet Archive, discusses the future of the #WaybackMachine in ABC Radio National (🇦🇺 Australia)’s “Wayback Machine: The internet’s archive in peril,” a look at how media companies are restricting the preservation of the web itself.
🎧 Listen ⤵️
https://www.abc.net.au/listen/programs/sundayextra/wayback-machine/106604988#InternetHistory #WebArchiving @abcaustraliarss @brewsterkahle
-
"Common Crawl mirrors its monthly crawl archive to the Hugging Face Hub as a Storage Bucket. Alongside the raw pages, it now publishes the columnar URL index — one parquet row per crawled page (host, language, MIME type, fetch status, and a pointer to the page's bytes). That makes the whole crawl queryable without touching the petabytes of underlying WARCs."
https://huggingface.co/spaces/davanstrien/common-crawl-april-2026
#webarchiving -
NiemanLab: More than 340 local news outlets are limiting the Internet Archive’s access to their journalism. “Our new analysis shows that more than 340 local news sites across the United States are now limiting the Internet Archive’s ability to access and preserve their stories. Many sites in our sample are owned by five of the seven largest local news publishers in the country: USA Today […]
https://rbfirehose.com/2026/05/21/niemanlab-more-than-340-local-news-outlets-are-limiting-the-internet-archives-access-to-their-journalism/