#httrack — Public Fediverse posts
Live and recent posts from across the Fediverse tagged #httrack, aggregated by home.social.
-
the biggest things i need ai to do for me is to have a high initial elo ranking but also be trainable to scan all local docs and then also bring in a lots of real time data and open datasets 24/7, display results on series of dashboards #rag #pydantic #yacy #httrack #cached version #best stacks #free for commercial use #competitive intel #tailored data
-
the biggest things i need ai to do for me is to have a high initial elo ranking but also be trainable to scan all local docs and then also bring in a lots of real time data and open datasets 24/7, display results on series of dashboards #rag #pydantic #yacy #httrack #cached version #best stacks #free for commercial use #competitive intel #tailored data
-
🌐🤦♂️ "Look Ma, I copied the entire internet! #HTTrack, the digital hoarder's dream, lets you download the web so you can finally browse those cat memes offline. Because nothing screams cutting-edge technology like reading 2005 forum threads in 2023." 📂😂
https://www.httrack.com/ #OfflineBrowsing #DigitalHoarding #InternetArchive #CatMemes #Nostalgia #HackerNews #ngated -
🌐🤦♂️ "Look Ma, I copied the entire internet! #HTTrack, the digital hoarder's dream, lets you download the web so you can finally browse those cat memes offline. Because nothing screams cutting-edge technology like reading 2005 forum threads in 2023." 📂😂
https://www.httrack.com/ #OfflineBrowsing #DigitalHoarding #InternetArchive #CatMemes #Nostalgia #HackerNews #ngated -
HTTrack Website Copier
#HackerNews #HTTrack #Website #Copier #website #cloning #webdevelopment #open-source #tools #technews
-
HTTrack Website Copier
#HackerNews #HTTrack #Website #Copier #website #cloning #webdevelopment #open-source #tools #technews
-
HTTrack - Der Website Downloader
In diesem Tutorial zeige ich dir, wie du ganze Websites mit HTTrack für den Offline-Zugriff speichern kannst. Egal, ob für die eigene Sicherung oder einfach zum Stöbern ohne Internet – ich zeige dir Schritt für Schritt, wie es funktioniert.
-
HTTrack - Der Website Downloader
In diesem Tutorial zeige ich dir, wie du ganze Websites mit HTTrack für den Offline-Zugriff speichern kannst. Egal, ob für die eigene Sicherung oder einfach zum Stöbern ohne Internet – ich zeige dir Schritt für Schritt, wie es funktioniert.
-
I am looking for archive.org as a self hosted service.
I want to have an automated static copy of a website, which preserves old copied versions.
It should provide a #crawler and a web interface to access the archived versions of the website.
The use case is a lousy CMS which often destroys content. I want to be able to restore content from the archive and to have a static website copy in the worst case.
#SelfHosting #WebsiteArchive #Archive #OffsiteBackp #Backup #HTTrack #WebsiteCopy
-
I am looking for archive.org as a self hosted service.
I want to have an automated static copy of a website, which preserves old copied versions.
It should provide a #crawler and a web interface to access the archived versions of the website.
The use case is a lousy CMS which often destroys content. I want to be able to restore content from the archive and to have a static website copy in the worst case.
#SelfHosting #WebsiteArchive #Archive #OffsiteBackp #Backup #HTTrack #WebsiteCopy
-
Actually, lemme think out loud about what I need #HTTrack to do, before I forget. It needs to pull jpg, png and svg images, javascript and any external CSS*) from any level within the Comicfury.com domain, but external links need to be skipped for mirroring.
*)AFAIK all CSS within ComicFury is inline! A baffling decision but one that will make my life easier with this. But I may be mistaken.
-
Actually, lemme think out loud about what I need #HTTrack to do, before I forget. It needs to pull jpg, png and svg images, javascript and any external CSS*) from any level within the Comicfury.com domain, but external links need to be skipped for mirroring.
*)AFAIK all CSS within ComicFury is inline! A baffling decision but one that will make my life easier with this. But I may be mistaken.
-
search engine on a stick would be a fun project 1tb nvme enc persistent bootable and you can spider your own sites in addition to top 10k sites already crawled and indexed - yacy could stand to be much more automated - it is a bit of work to get it set - not the config just all the sites loaded #httrack
-
search engine on a stick would be a fun project 1tb nvme enc persistent bootable and you can spider your own sites in addition to top 10k sites already crawled and indexed - yacy could stand to be much more automated - it is a bit of work to get it set - not the config just all the sites loaded #httrack
-
#HTTrack seems to be unmaintained (last release in 2017).
Any maintained recent opensource mirroring solution than can offload auth to a browser (for example, like #destreamer can)?
-
#HTTrack seems to be unmaintained (last release in 2017).
Any maintained recent opensource mirroring solution than can offload auth to a browser (for example, like #destreamer can)?
-
Manchmal will man ja auch eine ganze Webpräsenz sichern. #Httrack ist dafür auch ein gutes Tool, aber die Voreinstellungen müssen angepasst werden. #OSINT https://bashinho.de/2024/01/18/webseiten-mit-httrack-herunterladen/
-
Manchmal will man ja auch eine ganze Webpräsenz sichern. #Httrack ist dafür auch ein gutes Tool, aber die Voreinstellungen müssen angepasst werden. #OSINT https://bashinho.de/2024/01/18/webseiten-mit-httrack-herunterladen/
-
В очередной раз убеждаюсь, что #wget великая вещь!
Одна мелкая бура сообщила о своём закрытии, и я решил её сохранить себе.
Попробовал сначала #HTTrack, он пыхтел полдня и сохранил только html файлы.
wget сначала отказывался зеркалить сайт, но я добавил-Uи всё заработало. Примерно за 2 два часа он скачал весь сайт и все картинки.
Теперь я обладаю ~1800 картинками среднего качества и не знаю что с этим делать. :blobcatshrug: -
I would cancel one web server from an old-company, but wanted to keep the site somewhere (a php one).
Using httrack and gitlab pages I could do it quite easily! The site looks exactly the same, now static, and no cost to keep it running (only need to pay the domain).Some days I like technology, mainly the free/libre ones :)
-
Puras broncas al tratar de hacer un #WebScraping de un Google site, ni con el famoso #httrack
-
-
#httrack is running since 17 & 18 hours, #mediawiki really produce a hell lot of pages!
#website #carboncopy #websiterendering #html #static -
I'm going slightly #mad about #httrack and #wget to load all *.opus files from @chaosradio.
So it will be the good old fashioned +click-wait-click-wait-click-load+ way...
But maybe I'm just holding it wrong.
🙄 -
Today I learned that #httrack requires the input address as a URL. So if you have a local file with links, you need to specify the file protocol. And apparently you need a complete absolute path, too.
So for me that became:
$ httrack file://mnt/c/Users/aveeltstra/AppData/Local/Temp/mirror/source.html -O ./ "+*.JPG" -Y
And yes: that is accessing an MSWindows path from Debian on the Microsoft Linux Subsystem.