home.social

#httrack — Public Fediverse posts

Live and recent posts from across the Fediverse tagged #httrack, aggregated by home.social.

fetched live
  1. the biggest things i need ai to do for me is to have a high initial elo ranking but also be trainable to scan all local docs and then also bring in a lots of real time data and open datasets 24/7, display results on series of dashboards #rag #pydantic #yacy #httrack #cached version #best stacks #free for commercial use #competitive intel #tailored data

  2. the biggest things i need ai to do for me is to have a high initial elo ranking but also be trainable to scan all local docs and then also bring in a lots of real time data and open datasets 24/7, display results on series of dashboards #rag #pydantic #yacy #httrack #cached version #best stacks #free for commercial use #competitive intel #tailored data

  3. 🌐🤦‍♂️ "Look Ma, I copied the entire internet! #HTTrack, the digital hoarder's dream, lets you download the web so you can finally browse those cat memes offline. Because nothing screams cutting-edge technology like reading 2005 forum threads in 2023." 📂😂
    httrack.com/ #OfflineBrowsing #DigitalHoarding #InternetArchive #CatMemes #Nostalgia #HackerNews #ngated

  4. 🌐🤦‍♂️ "Look Ma, I copied the entire internet! #HTTrack, the digital hoarder's dream, lets you download the web so you can finally browse those cat memes offline. Because nothing screams cutting-edge technology like reading 2005 forum threads in 2023." 📂😂
    httrack.com/ #OfflineBrowsing #DigitalHoarding #InternetArchive #CatMemes #Nostalgia #HackerNews #ngated

  5. HTTrack - Der Website Downloader

    In diesem Tutorial zeige ich dir, wie du ganze Websites mit HTTrack für den Offline-Zugriff speichern kannst. Egal, ob für die eigene Sicherung oder einfach zum Stöbern ohne Internet – ich zeige dir Schritt für Schritt, wie es funktioniert.

    #httrack #Curl #wget #Website #Linux

    gnulinux.ch/httrack-der-websit

  6. HTTrack - Der Website Downloader

    In diesem Tutorial zeige ich dir, wie du ganze Websites mit HTTrack für den Offline-Zugriff speichern kannst. Egal, ob für die eigene Sicherung oder einfach zum Stöbern ohne Internet – ich zeige dir Schritt für Schritt, wie es funktioniert.

    #httrack #Curl #wget #Website #Linux

    gnulinux.ch/httrack-der-websit

  7. I am looking for archive.org as a self hosted service.

    I want to have an automated static copy of a website, which preserves old copied versions.

    It should provide a #crawler and a web interface to access the archived versions of the website.

    The use case is a lousy CMS which often destroys content. I want to be able to restore content from the archive and to have a static website copy in the worst case.

    #SelfHosting #WebsiteArchive #Archive #OffsiteBackp #Backup #HTTrack #WebsiteCopy

  8. I am looking for archive.org as a self hosted service.

    I want to have an automated static copy of a website, which preserves old copied versions.

    It should provide a #crawler and a web interface to access the archived versions of the website.

    The use case is a lousy CMS which often destroys content. I want to be able to restore content from the archive and to have a static website copy in the worst case.

    #SelfHosting #WebsiteArchive #Archive #OffsiteBackp #Backup #HTTrack #WebsiteCopy

  9. Actually, lemme think out loud about what I need #HTTrack to do, before I forget. It needs to pull jpg, png and svg images, javascript and any external CSS*) from any level within the Comicfury.com domain, but external links need to be skipped for mirroring.

    *)AFAIK all CSS within ComicFury is inline! A baffling decision but one that will make my life easier with this. But I may be mistaken.

  10. Actually, lemme think out loud about what I need #HTTrack to do, before I forget. It needs to pull jpg, png and svg images, javascript and any external CSS*) from any level within the Comicfury.com domain, but external links need to be skipped for mirroring.

    *)AFAIK all CSS within ComicFury is inline! A baffling decision but one that will make my life easier with this. But I may be mistaken.

  11. search engine on a stick would be a fun project 1tb nvme enc persistent bootable and you can spider your own sites in addition to top 10k sites already crawled and indexed - yacy could stand to be much more automated - it is a bit of work to get it set - not the config just all the sites loaded #httrack

  12. search engine on a stick would be a fun project 1tb nvme enc persistent bootable and you can spider your own sites in addition to top 10k sites already crawled and indexed - yacy could stand to be much more automated - it is a bit of work to get it set - not the config just all the sites loaded #httrack

  13. #HTTrack seems to be unmaintained (last release in 2017).

    Any maintained recent opensource mirroring solution than can offload auth to a browser (for example, like #destreamer can)?

    httrack.com/

  14. #HTTrack seems to be unmaintained (last release in 2017).

    Any maintained recent opensource mirroring solution than can offload auth to a browser (for example, like #destreamer can)?

    httrack.com/

  15. Manchmal will man ja auch eine ganze Webpräsenz sichern. #Httrack ist dafür auch ein gutes Tool, aber die Voreinstellungen müssen angepasst werden. #OSINT bashinho.de/2024/01/18/webseit

  16. Manchmal will man ja auch eine ganze Webpräsenz sichern. #Httrack ist dafür auch ein gutes Tool, aber die Voreinstellungen müssen angepasst werden. #OSINT bashinho.de/2024/01/18/webseit

  17. В очередной раз убеждаюсь, что #wget великая вещь!

    Одна мелкая бура сообщила о своём закрытии, и я решил её сохранить себе.

    Попробовал сначала
    #HTTrack, он пыхтел полдня и сохранил только html файлы.

    wget сначала отказывался зеркалить сайт, но я добавил
    -U и всё заработало. Примерно за 2 два часа он скачал весь сайт и все картинки.

    Теперь я обладаю ~1800 картинками среднего качества и не знаю что с этим делать.
    ​:blobcatshrug:​

  18. I would cancel one web server from an old-company, but wanted to keep the site somewhere (a php one).
    Using httrack and gitlab pages I could do it quite easily! The site looks exactly the same, now static, and no cost to keep it running (only need to pay the domain).

    Some days I like technology, mainly the free/libre ones :)

    #gitlab #floss #httrack

  19. Had to use #HTTrack to backup a website.
    So here is my 2 cents command line to download a little bit faster than the default options :

    humanize.me/nerd/httrack.html

    #backup #mirror #website

  20. Had to use #HTTrack to backup a website.
    So here is my 2 cents command line to download a little bit faster than the default options :

    humanize.me/nerd/httrack.html

    #backup #mirror #website

  21. Puras broncas al tratar de hacer un #WebScraping de un Google site, ni con el famoso #httrack

  22. @jgoerzen Sounds like #httrack but a little better, although in my experience html sites are a PITA -- It would be nice if it was literally a .tar archive containing the website & files like how many file formats are archives in disguise.

  23. I'm going slightly #mad about #httrack and #wget to load all *.opus files from @chaosradio.

    So it will be the good old fashioned +click-wait-click-wait-click-load+ way...

    But maybe I'm just holding it wrong.
    🙄

  24. Today I learned that #httrack requires the input address as a URL. So if you have a local file with links, you need to specify the file protocol. And apparently you need a complete absolute path, too.

    So for me that became:

    $ httrack file://mnt/c/Users/aveeltstra/AppData/Local/Temp/mirror/source.html -O ./ "+*.JPG" -Y

    And yes: that is accessing an MSWindows path from Debian on the Microsoft Linux Subsystem.