home.social

#searchengine — Public Fediverse posts

Live and recent posts from across the Fediverse tagged #searchengine, aggregated by home.social.

  1. Yesterday we were talking with Viktor on #Marginalia Discord. I was thinking the bottleneck in crawling is bandwidth. Viktor said something else:

    The bottleneck is not BW, not most of the time. But the problem's that a) the crawler is polite. So I don't send so many requests to the same host concurrently. For instance, so many websites on wordpress.com, geocities, github.io and similar.

    And b) Too many open connections will drive the network switch mad and it'll be problematic.

    Since yesterday, I'm using #wget to download websites for offline usage and later index them using #yacy. By default, wget is a polite crawler. I am fetching about ~10 websites concurrently. And the average BW usage is just 1MB/s. Fetching some websites is not done since yesterday till now...

    #SearchEngine #search_engine #crawling #web #opensource

  2. Yesterday we were talking with Viktor on #Marginalia Discord. I was thinking the bottleneck in crawling is bandwidth. Viktor said something else:

    The bottleneck is not BW, not most of the time. But the problem's that a) the crawler is polite. So I don't send so many requests to the same host concurrently. For instance, so many websites on wordpress.com, geocities, github.io and similar.

    And b) Too many open connections will drive the network switch mad and it'll be problematic.

    Since yesterday, I'm using #wget to download websites for offline usage and later index them using #yacy. By default, wget is a polite crawler. I am fetching about ~10 websites concurrently. And the average BW usage is just 1MB/s. Fetching some websites is not done since yesterday till now...

    #SearchEngine #search_engine #crawling #web #opensource

  3. Yesterday we were talking with Viktor on #Marginalia Discord. I was thinking the bottleneck in crawling is bandwidth. Viktor said something else:

    The bottleneck is not BW, not most of the time. But the problem's that a) the crawler is polite. So I don't send so many requests to the same host concurrently. For instance, so many websites on wordpress.com, geocities, github.io and similar.

    And b) Too many open connections will drive the network switch mad and it'll be problematic.

    Since yesterday, I'm using #wget to download websites for offline usage and later index them using #yacy. By default, wget is a polite crawler. I am fetching about ~10 websites concurrently. And the average BW usage is just 1MB/s. Fetching some websites is not done since yesterday till now...

    #SearchEngine #search_engine #crawling #web #opensource

  4. Yesterday we were talking with Viktor on #Marginalia Discord. I was thinking the bottleneck in crawling is bandwidth. Viktor said something else:

    The bottleneck is not BW, not most of the time. But the problem's that a) the crawler is polite. So I don't send so many requests to the same host concurrently. For instance, so many websites on wordpress.com, geocities, github.io and similar.

    And b) Too many open connections will drive the network switch mad and it'll be problematic.

    Since yesterday, I'm using #wget to download websites for offline usage and later index them using #yacy. By default, wget is a polite crawler. I am fetching about ~10 websites concurrently. And the average BW usage is just 1MB/s. Fetching some websites is not done since yesterday till now...

    #SearchEngine #search_engine #crawling #web #opensource

  5. Yesterday we were talking with Viktor on #Marginalia Discord. I was thinking the bottleneck in crawling is bandwidth. Viktor said something else:

    The bottleneck is not BW, not most of the time. But the problem's that a) the crawler is polite. So I don't send so many requests to the same host concurrently. For instance, so many websites on wordpress.com, geocities, github.io and similar.

    And b) Too many open connections will drive the network switch mad and it'll be problematic.

    Since yesterday, I'm using #wget to download websites for offline usage and later index them using #yacy. By default, wget is a polite crawler. I am fetching about ~10 websites concurrently. And the average BW usage is just 1MB/s. Fetching some websites is not done since yesterday till now...

    #SearchEngine #search_engine #crawling #web #opensource