home.social

#openwebsearcheu — Public Fediverse posts

Live and recent posts from across the Fediverse tagged #openwebsearcheu, aggregated by home.social.

  1. Gijs Hendriksen presenting our work on "remote querying" to provide access to huge Web resources through de facto standard tech: Parquet files in S3 queried using #DuckDB to facilitate IR research at very acceptable latencies.

    Run your ClueWeb experiment in 10 minutes or so, and repeat your experiments on recent Web data from the #openwebsearcheu Open Web Index.

  2. Gijs Hendriksen presenting our work on "remote querying" to provide access to huge Web resources through de facto standard tech: Parquet files in S3 queried using #DuckDB to facilitate IR research at very acceptable latencies.

    Run your ClueWeb experiment in 10 minutes or so, and repeat your experiments on recent Web data from the #openwebsearcheu Open Web Index.

  3. Gijs Hendriksen presenting our work on "remote querying" to provide access to huge Web resources through de facto standard tech: Parquet files in S3 queried using #DuckDB to facilitate IR research at very acceptable latencies.

    Run your ClueWeb experiment in 10 minutes or so, and repeat your experiments on recent Web data from the #openwebsearcheu Open Web Index.

  4. Gijs Hendriksen presenting our work on "remote querying" to provide access to huge Web resources through de facto standard tech: Parquet files in S3 queried using #DuckDB to facilitate IR research at very acceptable latencies.

    Run your ClueWeb experiment in 10 minutes or so, and repeat your experiments on recent Web data from the #openwebsearcheu Open Web Index.

  5. Gijs Hendriksen presenting our work on "remote querying" to provide access to huge Web resources through de facto standard tech: Parquet files in S3 queried using #DuckDB to facilitate IR research at very acceptable latencies.

    Run your ClueWeb experiment in 10 minutes or so, and repeat your experiments on recent Web data from the #openwebsearcheu Open Web Index.

  6. #ECIR2026 notifications were friendly to me 🤗

    1. Full paper "Open Web Indexes for Remote Querying" with @gijs and @djoerd.

    Can we let ppl query the Terabytes of Web Index we collect in #OpenWebSearchEU in new ways, making good use of Parquet, S3, DuckDB?

    Turns out the answer is a big YES!

    Pre-print of the paper w/ code coming soon!

    1/4

  7. #ECIR2026 notifications were friendly to me 🤗

    1. Full paper "Open Web Indexes for Remote Querying" with @gijs and @djoerd.

    Can we let ppl query the Terabytes of Web Index we collect in #OpenWebSearchEU in new ways, making good use of Parquet, S3, DuckDB?

    Turns out the answer is a big YES!

    Pre-print of the paper w/ code coming soon!

    1/4

  8. #ECIR2026 notifications were friendly to me 🤗

    1. Full paper "Open Web Indexes for Remote Querying" with @gijs and @djoerd.

    Can we let ppl query the Terabytes of Web Index we collect in #OpenWebSearchEU in new ways, making good use of Parquet, S3, DuckDB?

    Turns out the answer is a big YES!

    Pre-print of the paper w/ code coming soon!

    1/4

  9. #ECIR2026 notifications were friendly to me 🤗

    1. Full paper "Open Web Indexes for Remote Querying" with @gijs and @djoerd.

    Can we let ppl query the Terabytes of Web Index we collect in #OpenWebSearchEU in new ways, making good use of Parquet, S3, DuckDB?

    Turns out the answer is a big YES!

    Pre-print of the paper w/ code coming soon!

    1/4

  10. #ECIR2026 notifications were friendly to me 🤗

    1. Full paper "Open Web Indexes for Remote Querying" with @gijs and @djoerd.

    Can we let ppl query the Terabytes of Web Index we collect in #OpenWebSearchEU in new ways, making good use of Parquet, S3, DuckDB?

    Turns out the answer is a big YES!

    Pre-print of the paper w/ code coming soon!

    1/4

  11. Open web index #OWI update:

    4 billion URLs crawled
    185 different languages
    28 million Hosts
    750 TB crawled
    1 TB crawled per day
    147 WARC Datasets
    17.5 TB size of Open Web Index
    28.8 TB size of WARC datasets
    346 public datasets

    #OpenWebSearchEU #OpenWebSearch

    ows.eu

  12. Open web index #OWI update:

    4 billion URLs crawled
    185 different languages
    28 million Hosts
    750 TB crawled
    1 TB crawled per day
    147 WARC Datasets
    17.5 TB size of Open Web Index
    28.8 TB size of WARC datasets
    346 public datasets

    #OpenWebSearchEU #OpenWebSearch

    ows.eu

  13. Open web index #OWI update:

    4 billion URLs crawled
    185 different languages
    28 million Hosts
    750 TB crawled
    1 TB crawled per day
    147 WARC Datasets
    17.5 TB size of Open Web Index
    28.8 TB size of WARC datasets
    346 public datasets

    #OpenWebSearchEU #OpenWebSearch

    ows.eu

  14. Open web index #OWI update:

    4 billion URLs crawled
    185 different languages
    28 million Hosts
    750 TB crawled
    1 TB crawled per day
    147 WARC Datasets
    17.5 TB size of Open Web Index
    28.8 TB size of WARC datasets
    346 public datasets

    #OpenWebSearchEU #OpenWebSearch

    ows.eu

  15. Open web index #OWI update:

    4 billion URLs crawled
    185 different languages
    28 million Hosts
    750 TB crawled
    1 TB crawled per day
    147 WARC Datasets
    17.5 TB size of Open Web Index
    28.8 TB size of WARC datasets
    346 public datasets

    #OpenWebSearchEU #OpenWebSearch

    ows.eu

  16. @heinragas Thanks for offering! #OpenWebSearchEU provides the data and index. I hope there will be a fully functional search engine at the end of the project. (build by anyone!) @openwebsearcheu @Negin

  17. @heinragas Thanks for offering! #OpenWebSearchEU provides the data and index. I hope there will be a fully functional search engine at the end of the project. (build by anyone!) @openwebsearcheu @Negin

  18. @heinragas Thanks for offering! #OpenWebSearchEU provides the data and index. I hope there will be a fully functional search engine at the end of the project. (build by anyone!) @openwebsearcheu @Negin

  19. @heinragas Thanks for offering! #OpenWebSearchEU provides the data and index. I hope there will be a fully functional search engine at the end of the project. (build by anyone!) @openwebsearcheu @Negin