#openwebsearcheu — Public Fediverse posts
Live and recent posts from across the Fediverse tagged #openwebsearcheu, aggregated by home.social.
-
Paper:
https://doi.org/10.1007/978-3-032-21289-4_25Want to try it yourself?
#openwebsearcheu book chapter:
https://openwebsearcheu-public.pages.it4i.eu/ows-the-book/content/dnt/remote-index.htmlCode repository for running remote queries using #DuckDB
https://gitlab.science.ru.nl/informagus/remote-querying -
Paper:
https://doi.org/10.1007/978-3-032-21289-4_25Want to try it yourself?
#openwebsearcheu book chapter:
https://openwebsearcheu-public.pages.it4i.eu/ows-the-book/content/dnt/remote-index.htmlCode repository for running remote queries using #DuckDB
https://gitlab.science.ru.nl/informagus/remote-querying -
Paper:
https://doi.org/10.1007/978-3-032-21289-4_25Want to try it yourself?
#openwebsearcheu book chapter:
https://openwebsearcheu-public.pages.it4i.eu/ows-the-book/content/dnt/remote-index.htmlCode repository for running remote queries using #DuckDB
https://gitlab.science.ru.nl/informagus/remote-querying -
Paper:
https://doi.org/10.1007/978-3-032-21289-4_25Want to try it yourself?
#openwebsearcheu book chapter:
https://openwebsearcheu-public.pages.it4i.eu/ows-the-book/content/dnt/remote-index.htmlCode repository for running remote queries using #DuckDB
https://gitlab.science.ru.nl/informagus/remote-querying -
Paper:
https://doi.org/10.1007/978-3-032-21289-4_25Want to try it yourself?
#openwebsearcheu book chapter:
https://openwebsearcheu-public.pages.it4i.eu/ows-the-book/content/dnt/remote-index.htmlCode repository for running remote queries using #DuckDB
https://gitlab.science.ru.nl/informagus/remote-querying -
Gijs Hendriksen presenting our work on "remote querying" to provide access to huge Web resources through de facto standard tech: Parquet files in S3 queried using #DuckDB to facilitate IR research at very acceptable latencies.
Run your ClueWeb experiment in 10 minutes or so, and repeat your experiments on recent Web data from the #openwebsearcheu Open Web Index.
-
Gijs Hendriksen presenting our work on "remote querying" to provide access to huge Web resources through de facto standard tech: Parquet files in S3 queried using #DuckDB to facilitate IR research at very acceptable latencies.
Run your ClueWeb experiment in 10 minutes or so, and repeat your experiments on recent Web data from the #openwebsearcheu Open Web Index.
-
Gijs Hendriksen presenting our work on "remote querying" to provide access to huge Web resources through de facto standard tech: Parquet files in S3 queried using #DuckDB to facilitate IR research at very acceptable latencies.
Run your ClueWeb experiment in 10 minutes or so, and repeat your experiments on recent Web data from the #openwebsearcheu Open Web Index.
-
Gijs Hendriksen presenting our work on "remote querying" to provide access to huge Web resources through de facto standard tech: Parquet files in S3 queried using #DuckDB to facilitate IR research at very acceptable latencies.
Run your ClueWeb experiment in 10 minutes or so, and repeat your experiments on recent Web data from the #openwebsearcheu Open Web Index.
-
Gijs Hendriksen presenting our work on "remote querying" to provide access to huge Web resources through de facto standard tech: Parquet files in S3 queried using #DuckDB to facilitate IR research at very acceptable latencies.
Run your ClueWeb experiment in 10 minutes or so, and repeat your experiments on recent Web data from the #openwebsearcheu Open Web Index.
-
I will be giving an invited talk at the #ECIR2026 IR4Good track about #OpenWebSearchEU: "Towards a shared infrastructure for assembling web search engines"
https://djoerdhiemstra.com/2026/towards-a-shared-infrastructure-for-assembling-web-search-engines/
-
I will be giving an invited talk at the #ECIR2026 IR4Good track about #OpenWebSearchEU: "Towards a shared infrastructure for assembling web search engines"
https://djoerdhiemstra.com/2026/towards-a-shared-infrastructure-for-assembling-web-search-engines/
-
I will be giving an invited talk at the #ECIR2026 IR4Good track about #OpenWebSearchEU: "Towards a shared infrastructure for assembling web search engines"
https://djoerdhiemstra.com/2026/towards-a-shared-infrastructure-for-assembling-web-search-engines/
-
I will be giving an invited talk at the #ECIR2026 IR4Good track about #OpenWebSearchEU: "Towards a shared infrastructure for assembling web search engines"
https://djoerdhiemstra.com/2026/towards-a-shared-infrastructure-for-assembling-web-search-engines/
-
I will be giving an invited talk at the #ECIR2026 IR4Good track about #OpenWebSearchEU: "Towards a shared infrastructure for assembling web search engines"
https://djoerdhiemstra.com/2026/towards-a-shared-infrastructure-for-assembling-web-search-engines/
-
Pour avoir une chance de devenir souverain @Qwant , @StartpageSearch , @ecosia et #OpenWebSearchEU devrait travailler ensemble.
#EuropeFirstUnitedSovereign #logicielLibre #technology #europe
-
Pour avoir une chance de devenir souverain @Qwant , @StartpageSearch , @ecosia et #OpenWebSearchEU devrait travailler ensemble.
#EuropeFirstUnitedSovereign #logicielLibre #technology #europe
-
Pour avoir une chance de devenir souverain @Qwant , @StartpageSearch , @ecosia et #OpenWebSearchEU devrait travailler ensemble.
#EuropeFirstUnitedSovereign #logicielLibre #technology #europe
-
Pour avoir une chance de devenir souverain @Qwant , @StartpageSearch , @ecosia et #OpenWebSearchEU devrait travailler ensemble.
#EuropeFirstUnitedSovereign #logicielLibre #technology #europe
-
Pour avoir une chance de devenir souverain @Qwant , @StartpageSearch , @ecosia et #OpenWebSearchEU devrait travailler ensemble.
#EuropeFirstUnitedSovereign #logicielLibre #technology #europe
-
#ECIR2026 notifications were friendly to me 🤗
1. Full paper "Open Web Indexes for Remote Querying" with @gijs and @djoerd.
Can we let ppl query the Terabytes of Web Index we collect in #OpenWebSearchEU in new ways, making good use of Parquet, S3, DuckDB?
Turns out the answer is a big YES!
Pre-print of the paper w/ code coming soon!
1/4
-
#ECIR2026 notifications were friendly to me 🤗
1. Full paper "Open Web Indexes for Remote Querying" with @gijs and @djoerd.
Can we let ppl query the Terabytes of Web Index we collect in #OpenWebSearchEU in new ways, making good use of Parquet, S3, DuckDB?
Turns out the answer is a big YES!
Pre-print of the paper w/ code coming soon!
1/4
-
#ECIR2026 notifications were friendly to me 🤗
1. Full paper "Open Web Indexes for Remote Querying" with @gijs and @djoerd.
Can we let ppl query the Terabytes of Web Index we collect in #OpenWebSearchEU in new ways, making good use of Parquet, S3, DuckDB?
Turns out the answer is a big YES!
Pre-print of the paper w/ code coming soon!
1/4
-
#ECIR2026 notifications were friendly to me 🤗
1. Full paper "Open Web Indexes for Remote Querying" with @gijs and @djoerd.
Can we let ppl query the Terabytes of Web Index we collect in #OpenWebSearchEU in new ways, making good use of Parquet, S3, DuckDB?
Turns out the answer is a big YES!
Pre-print of the paper w/ code coming soon!
1/4
-
#ECIR2026 notifications were friendly to me 🤗
1. Full paper "Open Web Indexes for Remote Querying" with @gijs and @djoerd.
Can we let ppl query the Terabytes of Web Index we collect in #OpenWebSearchEU in new ways, making good use of Parquet, S3, DuckDB?
Turns out the answer is a big YES!
Pre-print of the paper w/ code coming soon!
1/4
-
#OpenWebSearchEU is a silver sponsor of #ECIR2025!
-
#OpenWebSearchEU is a silver sponsor of #ECIR2025!
-
#OpenWebSearchEU is a silver sponsor of #ECIR2025!
-
#OpenWebSearchEU is a silver sponsor of #ECIR2025!
-
Open web index #OWI update:
4 billion URLs crawled
185 different languages
28 million Hosts
750 TB crawled
1 TB crawled per day
147 WARC Datasets
17.5 TB size of Open Web Index
28.8 TB size of WARC datasets
346 public datasets -
Open web index #OWI update:
4 billion URLs crawled
185 different languages
28 million Hosts
750 TB crawled
1 TB crawled per day
147 WARC Datasets
17.5 TB size of Open Web Index
28.8 TB size of WARC datasets
346 public datasets -
Open web index #OWI update:
4 billion URLs crawled
185 different languages
28 million Hosts
750 TB crawled
1 TB crawled per day
147 WARC Datasets
17.5 TB size of Open Web Index
28.8 TB size of WARC datasets
346 public datasets -
Open web index #OWI update:
4 billion URLs crawled
185 different languages
28 million Hosts
750 TB crawled
1 TB crawled per day
147 WARC Datasets
17.5 TB size of Open Web Index
28.8 TB size of WARC datasets
346 public datasets -
Open web index #OWI update:
4 billion URLs crawled
185 different languages
28 million Hosts
750 TB crawled
1 TB crawled per day
147 WARC Datasets
17.5 TB size of Open Web Index
28.8 TB size of WARC datasets
346 public datasets -
Today, Open Web Search consortium meeting at LRZ. #OpenWebSearchEU
-
Today, Open Web Search consortium meeting at LRZ. #OpenWebSearchEU
-
Today, Open Web Search consortium meeting at LRZ. #OpenWebSearchEU
-
Today, Open Web Search consortium meeting at LRZ. #OpenWebSearchEU
-
@heinragas Thanks for offering! #OpenWebSearchEU provides the data and index. I hope there will be a fully functional search engine at the end of the project. (build by anyone!) @openwebsearcheu @Negin
-
@heinragas Thanks for offering! #OpenWebSearchEU provides the data and index. I hope there will be a fully functional search engine at the end of the project. (build by anyone!) @openwebsearcheu @Negin
-
@heinragas Thanks for offering! #OpenWebSearchEU provides the data and index. I hope there will be a fully functional search engine at the end of the project. (build by anyone!) @openwebsearcheu @Negin
-
@heinragas Thanks for offering! #OpenWebSearchEU provides the data and index. I hope there will be a fully functional search engine at the end of the project. (build by anyone!) @openwebsearcheu @Negin
-
-
-
-
-
This week: #OpenWebSearchEU consortium meeting in Nijmegen!
-
This week: #OpenWebSearchEU consortium meeting in Nijmegen!
-
This week: #OpenWebSearchEU consortium meeting in Nijmegen!
-
This week: #OpenWebSearchEU consortium meeting in Nijmegen!