#stormcrawler — Public Fediverse posts
Live and recent posts from across the Fediverse tagged #stormcrawler, aggregated by home.social.
-
And the vote is open for #StormCrawler #ApacheStormCrawler 3.2.0!
Many thanks to Richard Zowalla and team for helping me through the release process!
https://lists.apache.org/thread/2ftf254mdwks4llogxfb8w0qx5fj1qgh
-
#stormcrawler has now left the DigitalPebble organisation on GitHub on its way to Apache.
I am not going to lie, I am feeling something: I created SC more than 10 years ago and it became the focal point of my professional activities since.
I am also 100% convinced that it is the right move, at the right time and hope that being an ASF project will help its adoption by users and new contributors -
-
@jhy thanks! I'll give it a try and upgrade to it before releasing the next version of #StormCrawler
-
Meet the #StormCrawler users is back on our blog! We are delighted to share this Q&A with members of the OpenWebSearch.eu team. Come and read about their project and how they use both #StormCrawler and #URLFrontier to help deliver a truly open, transparent and legally compliant alternative to the big search engines.
#opensource #openwebsearch #opendata #innovation
https://digitalpebble.blogspot.com/2023/11/meet-stormcrawler-users-q-with-open-web.html
-
#StormCrawler 2.10 is out!
https://github.com/DigitalPebble/storm-crawler/releases/tag/2.10
We have also written a short blog detailing the improvements to the protocol implementations
https://digitalpebble.blogspot.com/2023/10/focus-on-protocol-improvements-in.html
-
Really proud to see both #stormcrawler and #URLFrontier used by OWLer
#OSSYM23 -
If you use #StormCrawler, it makes you an #ApacheStorm user.
Please help the project by filling the survey on https://terminplaner4.dfn.de/EYNJzD9U64UFGOGq -
@tallison @OpenSearchProject
Or just use #StormCrawler ? :mastoinnocent: -
@digitalpebble sadly homegrown one off: https://github.com/tballison/file-observatory/tree/main/commoncrawl-fetcher
If I were to do it again, I’d use #ApacheNutch or #StormCrawler
-
Have added test coverage for #StormCrawler
https://coveralls.io/github/DigitalPebble/storm-crawler?branch=master
As expected pretty low on average, partly explained by the fact that writing tests for Bolts is not trivial but at least we can now see where new tests should be added.
BTW #tests are great #opensource #contributions
-
Hoping to benchmark #StormCrawler + #opensearch with segment replication
-
@davidshq
Definitely. To give an example, one of the top EU online retailers use #StormCrawler but won't publicise (or sponsor) it. Their legal department advised them not to because it would expose the way they use it and that is seen as a risk. -
What are your favorite / the best #WebCrawlers for broad / #WebScale #crawling?
I've built a list but am looking for anything I missed: https://github.com/davidshq/awesome-search-engines/blob/main/WebCrawlers.md
Main options I've found include #Apache #Nutch, #StormCrawler, #Scrapy, #Norconex, #PulsarR, #Heritrix, and #sparkler
-
@elan also see #Heritrix and of course #StormCrawler as alternatives to #ApacheNutch
-
We're super excited about #StormCrawler being used by the #OpenWebSearch project.
-
Should we support tracing in #stormcrawler? Anyone using tools like Datadog when crawling to track slow URLs and bottlenecks?
-
Missing Link: Offener Web-Index soll Europa bei der Suche unabhängig machen
Mit der von der EU geförderten Entwicklung eines Open Web Index wollen Forscher die Dominanz von Google & Co. brechen und das menschliche Wissen verbreitern.
#SearchEngine #eu #OWI #EuropeanOpenWebIndex #OWSAI #OpenWebSearchAndAnalysisInfrastructure #OpenWebSearch #Suma #OSF #Serci #StormCrawler
Ferner #Gigablast #FindX #Quaero #Theseus #CommonCrawls
-
We are pleased to announce that DigitalPebble Ltd is a partner of the OpenSearch Project.
In case you have missed it, #StormCrawler has a module for #OpenSearch since its latest release and hopefully there will be more good things to come!
-
Call to all #StormCrawler users: we will release a new version shortly so that people can benefit from the latest additions (#Opensearch) and improvements (#WARC). Any chance you could test some crawls with the latest code in the main branch and report any issues? Thanks
-
Just committed a Maven #archetype for crawling with the #OpenSearch module of #StormCrawler.
-
Just opened a PR to port the content of the #Elasticsearch module of #StormCrawler to #OpenSearch
includes simple #dashboards
Feedback welcome as usual
-
A very nice contribution to #StormCrawler improving the generation of #WARC files
-
#StormCrawler 2.6 released
https://github.com/DigitalPebble/storm-crawler/releases/tag/2.6
Thanks to our contributors and users
-
There is a paradox with the sponsoring of #StormCrawler: the only organisations who have financially supported our work are very small, typically less than 5 employees. Meanwhile, larger ones (some of which have multi-million $£€ budgets and use SC on a large scale) do not donate at all, nor contribute any code. Most of them are also very reluctant to acknowledging publicly their use of it. Is it down to the bureaucratic hassle of convincing ppl up the decision ladder? What do you think?
-
Now merged. This will be in the next release of #StormCrawler
-
Fancy trying the new version of the #StormCrawler archetype which uses #URLFrontier as a backend?