#web-scraping — Public Fediverse posts
Live and recent posts from across the Fediverse tagged #web-scraping, aggregated by home.social.
-
Stephen Follows: What just happened to TheNumbers.com should worry us all. “Its hand-researched data is the highest quality, tracking box office grosses, budgets, home video and streaming across more than 78,000 films and 236,000 people. It gets north of eight million visitors a year, and is treated as THE definitive authority by journalists, academics, filmmakers, prediction markets, and even […]
https://rbfirehose.com/2026/07/24/stephen-follows-what-just-happened-to-thenumbers-com-should-worry-us-all/ -
Stephen Follows: What just happened to TheNumbers.com should worry us all. “Its hand-researched data is the highest quality, tracking box office grosses, budgets, home video and streaming across more than 78,000 films and 236,000 people. It gets north of eight million visitors a year, and is treated as THE definitive authority by journalists, academics, filmmakers, prediction markets, and even […]
https://rbfirehose.com/2026/07/24/stephen-follows-what-just-happened-to-thenumbers-com-should-worry-us-all/ -
Stephen Follows: What just happened to TheNumbers.com should worry us all. “Its hand-researched data is the highest quality, tracking box office grosses, budgets, home video and streaming across more than 78,000 films and 236,000 people. It gets north of eight million visitors a year, and is treated as THE definitive authority by journalists, academics, filmmakers, prediction markets, and even […]
https://rbfirehose.com/2026/07/24/stephen-follows-what-just-happened-to-thenumbers-com-should-worry-us-all/ -
Stephen Follows: What just happened to TheNumbers.com should worry us all. “Its hand-researched data is the highest quality, tracking box office grosses, budgets, home video and streaming across more than 78,000 films and 236,000 people. It gets north of eight million visitors a year, and is treated as THE definitive authority by journalists, academics, filmmakers, prediction markets, and even […]
https://rbfirehose.com/2026/07/24/stephen-follows-what-just-happened-to-thenumbers-com-should-worry-us-all/ -
Stephen Follows: What just happened to TheNumbers.com should worry us all. “Its hand-researched data is the highest quality, tracking box office grosses, budgets, home video and streaming across more than 78,000 films and 236,000 people. It gets north of eight million visitors a year, and is treated as THE definitive authority by journalists, academics, filmmakers, prediction markets, and even […]
https://rbfirehose.com/2026/07/24/stephen-follows-what-just-happened-to-thenumbers-com-should-worry-us-all/ -
Reuters: News Corp countersues Brave for allegedly ‘scraping’ articles for AI . “News Corp, facing a lawsuit by search engine Brave Software, has filed a countersuit accusing it of “flagrant theft” in distributing and selling versions of articles from the Wall Street Journal and New York Post to AI companies.”
https://rbfirehose.com/2026/07/23/reuters-news-corp-countersues-brave-for-allegedly-scraping-articles-for-ai/ -
Reuters: News Corp countersues Brave for allegedly ‘scraping’ articles for AI . “News Corp, facing a lawsuit by search engine Brave Software, has filed a countersuit accusing it of “flagrant theft” in distributing and selling versions of articles from the Wall Street Journal and New York Post to AI companies.”
https://rbfirehose.com/2026/07/23/reuters-news-corp-countersues-brave-for-allegedly-scraping-articles-for-ai/ -
Reuters: News Corp countersues Brave for allegedly ‘scraping’ articles for AI . “News Corp, facing a lawsuit by search engine Brave Software, has filed a countersuit accusing it of “flagrant theft” in distributing and selling versions of articles from the Wall Street Journal and New York Post to AI companies.”
https://rbfirehose.com/2026/07/23/reuters-news-corp-countersues-brave-for-allegedly-scraping-articles-for-ai/ -
Reuters: News Corp countersues Brave for allegedly ‘scraping’ articles for AI . “News Corp, facing a lawsuit by search engine Brave Software, has filed a countersuit accusing it of “flagrant theft” in distributing and selling versions of articles from the Wall Street Journal and New York Post to AI companies.”
https://rbfirehose.com/2026/07/23/reuters-news-corp-countersues-brave-for-allegedly-scraping-articles-for-ai/ -
Reuters: News Corp countersues Brave for allegedly ‘scraping’ articles for AI . “News Corp, facing a lawsuit by search engine Brave Software, has filed a countersuit accusing it of “flagrant theft” in distributing and selling versions of articles from the Wall Street Journal and New York Post to AI companies.”
https://rbfirehose.com/2026/07/23/reuters-news-corp-countersues-brave-for-allegedly-scraping-articles-for-ai/ -
Turns out you can't just ask an LLM for CSS selectors and ship them. In our scraping system, first-attempt selectors returned nothing 30 to 40% of the time. The trick that made it work: check JSON-LD first, then run every generated selector through a validation loop against the real DOM before trusting it. https://go.upgradejs.com/qru #LLM #WebScraping #AI
-
Turns out you can't just ask an LLM for CSS selectors and ship them. In our scraping system, first-attempt selectors returned nothing 30 to 40% of the time. The trick that made it work: check JSON-LD first, then run every generated selector through a validation loop against the real DOM before trusting it. https://go.upgradejs.com/qru #LLM #WebScraping #AI
-
Turns out you can't just ask an LLM for CSS selectors and ship them. In our scraping system, first-attempt selectors returned nothing 30 to 40% of the time. The trick that made it work: check JSON-LD first, then run every generated selector through a validation loop against the real DOM before trusting it. https://go.upgradejs.com/qru #LLM #WebScraping #AI
-
Turns out you can't just ask an LLM for CSS selectors and ship them. In our scraping system, first-attempt selectors returned nothing 30 to 40% of the time. The trick that made it work: check JSON-LD first, then run every generated selector through a validation loop against the real DOM before trusting it. https://go.upgradejs.com/qru #LLM #WebScraping #AI
-
MediaPost: Judge Dismisses Google Complaint Against SerpApi Over Scraping. “A federal judge has dismissed Google’s complaint against the Texas-based company SerpApi, which allegedly circumvented attempts to prevent it from scraping search results. The ruling, issued Monday by U.S. District Court Judge Yvonne Gonzalez Rogers, allows Google to amend its complaint and bring it again.”
https://rbfirehose.com/2026/07/22/mediapost-judge-dismisses-google-complaint-against-serpapi-over-scraping/ -
MediaPost: Judge Dismisses Google Complaint Against SerpApi Over Scraping. “A federal judge has dismissed Google’s complaint against the Texas-based company SerpApi, which allegedly circumvented attempts to prevent it from scraping search results. The ruling, issued Monday by U.S. District Court Judge Yvonne Gonzalez Rogers, allows Google to amend its complaint and bring it again.”
https://rbfirehose.com/2026/07/22/mediapost-judge-dismisses-google-complaint-against-serpapi-over-scraping/ -
MediaPost: Judge Dismisses Google Complaint Against SerpApi Over Scraping. “A federal judge has dismissed Google’s complaint against the Texas-based company SerpApi, which allegedly circumvented attempts to prevent it from scraping search results. The ruling, issued Monday by U.S. District Court Judge Yvonne Gonzalez Rogers, allows Google to amend its complaint and bring it again.”
https://rbfirehose.com/2026/07/22/mediapost-judge-dismisses-google-complaint-against-serpapi-over-scraping/ -
MediaPost: Judge Dismisses Google Complaint Against SerpApi Over Scraping. “A federal judge has dismissed Google’s complaint against the Texas-based company SerpApi, which allegedly circumvented attempts to prevent it from scraping search results. The ruling, issued Monday by U.S. District Court Judge Yvonne Gonzalez Rogers, allows Google to amend its complaint and bring it again.”
https://rbfirehose.com/2026/07/22/mediapost-judge-dismisses-google-complaint-against-serpapi-over-scraping/ -
MediaPost: Judge Dismisses Google Complaint Against SerpApi Over Scraping. “A federal judge has dismissed Google’s complaint against the Texas-based company SerpApi, which allegedly circumvented attempts to prevent it from scraping search results. The ruling, issued Monday by U.S. District Court Judge Yvonne Gonzalez Rogers, allows Google to amend its complaint and bring it again.”
https://rbfirehose.com/2026/07/22/mediapost-judge-dismisses-google-complaint-against-serpapi-over-scraping/ -
🕷️ Anakin-Inc/anakin
Converts websites to clean markdown or JSON with fallback scraping handlers and proxy auto-selection
⭐ Stars: 687
📅 Last Update: Jul 20, 2026https://github.com/Anakin-Inc/anakin
#selfhosted #homelab #selfhost #selfhosting #opensource #webscraping #api
-
🕷️ Anakin-Inc/anakin
Converts websites to clean markdown or JSON with fallback scraping handlers and proxy auto-selection
⭐ Stars: 687
📅 Last Update: Jul 20, 2026https://github.com/Anakin-Inc/anakin
#selfhosted #homelab #selfhost #selfhosting #opensource #webscraping #api
-
🕷️ Anakin-Inc/anakin
Converts websites to clean markdown or JSON with fallback scraping handlers and proxy auto-selection
⭐ Stars: 687
📅 Last Update: Jul 20, 2026https://github.com/Anakin-Inc/anakin
#selfhosted #homelab #selfhost #selfhosting #opensource #webscraping #api
-
Mixfont: Decoy Font. “Decoy font is a font that prints a decoy for every letter, making it more difficult for AI to read what you type. The font works by using separate spatial frequencies to communicate two different letters in the same space.”
https://rbfirehose.com/2026/07/18/mixfont-decoy-font/ -
Mixfont: Decoy Font. “Decoy font is a font that prints a decoy for every letter, making it more difficult for AI to read what you type. The font works by using separate spatial frequencies to communicate two different letters in the same space.”
https://rbfirehose.com/2026/07/18/mixfont-decoy-font/ -
Mixfont: Decoy Font. “Decoy font is a font that prints a decoy for every letter, making it more difficult for AI to read what you type. The font works by using separate spatial frequencies to communicate two different letters in the same space.”
https://rbfirehose.com/2026/07/18/mixfont-decoy-font/ -
Mixfont: Decoy Font. “Decoy font is a font that prints a decoy for every letter, making it more difficult for AI to read what you type. The font works by using separate spatial frequencies to communicate two different letters in the same space.”
https://rbfirehose.com/2026/07/18/mixfont-decoy-font/ -
Mixfont: Decoy Font. “Decoy font is a font that prints a decoy for every letter, making it more difficult for AI to read what you type. The font works by using separate spatial frequencies to communicate two different letters in the same space.”
https://rbfirehose.com/2026/07/18/mixfont-decoy-font/ -
Launch HN: Context.dev (YC S26) – API to get structured data from any website
Comments: https://news.ycombinator.com/item?id=48847562
#HackerNews #LaunchHN #ContextDev #API #WebScraping #YC #S26
-
Launch HN: Context.dev (YC S26) – API to get structured data from any website
Comments: https://news.ycombinator.com/item?id=48847562
#HackerNews #LaunchHN #ContextDev #API #WebScraping #YC #S26
-
Launch HN: Context.dev (YC S26) – API to get structured data from any website
Comments: https://news.ycombinator.com/item?id=48847562
#HackerNews #LaunchHN #ContextDev #API #WebScraping #YC #S26
-
Launch HN: Context.dev (YC S26) – API to get structured data from any website
Comments: https://news.ycombinator.com/item?id=48847562
#HackerNews #LaunchHN #ContextDev #API #WebScraping #YC #S26
-
Launch HN: Context.dev (YC S26) – API to get structured data from any website
Comments: https://news.ycombinator.com/item?id=48847562
#HackerNews #LaunchHN #ContextDev #API #WebScraping #YC #S26
-
Engadget: Cloudflare will filter out web crawlers that serve AI companies . “Cloudflare has announced plans to automatically block mixed-use web crawlers that index websites for search engines and act as AI agents and trainers at the same time. The company previously offered its customers the optional ability to prevent crawlers from scraping their sites for AI chatbots, but now Cloudflare’s […]
https://rbfirehose.com/2026/07/06/engadget-cloudflare-will-filter-out-web-crawlers-that-serve-ai-companies/ -
Engadget: Cloudflare will filter out web crawlers that serve AI companies . “Cloudflare has announced plans to automatically block mixed-use web crawlers that index websites for search engines and act as AI agents and trainers at the same time. The company previously offered its customers the optional ability to prevent crawlers from scraping their sites for AI chatbots, but now Cloudflare’s […]
https://rbfirehose.com/2026/07/06/engadget-cloudflare-will-filter-out-web-crawlers-that-serve-ai-companies/ -
Engadget: Cloudflare will filter out web crawlers that serve AI companies . “Cloudflare has announced plans to automatically block mixed-use web crawlers that index websites for search engines and act as AI agents and trainers at the same time. The company previously offered its customers the optional ability to prevent crawlers from scraping their sites for AI chatbots, but now Cloudflare’s […]
https://rbfirehose.com/2026/07/06/engadget-cloudflare-will-filter-out-web-crawlers-that-serve-ai-companies/ -
Engadget: Cloudflare will filter out web crawlers that serve AI companies . “Cloudflare has announced plans to automatically block mixed-use web crawlers that index websites for search engines and act as AI agents and trainers at the same time. The company previously offered its customers the optional ability to prevent crawlers from scraping their sites for AI chatbots, but now Cloudflare’s […]
https://rbfirehose.com/2026/07/06/engadget-cloudflare-will-filter-out-web-crawlers-that-serve-ai-companies/ -
Engadget: Cloudflare will filter out web crawlers that serve AI companies . “Cloudflare has announced plans to automatically block mixed-use web crawlers that index websites for search engines and act as AI agents and trainers at the same time. The company previously offered its customers the optional ability to prevent crawlers from scraping their sites for AI chatbots, but now Cloudflare’s […]
https://rbfirehose.com/2026/07/06/engadget-cloudflare-will-filter-out-web-crawlers-that-serve-ai-companies/ -
RT @NousResearch: Der Hermes-Agent liest die Webinhalte nun bis zu 60-mal schneller und zu 49-mal niedrigeren Kosten. Scraping-Backends übergeben saubere Inhalte direkt an den Agenten, ohne redundante Verarbeitungsschritte; große Seiten werden lokal gespeichert und bei Bedarf seitenweise abgerufen, sodass Sie die gleiche Qualität zu einem Bruchteil der Zeit und Kosten erhalten. Video
mehr auf Arint.info
#Effizienz #HermesAgent #Kostensenkung #Video #WebScraping #arint_info
-
RT @NousResearch: Der Hermes-Agent liest die Webinhalte nun bis zu 60-mal schneller und zu 49-mal niedrigeren Kosten. Scraping-Backends übergeben saubere Inhalte direkt an den Agenten, ohne redundante Verarbeitungsschritte; große Seiten werden lokal gespeichert und bei Bedarf seitenweise abgerufen, sodass Sie die gleiche Qualität zu einem Bruchteil der Zeit und Kosten erhalten. Video
mehr auf Arint.info
#Effizienz #HermesAgent #Kostensenkung #Video #WebScraping #arint_info
-
RT @NousResearch: Der Hermes-Agent liest die Webinhalte nun bis zu 60-mal schneller und zu 49-mal niedrigeren Kosten. Scraping-Backends übergeben saubere Inhalte direkt an den Agenten, ohne redundante Verarbeitungsschritte; große Seiten werden lokal gespeichert und bei Bedarf seitenweise abgerufen, sodass Sie die gleiche Qualität zu einem Bruchteil der Zeit und Kosten erhalten. Video
mehr auf Arint.info
#Effizienz #HermesAgent #Kostensenkung #Video #WebScraping #arint_info
-
I tested every way to scrape Amazon in 2026 — plain requests, Selenium, Playwright, free proxies, paid proxies. https://hackernoon.com/i-tried-every-way-to-scrape-amazon-in-2026-here-is-what-actually-works #webscraping
-
I tested every way to scrape Amazon in 2026 — plain requests, Selenium, Playwright, free proxies, paid proxies. https://hackernoon.com/i-tried-every-way-to-scrape-amazon-in-2026-here-is-what-actually-works #webscraping
-
I tested every way to scrape Amazon in 2026 — plain requests, Selenium, Playwright, free proxies, paid proxies. https://hackernoon.com/i-tried-every-way-to-scrape-amazon-in-2026-here-is-what-actually-works #webscraping
-
I tested every way to scrape Amazon in 2026 — plain requests, Selenium, Playwright, free proxies, paid proxies. https://hackernoon.com/i-tried-every-way-to-scrape-amazon-in-2026-here-is-what-actually-works #webscraping
-
I tested every way to scrape Amazon in 2026 — plain requests, Selenium, Playwright, free proxies, paid proxies. https://hackernoon.com/i-tried-every-way-to-scrape-amazon-in-2026-here-is-what-actually-works #webscraping
-
RT @NousResearch: Der Hermes-Agent liest das Web nun bis zu 60-mal schneller und 49-mal günstiger. Scraping-Backends übergeben saubere Inhalte direkt an den Agenten ohne redundante Verarbeitungsschritte; große Seiten werden lokal gespeichert und bei Bedarf aufgerufen, sodass Sie die gleiche Qualität zu einem Bruchteil der Zeit und Kosten erhalten. Video
mehr auf Arint.info
#HermesAgent #Kosteneffizienz #Performance #Technologie #WebScraping #arint_info
-
RT @NousResearch: Der Hermes-Agent liest das Web nun bis zu 60-mal schneller und 49-mal günstiger. Scraping-Backends übergeben saubere Inhalte direkt an den Agenten ohne redundante Verarbeitungsschritte; große Seiten werden lokal gespeichert und bei Bedarf aufgerufen, sodass Sie die gleiche Qualität zu einem Bruchteil der Zeit und Kosten erhalten. Video
mehr auf Arint.info
#HermesAgent #Kosteneffizienz #Performance #Technologie #WebScraping #arint_info
-
RT @NousResearch: Der Hermes-Agent liest das Web nun bis zu 60-mal schneller und 49-mal günstiger. Scraping-Backends übergeben saubere Inhalte direkt an den Agenten ohne redundante Verarbeitungsschritte; große Seiten werden lokal gespeichert und bei Bedarf aufgerufen, sodass Sie die gleiche Qualität zu einem Bruchteil der Zeit und Kosten erhalten. Video
mehr auf Arint.info
#HermesAgent #Kosteneffizienz #Performance #Technologie #WebScraping #arint_info
-
Why configure your AI Model Harness? #webscraping #podcast #ai #programming https://www.youtube.com/watch?v=xdVcK-XfxLo?utm_campaign=blog-posts&utm_activity=ORS&utm_medium=social&utm_source=mastodon
-
Why configure your AI Model Harness? #webscraping #podcast #ai #programming https://www.youtube.com/watch?v=xdVcK-XfxLo?utm_campaign=blog-posts&utm_activity=ORS&utm_medium=social&utm_source=mastodon
-
Why configure your AI Model Harness? #webscraping #podcast #ai #programming https://www.youtube.com/watch?v=xdVcK-XfxLo?utm_campaign=blog-posts&utm_activity=ORS&utm_medium=social&utm_source=mastodon
-
Keep your context window clean using these #webscraping #ai #podcast https://www.youtube.com/watch?v=wNbLzW1huZM?utm_campaign=blog-posts&utm_activity=ORS&utm_medium=social&utm_source=mastodon
-
Keep your context window clean using these #webscraping #ai #podcast https://www.youtube.com/watch?v=wNbLzW1huZM?utm_campaign=blog-posts&utm_activity=ORS&utm_medium=social&utm_source=mastodon
-
Keep your context window clean using these #webscraping #ai #podcast https://www.youtube.com/watch?v=wNbLzW1huZM?utm_campaign=blog-posts&utm_activity=ORS&utm_medium=social&utm_source=mastodon
-
New Jersey Globe: Nearly 400 local newspapers sue OpenAI, Microsoft over alleged copyright theft. “The massive coalition of local newspaper publishers filed a federal lawsuit today against OpenAI and Microsoft, alleging the technology companies systematically copied copyrighted reporting from nearly 400 local newspapers to train and develop commercial artificial intelligence products, including […]
https://rbfirehose.com/2026/06/25/new-jersey-globe-nearly-400-local-newspapers-sue-openai-microsoft-over-alleged-copyright-theft/ -
New Jersey Globe: Nearly 400 local newspapers sue OpenAI, Microsoft over alleged copyright theft. “The massive coalition of local newspaper publishers filed a federal lawsuit today against OpenAI and Microsoft, alleging the technology companies systematically copied copyrighted reporting from nearly 400 local newspapers to train and develop commercial artificial intelligence products, including […]
https://rbfirehose.com/2026/06/25/new-jersey-globe-nearly-400-local-newspapers-sue-openai-microsoft-over-alleged-copyright-theft/ -
New Jersey Globe: Nearly 400 local newspapers sue OpenAI, Microsoft over alleged copyright theft. “The massive coalition of local newspaper publishers filed a federal lawsuit today against OpenAI and Microsoft, alleging the technology companies systematically copied copyrighted reporting from nearly 400 local newspapers to train and develop commercial artificial intelligence products, including […]
https://rbfirehose.com/2026/06/25/new-jersey-globe-nearly-400-local-newspapers-sue-openai-microsoft-over-alleged-copyright-theft/ -
New Jersey Globe: Nearly 400 local newspapers sue OpenAI, Microsoft over alleged copyright theft. “The massive coalition of local newspaper publishers filed a federal lawsuit today against OpenAI and Microsoft, alleging the technology companies systematically copied copyrighted reporting from nearly 400 local newspapers to train and develop commercial artificial intelligence products, including […]
https://rbfirehose.com/2026/06/25/new-jersey-globe-nearly-400-local-newspapers-sue-openai-microsoft-over-alleged-copyright-theft/ -
New Jersey Globe: Nearly 400 local newspapers sue OpenAI, Microsoft over alleged copyright theft. “The massive coalition of local newspaper publishers filed a federal lawsuit today against OpenAI and Microsoft, alleging the technology companies systematically copied copyrighted reporting from nearly 400 local newspapers to train and develop commercial artificial intelligence products, including […]
https://rbfirehose.com/2026/06/25/new-jersey-globe-nearly-400-local-newspapers-sue-openai-microsoft-over-alleged-copyright-theft/ -
Search Engine Journal: US Publishers Demand Common Crawl Stop Scraping Their Content. “Digital Content Next, a trade body representing US digital publishers, has sent a cease and desist letter to the Common Crawl Foundation. The letter demands Common Crawl stop collecting publisher content and remove material already in its datasets.”
https://rbfirehose.com/2026/06/11/search-engine-journal-us-publishers-demand-common-crawl-stop-scraping-their-content/