home.social

#web-scraping — Public Fediverse posts

Live and recent posts from across the Fediverse tagged #web-scraping, aggregated by home.social.

fetched live
  1. Stephen Follows: What just happened to TheNumbers.com should worry us all. “Its hand-researched data is the highest quality, tracking box office grosses, budgets, home video and streaming across more than 78,000 films and 236,000 people. It gets north of eight million visitors a year, and is treated as THE definitive authority by journalists, academics, filmmakers, prediction markets, and even […]

    https://rbfirehose.com/2026/07/24/stephen-follows-what-just-happened-to-thenumbers-com-should-worry-us-all/
  2. Stephen Follows: What just happened to TheNumbers.com should worry us all. “Its hand-researched data is the highest quality, tracking box office grosses, budgets, home video and streaming across more than 78,000 films and 236,000 people. It gets north of eight million visitors a year, and is treated as THE definitive authority by journalists, academics, filmmakers, prediction markets, and even […]

    https://rbfirehose.com/2026/07/24/stephen-follows-what-just-happened-to-thenumbers-com-should-worry-us-all/
  3. Stephen Follows: What just happened to TheNumbers.com should worry us all. “Its hand-researched data is the highest quality, tracking box office grosses, budgets, home video and streaming across more than 78,000 films and 236,000 people. It gets north of eight million visitors a year, and is treated as THE definitive authority by journalists, academics, filmmakers, prediction markets, and even […]

    https://rbfirehose.com/2026/07/24/stephen-follows-what-just-happened-to-thenumbers-com-should-worry-us-all/
  4. Stephen Follows: What just happened to TheNumbers.com should worry us all. “Its hand-researched data is the highest quality, tracking box office grosses, budgets, home video and streaming across more than 78,000 films and 236,000 people. It gets north of eight million visitors a year, and is treated as THE definitive authority by journalists, academics, filmmakers, prediction markets, and even […]

    https://rbfirehose.com/2026/07/24/stephen-follows-what-just-happened-to-thenumbers-com-should-worry-us-all/
  5. Stephen Follows: What just happened to TheNumbers.com should worry us all. “Its hand-researched data is the highest quality, tracking box office grosses, budgets, home video and streaming across more than 78,000 films and 236,000 people. It gets north of eight million visitors a year, and is treated as THE definitive authority by journalists, academics, filmmakers, prediction markets, and even […]

    https://rbfirehose.com/2026/07/24/stephen-follows-what-just-happened-to-thenumbers-com-should-worry-us-all/
  6. Reuters: News Corp countersues Brave for allegedly ‘scraping’ articles for AI . “News Corp, facing a lawsuit by search engine Brave Software, has filed a countersuit accusing it ‌of “flagrant theft” in distributing and selling versions of articles from the Wall Street Journal and New York Post to AI companies.”

    https://rbfirehose.com/2026/07/23/reuters-news-corp-countersues-brave-for-allegedly-scraping-articles-for-ai/
  7. Reuters: News Corp countersues Brave for allegedly ‘scraping’ articles for AI . “News Corp, facing a lawsuit by search engine Brave Software, has filed a countersuit accusing it ‌of “flagrant theft” in distributing and selling versions of articles from the Wall Street Journal and New York Post to AI companies.”

    https://rbfirehose.com/2026/07/23/reuters-news-corp-countersues-brave-for-allegedly-scraping-articles-for-ai/
  8. Reuters: News Corp countersues Brave for allegedly ‘scraping’ articles for AI . “News Corp, facing a lawsuit by search engine Brave Software, has filed a countersuit accusing it ‌of “flagrant theft” in distributing and selling versions of articles from the Wall Street Journal and New York Post to AI companies.”

    https://rbfirehose.com/2026/07/23/reuters-news-corp-countersues-brave-for-allegedly-scraping-articles-for-ai/
  9. Reuters: News Corp countersues Brave for allegedly ‘scraping’ articles for AI . “News Corp, facing a lawsuit by search engine Brave Software, has filed a countersuit accusing it ‌of “flagrant theft” in distributing and selling versions of articles from the Wall Street Journal and New York Post to AI companies.”

    https://rbfirehose.com/2026/07/23/reuters-news-corp-countersues-brave-for-allegedly-scraping-articles-for-ai/
  10. Reuters: News Corp countersues Brave for allegedly ‘scraping’ articles for AI . “News Corp, facing a lawsuit by search engine Brave Software, has filed a countersuit accusing it ‌of “flagrant theft” in distributing and selling versions of articles from the Wall Street Journal and New York Post to AI companies.”

    https://rbfirehose.com/2026/07/23/reuters-news-corp-countersues-brave-for-allegedly-scraping-articles-for-ai/
  11. Turns out you can't just ask an LLM for CSS selectors and ship them. In our scraping system, first-attempt selectors returned nothing 30 to 40% of the time. The trick that made it work: check JSON-LD first, then run every generated selector through a validation loop against the real DOM before trusting it. go.upgradejs.com/qru #LLM #WebScraping #AI

  12. Turns out you can't just ask an LLM for CSS selectors and ship them. In our scraping system, first-attempt selectors returned nothing 30 to 40% of the time. The trick that made it work: check JSON-LD first, then run every generated selector through a validation loop against the real DOM before trusting it. go.upgradejs.com/qru #LLM #WebScraping #AI

  13. Turns out you can't just ask an LLM for CSS selectors and ship them. In our scraping system, first-attempt selectors returned nothing 30 to 40% of the time. The trick that made it work: check JSON-LD first, then run every generated selector through a validation loop against the real DOM before trusting it. go.upgradejs.com/qru #LLM #WebScraping #AI

  14. Turns out you can't just ask an LLM for CSS selectors and ship them. In our scraping system, first-attempt selectors returned nothing 30 to 40% of the time. The trick that made it work: check JSON-LD first, then run every generated selector through a validation loop against the real DOM before trusting it. go.upgradejs.com/qru #LLM #WebScraping #AI

  15. MediaPost: Judge Dismisses Google Complaint Against SerpApi Over Scraping. “A federal judge has dismissed Google’s complaint against the Texas-based company SerpApi, which allegedly circumvented attempts to prevent it from scraping search results. The ruling, issued Monday by U.S. District Court Judge Yvonne Gonzalez Rogers, allows Google to amend its complaint and bring it again.”

    https://rbfirehose.com/2026/07/22/mediapost-judge-dismisses-google-complaint-against-serpapi-over-scraping/
  16. MediaPost: Judge Dismisses Google Complaint Against SerpApi Over Scraping. “A federal judge has dismissed Google’s complaint against the Texas-based company SerpApi, which allegedly circumvented attempts to prevent it from scraping search results. The ruling, issued Monday by U.S. District Court Judge Yvonne Gonzalez Rogers, allows Google to amend its complaint and bring it again.”

    https://rbfirehose.com/2026/07/22/mediapost-judge-dismisses-google-complaint-against-serpapi-over-scraping/
  17. MediaPost: Judge Dismisses Google Complaint Against SerpApi Over Scraping. “A federal judge has dismissed Google’s complaint against the Texas-based company SerpApi, which allegedly circumvented attempts to prevent it from scraping search results. The ruling, issued Monday by U.S. District Court Judge Yvonne Gonzalez Rogers, allows Google to amend its complaint and bring it again.”

    https://rbfirehose.com/2026/07/22/mediapost-judge-dismisses-google-complaint-against-serpapi-over-scraping/
  18. MediaPost: Judge Dismisses Google Complaint Against SerpApi Over Scraping. “A federal judge has dismissed Google’s complaint against the Texas-based company SerpApi, which allegedly circumvented attempts to prevent it from scraping search results. The ruling, issued Monday by U.S. District Court Judge Yvonne Gonzalez Rogers, allows Google to amend its complaint and bring it again.”

    https://rbfirehose.com/2026/07/22/mediapost-judge-dismisses-google-complaint-against-serpapi-over-scraping/
  19. MediaPost: Judge Dismisses Google Complaint Against SerpApi Over Scraping. “A federal judge has dismissed Google’s complaint against the Texas-based company SerpApi, which allegedly circumvented attempts to prevent it from scraping search results. The ruling, issued Monday by U.S. District Court Judge Yvonne Gonzalez Rogers, allows Google to amend its complaint and bring it again.”

    https://rbfirehose.com/2026/07/22/mediapost-judge-dismisses-google-complaint-against-serpapi-over-scraping/
  20. 🕷️ Anakin-Inc/anakin

    Converts websites to clean markdown or JSON with fallback scraping handlers and proxy auto-selection

    ⭐ Stars: 687
    📅 Last Update: Jul 20, 2026

    github.com/Anakin-Inc/anakin

    #selfhosted #homelab #selfhost #selfhosting #opensource #webscraping #api

  21. 🕷️ Anakin-Inc/anakin

    Converts websites to clean markdown or JSON with fallback scraping handlers and proxy auto-selection

    ⭐ Stars: 687
    📅 Last Update: Jul 20, 2026

    github.com/Anakin-Inc/anakin

    #selfhosted #homelab #selfhost #selfhosting #opensource #webscraping #api

  22. 🕷️ Anakin-Inc/anakin

    Converts websites to clean markdown or JSON with fallback scraping handlers and proxy auto-selection

    ⭐ Stars: 687
    📅 Last Update: Jul 20, 2026

    github.com/Anakin-Inc/anakin

    #selfhosted #homelab #selfhost #selfhosting #opensource #webscraping #api

  23. Mixfont: Decoy Font. “Decoy font is a font that prints a decoy for every letter, making it more difficult for AI to read what you type. The font works by using separate spatial frequencies to communicate two different letters in the same space.”

    https://rbfirehose.com/2026/07/18/mixfont-decoy-font/
  24. Mixfont: Decoy Font. “Decoy font is a font that prints a decoy for every letter, making it more difficult for AI to read what you type. The font works by using separate spatial frequencies to communicate two different letters in the same space.”

    https://rbfirehose.com/2026/07/18/mixfont-decoy-font/
  25. Mixfont: Decoy Font. “Decoy font is a font that prints a decoy for every letter, making it more difficult for AI to read what you type. The font works by using separate spatial frequencies to communicate two different letters in the same space.”

    https://rbfirehose.com/2026/07/18/mixfont-decoy-font/
  26. Mixfont: Decoy Font. “Decoy font is a font that prints a decoy for every letter, making it more difficult for AI to read what you type. The font works by using separate spatial frequencies to communicate two different letters in the same space.”

    https://rbfirehose.com/2026/07/18/mixfont-decoy-font/
  27. Mixfont: Decoy Font. “Decoy font is a font that prints a decoy for every letter, making it more difficult for AI to read what you type. The font works by using separate spatial frequencies to communicate two different letters in the same space.”

    https://rbfirehose.com/2026/07/18/mixfont-decoy-font/
  28. Engadget: Cloudflare will filter out web crawlers that serve AI companies . “Cloudflare has announced plans to automatically block mixed-use web crawlers that index websites for search engines and act as AI agents and trainers at the same time. The company previously offered its customers the optional ability to prevent crawlers from scraping their sites for AI chatbots, but now Cloudflare’s […]

    https://rbfirehose.com/2026/07/06/engadget-cloudflare-will-filter-out-web-crawlers-that-serve-ai-companies/
  29. Engadget: Cloudflare will filter out web crawlers that serve AI companies . “Cloudflare has announced plans to automatically block mixed-use web crawlers that index websites for search engines and act as AI agents and trainers at the same time. The company previously offered its customers the optional ability to prevent crawlers from scraping their sites for AI chatbots, but now Cloudflare’s […]

    https://rbfirehose.com/2026/07/06/engadget-cloudflare-will-filter-out-web-crawlers-that-serve-ai-companies/
  30. Engadget: Cloudflare will filter out web crawlers that serve AI companies . “Cloudflare has announced plans to automatically block mixed-use web crawlers that index websites for search engines and act as AI agents and trainers at the same time. The company previously offered its customers the optional ability to prevent crawlers from scraping their sites for AI chatbots, but now Cloudflare’s […]

    https://rbfirehose.com/2026/07/06/engadget-cloudflare-will-filter-out-web-crawlers-that-serve-ai-companies/
  31. Engadget: Cloudflare will filter out web crawlers that serve AI companies . “Cloudflare has announced plans to automatically block mixed-use web crawlers that index websites for search engines and act as AI agents and trainers at the same time. The company previously offered its customers the optional ability to prevent crawlers from scraping their sites for AI chatbots, but now Cloudflare’s […]

    https://rbfirehose.com/2026/07/06/engadget-cloudflare-will-filter-out-web-crawlers-that-serve-ai-companies/
  32. Engadget: Cloudflare will filter out web crawlers that serve AI companies . “Cloudflare has announced plans to automatically block mixed-use web crawlers that index websites for search engines and act as AI agents and trainers at the same time. The company previously offered its customers the optional ability to prevent crawlers from scraping their sites for AI chatbots, but now Cloudflare’s […]

    https://rbfirehose.com/2026/07/06/engadget-cloudflare-will-filter-out-web-crawlers-that-serve-ai-companies/
  33. RT @NousResearch: Der Hermes-Agent liest die Webinhalte nun bis zu 60-mal schneller und zu 49-mal niedrigeren Kosten. Scraping-Backends übergeben saubere Inhalte direkt an den Agenten, ohne redundante Verarbeitungsschritte; große Seiten werden lokal gespeichert und bei Bedarf seitenweise abgerufen, sodass Sie die gleiche Qualität zu einem Bruchteil der Zeit und Kosten erhalten. Video

    mehr auf Arint.info

    #Effizienz #HermesAgent #Kostensenkung #Video #WebScraping #arint_info

    https://x.com/NousResearch/status/2071974594961977727#m

  34. RT @NousResearch: Der Hermes-Agent liest die Webinhalte nun bis zu 60-mal schneller und zu 49-mal niedrigeren Kosten. Scraping-Backends übergeben saubere Inhalte direkt an den Agenten, ohne redundante Verarbeitungsschritte; große Seiten werden lokal gespeichert und bei Bedarf seitenweise abgerufen, sodass Sie die gleiche Qualität zu einem Bruchteil der Zeit und Kosten erhalten. Video

    mehr auf Arint.info

    #Effizienz #HermesAgent #Kostensenkung #Video #WebScraping #arint_info

    https://x.com/NousResearch/status/2071974594961977727#m

  35. RT @NousResearch: Der Hermes-Agent liest die Webinhalte nun bis zu 60-mal schneller und zu 49-mal niedrigeren Kosten. Scraping-Backends übergeben saubere Inhalte direkt an den Agenten, ohne redundante Verarbeitungsschritte; große Seiten werden lokal gespeichert und bei Bedarf seitenweise abgerufen, sodass Sie die gleiche Qualität zu einem Bruchteil der Zeit und Kosten erhalten. Video

    mehr auf Arint.info

    #Effizienz #HermesAgent #Kostensenkung #Video #WebScraping #arint_info

    https://x.com/NousResearch/status/2071974594961977727#m

  36. I tested every way to scrape Amazon in 2026 — plain requests, Selenium, Playwright, free proxies, paid proxies. hackernoon.com/i-tried-every-w #webscraping

  37. I tested every way to scrape Amazon in 2026 — plain requests, Selenium, Playwright, free proxies, paid proxies. hackernoon.com/i-tried-every-w #webscraping

  38. I tested every way to scrape Amazon in 2026 — plain requests, Selenium, Playwright, free proxies, paid proxies. hackernoon.com/i-tried-every-w #webscraping

  39. I tested every way to scrape Amazon in 2026 — plain requests, Selenium, Playwright, free proxies, paid proxies. hackernoon.com/i-tried-every-w

  40. I tested every way to scrape Amazon in 2026 — plain requests, Selenium, Playwright, free proxies, paid proxies. hackernoon.com/i-tried-every-w #webscraping

  41. RT @NousResearch: Der Hermes-Agent liest das Web nun bis zu 60-mal schneller und 49-mal günstiger. Scraping-Backends übergeben saubere Inhalte direkt an den Agenten ohne redundante Verarbeitungsschritte; große Seiten werden lokal gespeichert und bei Bedarf aufgerufen, sodass Sie die gleiche Qualität zu einem Bruchteil der Zeit und Kosten erhalten. Video

    mehr auf Arint.info

    #HermesAgent #Kosteneffizienz #Performance #Technologie #WebScraping #arint_info

    https://x.com/NousResearch/status/2071974594961977727#m

  42. RT @NousResearch: Der Hermes-Agent liest das Web nun bis zu 60-mal schneller und 49-mal günstiger. Scraping-Backends übergeben saubere Inhalte direkt an den Agenten ohne redundante Verarbeitungsschritte; große Seiten werden lokal gespeichert und bei Bedarf aufgerufen, sodass Sie die gleiche Qualität zu einem Bruchteil der Zeit und Kosten erhalten. Video

    mehr auf Arint.info

    #HermesAgent #Kosteneffizienz #Performance #Technologie #WebScraping #arint_info

    https://x.com/NousResearch/status/2071974594961977727#m

  43. RT @NousResearch: Der Hermes-Agent liest das Web nun bis zu 60-mal schneller und 49-mal günstiger. Scraping-Backends übergeben saubere Inhalte direkt an den Agenten ohne redundante Verarbeitungsschritte; große Seiten werden lokal gespeichert und bei Bedarf aufgerufen, sodass Sie die gleiche Qualität zu einem Bruchteil der Zeit und Kosten erhalten. Video

    mehr auf Arint.info

    #HermesAgent #Kosteneffizienz #Performance #Technologie #WebScraping #arint_info

    https://x.com/NousResearch/status/2071974594961977727#m

  44. New Jersey Globe: Nearly 400 local newspapers sue OpenAI, Microsoft over alleged copyright theft. “The massive coalition of local newspaper publishers filed a federal lawsuit today against OpenAI and Microsoft, alleging the technology companies systematically copied copyrighted reporting from nearly 400 local newspapers to train and develop commercial artificial intelligence products, including […]

    https://rbfirehose.com/2026/06/25/new-jersey-globe-nearly-400-local-newspapers-sue-openai-microsoft-over-alleged-copyright-theft/
  45. New Jersey Globe: Nearly 400 local newspapers sue OpenAI, Microsoft over alleged copyright theft. “The massive coalition of local newspaper publishers filed a federal lawsuit today against OpenAI and Microsoft, alleging the technology companies systematically copied copyrighted reporting from nearly 400 local newspapers to train and develop commercial artificial intelligence products, including […]

    https://rbfirehose.com/2026/06/25/new-jersey-globe-nearly-400-local-newspapers-sue-openai-microsoft-over-alleged-copyright-theft/
  46. New Jersey Globe: Nearly 400 local newspapers sue OpenAI, Microsoft over alleged copyright theft. “The massive coalition of local newspaper publishers filed a federal lawsuit today against OpenAI and Microsoft, alleging the technology companies systematically copied copyrighted reporting from nearly 400 local newspapers to train and develop commercial artificial intelligence products, including […]

    https://rbfirehose.com/2026/06/25/new-jersey-globe-nearly-400-local-newspapers-sue-openai-microsoft-over-alleged-copyright-theft/
  47. New Jersey Globe: Nearly 400 local newspapers sue OpenAI, Microsoft over alleged copyright theft. “The massive coalition of local newspaper publishers filed a federal lawsuit today against OpenAI and Microsoft, alleging the technology companies systematically copied copyrighted reporting from nearly 400 local newspapers to train and develop commercial artificial intelligence products, including […]

    https://rbfirehose.com/2026/06/25/new-jersey-globe-nearly-400-local-newspapers-sue-openai-microsoft-over-alleged-copyright-theft/
  48. New Jersey Globe: Nearly 400 local newspapers sue OpenAI, Microsoft over alleged copyright theft. “The massive coalition of local newspaper publishers filed a federal lawsuit today against OpenAI and Microsoft, alleging the technology companies systematically copied copyrighted reporting from nearly 400 local newspapers to train and develop commercial artificial intelligence products, including […]

    https://rbfirehose.com/2026/06/25/new-jersey-globe-nearly-400-local-newspapers-sue-openai-microsoft-over-alleged-copyright-theft/
  49. Search Engine Journal: US Publishers Demand Common Crawl Stop Scraping Their Content. “Digital Content Next, a trade body representing US digital publishers, has sent a cease and desist letter to the Common Crawl Foundation. The letter demands Common Crawl stop collecting publisher content and remove material already in its datasets.”

    https://rbfirehose.com/2026/06/11/search-engine-journal-us-publishers-demand-common-crawl-stop-scraping-their-content/