home.social

#web-scraping — Public Fediverse posts

Live and recent posts from across the Fediverse tagged #web-scraping, aggregated by home.social.

fetched live
  1. Stephen Follows: What just happened to TheNumbers.com should worry us all. “Its hand-researched data is the highest quality, tracking box office grosses, budgets, home video and streaming across more than 78,000 films and 236,000 people. It gets north of eight million visitors a year, and is treated as THE definitive authority by journalists, academics, filmmakers, prediction markets, and even […]

    https://rbfirehose.com/2026/07/24/stephen-follows-what-just-happened-to-thenumbers-com-should-worry-us-all/
  2. Stephen Follows: What just happened to TheNumbers.com should worry us all. “Its hand-researched data is the highest quality, tracking box office grosses, budgets, home video and streaming across more than 78,000 films and 236,000 people. It gets north of eight million visitors a year, and is treated as THE definitive authority by journalists, academics, filmmakers, prediction markets, and even […]

    https://rbfirehose.com/2026/07/24/stephen-follows-what-just-happened-to-thenumbers-com-should-worry-us-all/
  3. Reuters: News Corp countersues Brave for allegedly ‘scraping’ articles for AI . “News Corp, facing a lawsuit by search engine Brave Software, has filed a countersuit accusing it ‌of “flagrant theft” in distributing and selling versions of articles from the Wall Street Journal and New York Post to AI companies.”

    https://rbfirehose.com/2026/07/23/reuters-news-corp-countersues-brave-for-allegedly-scraping-articles-for-ai/
  4. Reuters: News Corp countersues Brave for allegedly ‘scraping’ articles for AI . “News Corp, facing a lawsuit by search engine Brave Software, has filed a countersuit accusing it ‌of “flagrant theft” in distributing and selling versions of articles from the Wall Street Journal and New York Post to AI companies.”

    https://rbfirehose.com/2026/07/23/reuters-news-corp-countersues-brave-for-allegedly-scraping-articles-for-ai/
  5. Turns out you can't just ask an LLM for CSS selectors and ship them. In our scraping system, first-attempt selectors returned nothing 30 to 40% of the time. The trick that made it work: check JSON-LD first, then run every generated selector through a validation loop against the real DOM before trusting it. go.upgradejs.com/qru #LLM #WebScraping #AI

  6. Turns out you can't just ask an LLM for CSS selectors and ship them. In our scraping system, first-attempt selectors returned nothing 30 to 40% of the time. The trick that made it work: check JSON-LD first, then run every generated selector through a validation loop against the real DOM before trusting it. go.upgradejs.com/qru #LLM #WebScraping #AI

  7. MediaPost: Judge Dismisses Google Complaint Against SerpApi Over Scraping. “A federal judge has dismissed Google’s complaint against the Texas-based company SerpApi, which allegedly circumvented attempts to prevent it from scraping search results. The ruling, issued Monday by U.S. District Court Judge Yvonne Gonzalez Rogers, allows Google to amend its complaint and bring it again.”

    https://rbfirehose.com/2026/07/22/mediapost-judge-dismisses-google-complaint-against-serpapi-over-scraping/
  8. MediaPost: Judge Dismisses Google Complaint Against SerpApi Over Scraping. “A federal judge has dismissed Google’s complaint against the Texas-based company SerpApi, which allegedly circumvented attempts to prevent it from scraping search results. The ruling, issued Monday by U.S. District Court Judge Yvonne Gonzalez Rogers, allows Google to amend its complaint and bring it again.”

    https://rbfirehose.com/2026/07/22/mediapost-judge-dismisses-google-complaint-against-serpapi-over-scraping/
  9. 🕷️ Anakin-Inc/anakin

    Converts websites to clean markdown or JSON with fallback scraping handlers and proxy auto-selection

    ⭐ Stars: 687
    📅 Last Update: Jul 20, 2026

    github.com/Anakin-Inc/anakin

    #selfhosted #homelab #selfhost #selfhosting #opensource #webscraping #api

  10. Mixfont: Decoy Font. “Decoy font is a font that prints a decoy for every letter, making it more difficult for AI to read what you type. The font works by using separate spatial frequencies to communicate two different letters in the same space.”

    https://rbfirehose.com/2026/07/18/mixfont-decoy-font/
  11. Mixfont: Decoy Font. “Decoy font is a font that prints a decoy for every letter, making it more difficult for AI to read what you type. The font works by using separate spatial frequencies to communicate two different letters in the same space.”

    https://rbfirehose.com/2026/07/18/mixfont-decoy-font/
  12. Engadget: Cloudflare will filter out web crawlers that serve AI companies . “Cloudflare has announced plans to automatically block mixed-use web crawlers that index websites for search engines and act as AI agents and trainers at the same time. The company previously offered its customers the optional ability to prevent crawlers from scraping their sites for AI chatbots, but now Cloudflare’s […]

    https://rbfirehose.com/2026/07/06/engadget-cloudflare-will-filter-out-web-crawlers-that-serve-ai-companies/
  13. Engadget: Cloudflare will filter out web crawlers that serve AI companies . “Cloudflare has announced plans to automatically block mixed-use web crawlers that index websites for search engines and act as AI agents and trainers at the same time. The company previously offered its customers the optional ability to prevent crawlers from scraping their sites for AI chatbots, but now Cloudflare’s […]

    https://rbfirehose.com/2026/07/06/engadget-cloudflare-will-filter-out-web-crawlers-that-serve-ai-companies/
  14. RT @NousResearch: Der Hermes-Agent liest die Webinhalte nun bis zu 60-mal schneller und zu 49-mal niedrigeren Kosten. Scraping-Backends übergeben saubere Inhalte direkt an den Agenten, ohne redundante Verarbeitungsschritte; große Seiten werden lokal gespeichert und bei Bedarf seitenweise abgerufen, sodass Sie die gleiche Qualität zu einem Bruchteil der Zeit und Kosten erhalten. Video

    mehr auf Arint.info

    #Effizienz #HermesAgent #Kostensenkung #Video #WebScraping #arint_info

    https://x.com/NousResearch/status/2071974594961977727#m

  15. I tested every way to scrape Amazon in 2026 — plain requests, Selenium, Playwright, free proxies, paid proxies. hackernoon.com/i-tried-every-w #webscraping

  16. I tested every way to scrape Amazon in 2026 — plain requests, Selenium, Playwright, free proxies, paid proxies. hackernoon.com/i-tried-every-w #webscraping

  17. RT @NousResearch: Der Hermes-Agent liest das Web nun bis zu 60-mal schneller und 49-mal günstiger. Scraping-Backends übergeben saubere Inhalte direkt an den Agenten ohne redundante Verarbeitungsschritte; große Seiten werden lokal gespeichert und bei Bedarf aufgerufen, sodass Sie die gleiche Qualität zu einem Bruchteil der Zeit und Kosten erhalten. Video

    mehr auf Arint.info

    #HermesAgent #Kosteneffizienz #Performance #Technologie #WebScraping #arint_info

    https://x.com/NousResearch/status/2071974594961977727#m

  18. New Jersey Globe: Nearly 400 local newspapers sue OpenAI, Microsoft over alleged copyright theft. “The massive coalition of local newspaper publishers filed a federal lawsuit today against OpenAI and Microsoft, alleging the technology companies systematically copied copyrighted reporting from nearly 400 local newspapers to train and develop commercial artificial intelligence products, including […]

    https://rbfirehose.com/2026/06/25/new-jersey-globe-nearly-400-local-newspapers-sue-openai-microsoft-over-alleged-copyright-theft/
  19. New Jersey Globe: Nearly 400 local newspapers sue OpenAI, Microsoft over alleged copyright theft. “The massive coalition of local newspaper publishers filed a federal lawsuit today against OpenAI and Microsoft, alleging the technology companies systematically copied copyrighted reporting from nearly 400 local newspapers to train and develop commercial artificial intelligence products, including […]

    https://rbfirehose.com/2026/06/25/new-jersey-globe-nearly-400-local-newspapers-sue-openai-microsoft-over-alleged-copyright-theft/
  20. Search Engine Journal: US Publishers Demand Common Crawl Stop Scraping Their Content. “Digital Content Next, a trade body representing US digital publishers, has sent a cease and desist letter to the Common Crawl Foundation. The letter demands Common Crawl stop collecting publisher content and remove material already in its datasets.”

    https://rbfirehose.com/2026/06/11/search-engine-journal-us-publishers-demand-common-crawl-stop-scraping-their-content/
  21. Search Engine Journal: US Publishers Demand Common Crawl Stop Scraping Their Content. “Digital Content Next, a trade body representing US digital publishers, has sent a cease and desist letter to the Common Crawl Foundation. The letter demands Common Crawl stop collecting publisher content and remove material already in its datasets.”

    https://rbfirehose.com/2026/06/11/search-engine-journal-us-publishers-demand-common-crawl-stop-scraping-their-content/
  22. Web scraping gets blocked by weak headers, broken sessions, poor IP reputation, fast requests, and careless proxy rotation. hackernoon.com/why-scrapers-fa #webscraping

  23. Web scraping gets blocked by weak headers, broken sessions, poor IP reputation, fast requests, and careless proxy rotation. hackernoon.com/why-scrapers-fa #webscraping

  24. 🎉 Behold, the future of browser automation, where writing code is for peasants! 🤖 Intuned claims to solve all your web scraping woes with magic #AI that doesn't need you, unless it breaks. But don't worry, you have a whole smorgasbord of buzzwords like "Playwright" and "RPAAI" to keep you entertained while you wait. 🚀✨
    intunedhq.com #browserautomation #webscraping #Playwright #RPAAI #HackerNews #ngated

  25. 🎉 Behold, the future of browser automation, where writing code is for peasants! 🤖 Intuned claims to solve all your web scraping woes with magic #AI that doesn't need you, unless it breaks. But don't worry, you have a whole smorgasbord of buzzwords like "Playwright" and "RPAAI" to keep you entertained while you wait. 🚀✨
    intunedhq.com #browserautomation #webscraping #Playwright #RPAAI #HackerNews #ngated

  26. Released hydrascrape — a free, self-healing scraper for sites behind Imperva/Distil bot walls.

    patched-Chrome (patchright) clears the JS challenge; a fleet of Tor exits beats per-IP limits and auto-rotates the bad ones. Resumable, with a live dashboard. Proven on a ~58k-page catalog from one laptop, $0.

    Full "what I tried and why" writeup included. MIT.

    🔗 github.com/philipposk/hydrascrape
    #OpenSource #WebScraping #Tor #Python #SelfHosting

  27. Я сошёл с ума и сдаю свой браузер ИИ-агентам

    Я совсем поехал кукухой — начал сдавать в аренду свой браузер за деньги. Началось всё с того, что мои ИИ-агенты не смогли нормально зарегаться из-за капчей и прочего, чужие расширения меня не устраивали — они плохо интегрировались в мой флоу и были завязаны на провайдера, что полный отстой. В итоге я интегрировал это в свой пет-проект, и в итоге сделал так, что браузер в аренду может взять любой желающий. Заодно сделал SDK, CLI и доки. Вот моя история погружения в пучину безумия. Погрузиться в пучину.

    habr.com/ru/articles/1041768/

    #aiагенты #browserautomation #chromeрасширения #mcp #петпроект #mcpserver #llm #webscraping #антидетект #криптоплатежи