home.social

#web-scraping — Public Fediverse posts

Live and recent posts from across the Fediverse tagged #web-scraping, aggregated by home.social.

fetched live
  1. ☕ Wie viel weiß eine Organisation eigentlich über sich selbst?

    Beim letzten #CivicDataLab Espresso Talk scrapten ehrenamtliche Data Scientists von Data Science for Social Good (@dssgberlin Berlin) die Webauftritte des AWO Bundesverband e.V. – und fanden deutlich mehr Einrichtungen, als die eigene Datenbank vermuten ließ.

    Was hinter der Methode steckt & wie ein Team aus 15 Freiwilligen eine Datenbank mit fast 21.000 Einträgen aufgebaut hat, gibt's im Blog + Talk-Video:

    👉 civic-data.de/blog/wenn-daten-

    Code offen auf #GitHub – Nachnutzung ausdrücklich erwünscht.

    #DataForGood #Webscraping #NonProfit #Digitalisierung

  2. Stephen Follows: What just happened to TheNumbers.com should worry us all. “Its hand-researched data is the highest quality, tracking box office grosses, budgets, home video and streaming across more than 78,000 films and 236,000 people. It gets north of eight million visitors a year, and is treated as THE definitive authority by journalists, academics, filmmakers, prediction markets, and even […]

    https://rbfirehose.com/2026/07/24/stephen-follows-what-just-happened-to-thenumbers-com-should-worry-us-all/
  3. Reuters: News Corp countersues Brave for allegedly ‘scraping’ articles for AI . “News Corp, facing a lawsuit by search engine Brave Software, has filed a countersuit accusing it ‌of “flagrant theft” in distributing and selling versions of articles from the Wall Street Journal and New York Post to AI companies.”

    https://rbfirehose.com/2026/07/23/reuters-news-corp-countersues-brave-for-allegedly-scraping-articles-for-ai/
  4. Turns out you can't just ask an LLM for CSS selectors and ship them. In our scraping system, first-attempt selectors returned nothing 30 to 40% of the time. The trick that made it work: check JSON-LD first, then run every generated selector through a validation loop against the real DOM before trusting it. go.upgradejs.com/qru #LLM #WebScraping #AI

  5. MediaPost: Judge Dismisses Google Complaint Against SerpApi Over Scraping. “A federal judge has dismissed Google’s complaint against the Texas-based company SerpApi, which allegedly circumvented attempts to prevent it from scraping search results. The ruling, issued Monday by U.S. District Court Judge Yvonne Gonzalez Rogers, allows Google to amend its complaint and bring it again.”

    https://rbfirehose.com/2026/07/22/mediapost-judge-dismisses-google-complaint-against-serpapi-over-scraping/
  6. 🕷️ Anakin-Inc/anakin

    Converts websites to clean markdown or JSON with fallback scraping handlers and proxy auto-selection

    ⭐ Stars: 687
    📅 Last Update: Jul 20, 2026

    github.com/Anakin-Inc/anakin

    #selfhosted #homelab #selfhost #selfhosting #opensource #webscraping #api

  7. Mixfont: Decoy Font. “Decoy font is a font that prints a decoy for every letter, making it more difficult for AI to read what you type. The font works by using separate spatial frequencies to communicate two different letters in the same space.”

    https://rbfirehose.com/2026/07/18/mixfont-decoy-font/
  8. Engadget: Cloudflare will filter out web crawlers that serve AI companies . “Cloudflare has announced plans to automatically block mixed-use web crawlers that index websites for search engines and act as AI agents and trainers at the same time. The company previously offered its customers the optional ability to prevent crawlers from scraping their sites for AI chatbots, but now Cloudflare’s […]

    https://rbfirehose.com/2026/07/06/engadget-cloudflare-will-filter-out-web-crawlers-that-serve-ai-companies/
  9. I tested every way to scrape Amazon in 2026 — plain requests, Selenium, Playwright, free proxies, paid proxies. hackernoon.com/i-tried-every-w #webscraping

  10. New Jersey Globe: Nearly 400 local newspapers sue OpenAI, Microsoft over alleged copyright theft. “The massive coalition of local newspaper publishers filed a federal lawsuit today against OpenAI and Microsoft, alleging the technology companies systematically copied copyrighted reporting from nearly 400 local newspapers to train and develop commercial artificial intelligence products, including […]

    https://rbfirehose.com/2026/06/25/new-jersey-globe-nearly-400-local-newspapers-sue-openai-microsoft-over-alleged-copyright-theft/
  11. Search Engine Journal: US Publishers Demand Common Crawl Stop Scraping Their Content. “Digital Content Next, a trade body representing US digital publishers, has sent a cease and desist letter to the Common Crawl Foundation. The letter demands Common Crawl stop collecting publisher content and remove material already in its datasets.”

    https://rbfirehose.com/2026/06/11/search-engine-journal-us-publishers-demand-common-crawl-stop-scraping-their-content/
  12. Web scraping gets blocked by weak headers, broken sessions, poor IP reputation, fast requests, and careless proxy rotation. hackernoon.com/why-scrapers-fa #webscraping

  13. 🎉 Behold, the future of browser automation, where writing code is for peasants! 🤖 Intuned claims to solve all your web scraping woes with magic #AI that doesn't need you, unless it breaks. But don't worry, you have a whole smorgasbord of buzzwords like "Playwright" and "RPAAI" to keep you entertained while you wait. 🚀✨
    intunedhq.com #browserautomation #webscraping #Playwright #RPAAI #HackerNews #ngated