home.social

#webscraping — Public Fediverse posts

Live and recent posts from across the Fediverse tagged #webscraping, aggregated by home.social.

fetched live
  1. PetaPixel: Tech Bro Scrapes Anti-AI Photo App Cara, Then Gloats About It. “Cara was started by photographer Jingna Zhang in 2023 as a direct result of big tech firms not taking action against AI bots scraping content from websites. In some cases, the platforms are scraping their users’ data for their own benefit, including Meta platforms. Zhang took to Instagram this morning to share news of […]

    https://rbfirehose.com/2026/08/17/petapixel-tech-bro-scrapes-anti-ai-photo-app-cara-then-gloats-about-it/
  2. San Francisco Standard: This BlackRock analyst wants to fix journalism. Step one? Skip the journalists. “In one example of The Dissent’s parasitic approach, [Dakota] Carrasco’s site regurgitated and reduced a 7,000-word investigation of a complex real-estate battle by The Standard’s Sam Mondros into a 400-word summary(opens in new tab) without mentioning the original work at all.”

    https://rbfirehose.com/2026/08/16/san-francisco-standard-this-blackrock-analyst-wants-to-fix-journalism-step-one-skip-the-journalists/
  3. Search Engine Journal: OpenAI Says Robots.txt May Not Apply To ChatGPT’s Fetch Bot. “ChatGPT’s page-fetching bot is disallowed by more sites than any other AI bot of its kind. It also reached disallowed pages on more sites than any other bot. OpenAI says robots.txt rules may not apply to it because a person asked for the page.”

    https://rbfirehose.com/2026/08/15/search-engine-journal-openai-says-robots-txt-may-not-apply-to-chatgpts-fetch-bot/
  4. 15% of AI page fetchers in Europe reached disallowed URLs, TollBit finds: ChatGPT-User hit blocked pages on nearly half the European sites naming it, while 9% of those sites disallow Claude-User. Cloudflare's defaults shift Sept 15. ppc.land/15-of-ai-page-fetcher #AI #MachineLearning #ChatGPT #ClaudeAI #WebScraping

  5. Fast Company: AI crawlers from Meta and Alibaba almost destroyed a volunteer-run LGBT history archive. “…the LGBT History Project… recently passed 50 million views since its launch in 2011 and has been archived by the British Library for posterity. But an onslaught of AI bots seeking to scrape its content nearly took it offline, bringing the site to a crawl while also making it more […]

    https://rbfirehose.com/2026/08/14/fast-company-ai-crawlers-from-meta-and-alibaba-almost-destroyed-a-volunteer-run-lgbt-history-archive/
  6. MediaPost: Google Renews Battle With SerpApi Over Scraping. “Renewing its battle with SerpApi, Google this week filed an amended complaint alleging that the Texas-based company — which provides data to other businesses — bypassed attempts to prevent it from scraping search results.”

    https://rbfirehose.com/2026/08/13/mediapost-google-renews-battle-with-serpapi-over-scraping/
  7. Before HTML hits the model, we strip styles, scripts, noscript, and svg tags. That alone shrinks pages 3 to 5x and buys back a lot of context window. Small, boring preprocessing beats a bigger model. go.upgradejs.com/gkz #WebScraping #LLM #AI

  8. Learn how to debug 403 errors when scraping websites by checking headers, sessions, cookies, IP reputation, rate limits, and request patterns. hackernoon.com/how-to-debug-40 #webscraping

  9. Internet Archive Blog: Internet Archive to New York: Don’t Kill the Good Bots in the Fight Against Bad Bots. “The problem isn’t anonymous bots. The problem is excessive, harmful scraping. We need targeted solutions for abusive AI practices, while actively protecting the rights of libraries, researchers, journalists, and readers alike. That’s why the Internet Archive has joined EFF and […]

    https://rbfirehose.com/2026/08/06/internet-archive-to-new-york-dont-kill-the-good-bots-in-the-fight-against-bad-bots-internet-archive/
  10. Ars Technica: Reddit keeps its strange DMCA fight over Google search results alive. “On Friday, a judge largely denied a motion to dismiss from a web scraper, SerpApi, which is accused of conspiring with Perplexity AI to illegally scrape copyrighted Reddit content from Google search results.”

    https://rbfirehose.com/2026/08/03/ars-technica-reddit-keeps-its-strange-dmca-fight-over-google-search-results-alive/
  11. Das Allerbeste* an unseren heutigen Zeit ist nebenbei die Tatsache, dass es sich kaum noch lohnt irgendwas online frei zur Verfügung zustellen, weil dann irgendein Techmilliardär kommt und dich einfach eiskalt beklaut und deinen Datensatz tot scrapt.
    #KünstlicheIntelligenz #WebScraping #Milliadäre #KI #AI #OpenSource

    *Achtung, Zynismus: Es ist einfach nur traurig.

  12. 🕸️ us/crw

    A fast web scraper, crawler and search API with MCP server for AI agents, Firecrawl-compatible and self-hostable

    ⭐ Stars: 511
    📅 Last Update: Aug 02, 2026

    github.com/us/crw

    #selfhosted #homelab #selfhost #selfhosting #opensource #webscraping #crawling

  13. The Register: Open source project fools AI scrapers with poisoned font. “Look at a web page written using a ShieldFont font and it’ll appear exactly as one would expect: All the content words (the nouns, verbs, adjectives and adverbs that give a sentence meaning) are the same as the writer originally wrote. Inspect the raw HTML that a scraper reads from a ShieldFonted page, however, and […]

    https://rbfirehose.com/2026/07/31/the-register-open-source-project-fools-ai-scrapers-with-poisoned-font/
  14. ☕ Wie viel weiß eine Organisation eigentlich über sich selbst?

    Beim letzten #CivicDataLab Espresso Talk scrapten ehrenamtliche Data Scientists von Data Science for Social Good (@dssgberlin Berlin) die Webauftritte des AWO Bundesverband e.V. – und fanden deutlich mehr Einrichtungen, als die eigene Datenbank vermuten ließ.

    Was hinter der Methode steckt & wie ein Team aus 15 Freiwilligen eine Datenbank mit fast 21.000 Einträgen aufgebaut hat, gibt's im Blog + Talk-Video:

    👉 civic-data.de/blog/wenn-daten-

    Code offen auf #GitHub – Nachnutzung ausdrücklich erwünscht.

    #DataForGood #Webscraping #NonProfit #Digitalisierung

  15. Stephen Follows: What just happened to TheNumbers.com should worry us all. “Its hand-researched data is the highest quality, tracking box office grosses, budgets, home video and streaming across more than 78,000 films and 236,000 people. It gets north of eight million visitors a year, and is treated as THE definitive authority by journalists, academics, filmmakers, prediction markets, and even […]

    https://rbfirehose.com/2026/07/24/stephen-follows-what-just-happened-to-thenumbers-com-should-worry-us-all/
  16. Reuters: News Corp countersues Brave for allegedly ‘scraping’ articles for AI . “News Corp, facing a lawsuit by search engine Brave Software, has filed a countersuit accusing it ‌of “flagrant theft” in distributing and selling versions of articles from the Wall Street Journal and New York Post to AI companies.”

    https://rbfirehose.com/2026/07/23/reuters-news-corp-countersues-brave-for-allegedly-scraping-articles-for-ai/
  17. Turns out you can't just ask an LLM for CSS selectors and ship them. In our scraping system, first-attempt selectors returned nothing 30 to 40% of the time. The trick that made it work: check JSON-LD first, then run every generated selector through a validation loop against the real DOM before trusting it. go.upgradejs.com/qru #LLM #WebScraping #AI

  18. MediaPost: Judge Dismisses Google Complaint Against SerpApi Over Scraping. “A federal judge has dismissed Google’s complaint against the Texas-based company SerpApi, which allegedly circumvented attempts to prevent it from scraping search results. The ruling, issued Monday by U.S. District Court Judge Yvonne Gonzalez Rogers, allows Google to amend its complaint and bring it again.”

    https://rbfirehose.com/2026/07/22/mediapost-judge-dismisses-google-complaint-against-serpapi-over-scraping/
  19. 🕷️ Anakin-Inc/anakin

    Converts websites to clean markdown or JSON with fallback scraping handlers and proxy auto-selection

    ⭐ Stars: 687
    📅 Last Update: Jul 20, 2026

    github.com/Anakin-Inc/anakin

    #selfhosted #homelab #selfhost #selfhosting #opensource #webscraping #api

  20. Mixfont: Decoy Font. “Decoy font is a font that prints a decoy for every letter, making it more difficult for AI to read what you type. The font works by using separate spatial frequencies to communicate two different letters in the same space.”

    https://rbfirehose.com/2026/07/18/mixfont-decoy-font/