home.social

#scraping — Public Fediverse posts

Live and recent posts from across the Fediverse tagged #scraping, aggregated by home.social.

fetched live
  1. "EDSA-Leitlinien 3/2026: #Scraping bleibt möglich – die „vernünftigen Erwartungen“ werden zum zentralen Prüfstein"

    -> "Damit macht der EDSA klar, dass es seiner Ansicht nach kein generelles Verbot des Scrapings gibt"

    delegedata.de/2026/08/edsa-lei

  2. It could be just a rant by Cara with no factual basis, but for all the discussion about consent here, that would be really disappointing if true. @Mastodon @staff care to clarify if it’s true that you allow AI scrapers to access our “real-time posts and images for free”?

    @Admin how about our server?

    #AI #scraping #consent #data #privacy #transparency #generativeAI #Cara #CaraApp #Art #MastoArt #Mastodon

  3. Friends of Web #scraping, #datahoarding and @internetarchive; did you know that can transform any #WARC archive in an easily viewable #ZIM archive?

    Check other warc2zim features at github.com/openzim/warc2zim

  4. "AI companies’ penchant for scraping through large swathes of the public web in search of valuable training data has already led to lawsuits and technical fixes aimed at stopping the practice. Now, a pair of designers are hoping to stymie these scrapers with a new font designed to offer people a perfectly readable webpage while serving scrapers a subtly edited, nonsensical version in the underlying HTML."

    arstechnica.com/ai/2026/08/new

    #ai #scraping #htmlfont

  5. #ZIM tools 3.8.0 has just been released!

    This new version brings many bug fixes, in particular to ZIMcheck, our ZIM Q&A command line tool!

    More details at:
    github.com/openzim/zim-tools/r

    #FOSS #offline #edtech #scraping

  6. Go-juggler и протокол Juggler

    ИИ-агентам нужно ходить по настоящему вебу. Но настоящий веб враждебен к автоматизации: Playwright блокируют, headless-версия Chrome палится по отпечаткам, а стелс-плагины просто становятся частью отпечатка. Если вы когда-нибудь писали скраперы или инструменты для агентов, вы знаете, как это бывает - всё работает локально, а потом продакшен разваливается за Cloudflare-челленджем. Несколько проектов пытались решить эту проблему со стороны браузера. Проект camoufox.com патчит Firefox на уровне C++, так что navigator.hardwareConcurrency, WebGL-рендереры, AudioContext, геометрия экрана и WebRTC подменяются ещё до того, как JavaScript их увидит. Браузер github.com/jo-inc/camofox-brow оборачивает этот движок в REST API, заточенный под агентов: снимки доступности вместо раздутого HTML, стабильные ссылки на элементы для кликов и изоляция сессий. Остаётся только одна дыра: инструментарий вокруг этой экосистемы завязан на JavaScript/Python. Если вы живёте в Go - а весь стек Go-агентов, взорвавшийся за последние пару лет, весомый аргумент в его пользу - вам оставалось писать сырые вызовы curl. Пакет github.com/yvv4git/go-juggler исправляет это. Это Go-клиент для протокола автоматизации Juggler (того самого, который патчит и расширяет Camoufox) под лицензией MIT. Он управляет Firefox/Camoufox из Go с единственной зависимостью и чистым, слоистым API.

    habr.com/ru/articles/1068462/

    #scraping #crawling #crawler #browser #go

  7. 🕷️ germondai/trawl

    Replaces FlareSolverr with a self-hosted engine that bypasses Cloudflare and CAPTCHAs using cached browser sessions and native solving

    ⭐ Stars: 533
    📅 Last Update: Jul 30, 2026

    github.com/germondai/trawl

    #selfhosted #homelab #selfhost #selfhosting #opensource #scraping #cloudflare

  8. @82mhz @stefano

    "The work at Include Security has us working with AI day in and day out (hacking it, using it, training it, etc). .."

    Smart TVs as residential proxies: a timeline

    billboard.bsd.cafe/post/911

    #AI #proxy #scraping #VPN

  9. Your smart TV may be scraping the web for AI – Janko Roettgers | Lowpass

    <lowpass.cc/p/smart-tv-web-scra> @jank0 (February 2026)

    "… The catch? With Bright’s SDK, a viewer’s smart TV becomes part of a massive global proxy network that crawls and scrapes the web. Including apps running on desktop PCs and mobile devices, the company claims to operate 150 million such residential proxies worldwide. Together, these devices gather petabytes of public web data from a wide range of different locations and IP addresses. This approach allows the company to capture localized versions of websites, but also helps to circumvent web crawler blacklists. The gathered data is then resold to companies to train AI models, among other things. …"

    #AI #scraping #proxy

  10. We know the scurge of "AI" scraping bots. But it has become almost impossible to tell the difference between scraping bots and pentest bots. This is from a domain I purchased months ago and put something on late on the 26th. 3 humans including me know the domain and I guarantee the two others have not visited more than once (which was out of sheer politeness).

    Look at this specific session, this is just a blitz probe, none of those URLs actually exist. The sheer volume is astounding. It's a Go application sitting behind a Caddy reverse proxy (both performing admirably) which reports the usage to Umami, which is what the screenshots are from.

    #AI #AIBots #Scraping #AIisGoingGreat #GoLang #Caddy #Umami

  11. @ddosecrets Bright Data also happens to be the operator of a residential proxy network that pays developers of games for smart TVs to embed its proxy library. A story about LG banning residential proxies on webOS shows a screenshot of a consent pop-up for Web Indexing by Bright Data inside a Pac-Man game.
    krebsonsecurity.com/2026/07/lg

    #BrightData #ResidentialProxy #scraping #PacMan #webOS

  12. RE: social.edu.nl/@wlaatje/1169708

    "The web is full of independent archives, hobby databases, local news sites, forums, reference works. Decades of accumulated human effort, running on old code, maintained by small teams or single individuals, quietly holding up far more of our shared knowledge than anyone acknowledges."

    And proponents of so-called 'AI' are speedrunning their destruction.

    #noAI #scrapers #scraping #crawlers #AI #genAI

  13. #BigTech #AI #TechBro companies crying to mommy to get the bad companies to stop (maybe) using their models' output for training is now a standard play from the #monopolies and wannabe monopolies.

    If a (Chinese) company gets too good at doing what American companies do, then force an embargo or force a sell. It was the #Meta/#Facebook with #Tiktok play, and now it's the #LLM play.

    If #scraping copyrighted works for training is legal, so is #distillation.

    #Capitalists sure do hate #Capitalism.

  14. #today I shall start to shut down all the creepy user data #scraping that #Win10 does by default, just so we can use #Ableton in peace again.

    Meanwhile, praying for a breakthrough in being able to use Ableton on #Linux without pulling all my teeth out. Is there a #Distro out there to do this yet?

    #infosec