#web-scraping — Public Fediverse posts
Live and recent posts from across the Fediverse tagged #web-scraping, aggregated by home.social.
-
AI didn't invent web scraping. It removed the one thing that kept it hard: needing to know how. Here's which defenses still hold, and which ones just broke. https://hackernoon.com/scraping-used-to-take-a-programmer-now-it-takes-a-sentence #webscraping
-
Parsing HTML tables via pandas? This 147-line beginner script extracts tables from any URL to CSV—just pass --url, --table-index, --out. Ideal for stats, financials, or gov data. Try it: https://www.valtersit.com/python/html-table-to-dataframe-extractor/ #Python #WebScraping #pan
-
Search Engine Roundtable: Google: Verifying Your Request To Continue To Search Results. “In early August, Google was testing requiring searchers to sign in to verify they are humans to continue to see search results. Now, Google is testing a similar feature without requiring you to sign in, but instead, Google is asking you to click a ‘Continue’ button after a few seconds, after Google verifies […]
https://rbfirehose.com/2026/09/07/google-verifying-your-request-to-continue-to-search-results-search-engine-roundtable/ -
Investing.com: Seattle Times, Newsday sue OpenAI and Microsoft over AI training. “The lawsuit, filed in the U.S. District Court for the Southern District of New York, alleges the companies scraped content from the newspapers’ websites, including articles behind paywalls, and incorporated the material into datasets used to train and operate AI products.”
https://rbfirehose.com/2026/09/06/investing-com-seattle-times-newsday-sue-openai-and-microsoft-over-ai-training/ -
🤖 #BrowserAct is an open-source browser automation CLI built for #AIagents — real browsers, anti-bot bypass and human handoff when the agent gets stuck #opensource #webscraping #automation #AI
🧵👇 -
In 2024 the bill for the wasteful web exploded: the software archive I direct buckled under AI crawlers, and our engineers spent their days blocking addresses instead of building. We are not Google. 👉 https://www.dicosmo.org/good-enough/ #GoodEnoughIsNotGoodEnough #WebScraping #SysAdmin #AI
-
AI crawlers now burden git.kernel.org with 6M daily requests, burning more CPU rendering commits for scrapers than on all legitimate access combined.
#AICrawlers #LinuxKernel #Anubis #WebScraping #OpenSource
https://securityonline.info/ai-crawlers-git-kernel/?utm_source=mastodon&utm_medium=jetpack_social
-
"There are no easy answers to questions of scraping in the AI era. As we learned when CDT hosted a multistakeholder event on this topic at Georgetown Law in the spring, 'Internet Scraping and the Future of the Open Web in the Internet Age', all sides represent values or capabilities that society should want to preserve."
Centre for Democracy and Technology, 2026
(1/?)
-
RT @ericciarla: Wir stellen vor: Firecrawl – kostenlos und ohne API-Schlüssel. Jetzt können Ihre KI-Agenten das Web zu 100 % kostenlos durchsuchen und scannen.
mehr auf Arint.info
-
RE: https://mastodon.archive.org/@internetarchive/117168452160850013
It's not the #webScraping itself that's the problem. It's how #LLMs separate #content from its #provenance, and the #plagiarism that directly follows from that separation.
-
RE: https://mastodon.archive.org/@internetarchive/117168452160850013
"...[W]hen should individuals and organizations be able to use automated tools or “bots” to collect data from or interact with openly published websites — whether we call it “scraping”, “crawling,” or simply automated data collection? Conversely, what kind of control should websites have over when that happens, or what use the collected data is put to?"
#AI #internetArchive #IETF #webscraping #freedom #copyright #archiving
-
"Nitter, an open source project that allowed people to read X posts without logging into or even opening the X app, has received cease-and-desist letters from X demanding that it shut down. The news was shared via a brief message posted to the project’s website, and follows X’s earlier attempts to knock Nitter offline by technical means.
The service also powers a number of other sites, including XCancel, that allow people to view X posts directly.
This isn’t X’s first attempt to shut down Nitter. In 2024, Nitter’s flagship instance, Nitter.net, went dark temporarily after X rolled out new API restrictions. Nitter worked by fetching public X posts and then stripping out the ads, tracking cookies, and JavaScript, giving people a clean, clutter-free way to read posts without an account or the app.
After that crackdown, those who wanted to host a Nitter instance had to connect it to a real X account, according to the project’s GitHub page. Despite the restrictions, development picked back up and Nitter instances came back online.
This time, X is working to shut down Nitter and its instances via legal means."
-
PetaPixel: Tech Bro Scrapes Anti-AI Photo App Cara, Then Gloats About It. “Cara was started by photographer Jingna Zhang in 2023 as a direct result of big tech firms not taking action against AI bots scraping content from websites. In some cases, the platforms are scraping their users’ data for their own benefit, including Meta platforms. Zhang took to Instagram this morning to share news of […]
https://rbfirehose.com/2026/08/17/petapixel-tech-bro-scrapes-anti-ai-photo-app-cara-then-gloats-about-it/ -
San Francisco Standard: This BlackRock analyst wants to fix journalism. Step one? Skip the journalists. “In one example of The Dissent’s parasitic approach, [Dakota] Carrasco’s site regurgitated and reduced a 7,000-word investigation of a complex real-estate battle by The Standard’s Sam Mondros into a 400-word summary(opens in new tab) without mentioning the original work at all.”
https://rbfirehose.com/2026/08/16/san-francisco-standard-this-blackrock-analyst-wants-to-fix-journalism-step-one-skip-the-journalists/ -
Search Engine Journal: OpenAI Says Robots.txt May Not Apply To ChatGPT’s Fetch Bot. “ChatGPT’s page-fetching bot is disallowed by more sites than any other AI bot of its kind. It also reached disallowed pages on more sites than any other bot. OpenAI says robots.txt rules may not apply to it because a person asked for the page.”
https://rbfirehose.com/2026/08/15/search-engine-journal-openai-says-robots-txt-may-not-apply-to-chatgpts-fetch-bot/ -
15% of AI page fetchers in Europe reached disallowed URLs, TollBit finds: ChatGPT-User hit blocked pages on nearly half the European sites naming it, while 9% of those sites disallow Claude-User. Cloudflare's defaults shift Sept 15. https://ppc.land/15-of-ai-page-fetchers-in-europe-reached-disallowed-urls-tollbit-finds/ #AI #MachineLearning #ChatGPT #ClaudeAI #WebScraping
-
ShieldFont: Bludgeoning AI Scrapers that Disrespect Robots.txt
-
Fast Company: AI crawlers from Meta and Alibaba almost destroyed a volunteer-run LGBT history archive. “…the LGBT History Project… recently passed 50 million views since its launch in 2011 and has been archived by the British Library for posterity. But an onslaught of AI bots seeking to scrape its content nearly took it offline, bringing the site to a crawl while also making it more […]
https://rbfirehose.com/2026/08/14/fast-company-ai-crawlers-from-meta-and-alibaba-almost-destroyed-a-volunteer-run-lgbt-history-archive/ -
MediaPost: Google Renews Battle With SerpApi Over Scraping. “Renewing its battle with SerpApi, Google this week filed an amended complaint alleging that the Texas-based company — which provides data to other businesses — bypassed attempts to prevent it from scraping search results.”
https://rbfirehose.com/2026/08/13/mediapost-google-renews-battle-with-serpapi-over-scraping/ -
Before HTML hits the model, we strip styles, scripts, noscript, and svg tags. That alone shrinks pages 3 to 5x and buys back a lot of context window. Small, boring preprocessing beats a bigger model. https://go.upgradejs.com/gkz #WebScraping #LLM #AI
-
Learn how to debug 403 errors when scraping websites by checking headers, sessions, cookies, IP reputation, rate limits, and request patterns. https://hackernoon.com/how-to-debug-403-errors-when-scraping-websites #webscraping
-
Internet Archive Blog: Internet Archive to New York: Don’t Kill the Good Bots in the Fight Against Bad Bots. “The problem isn’t anonymous bots. The problem is excessive, harmful scraping. We need targeted solutions for abusive AI practices, while actively protecting the rights of libraries, researchers, journalists, and readers alike. That’s why the Internet Archive has joined EFF and […]
https://rbfirehose.com/2026/08/06/internet-archive-to-new-york-dont-kill-the-good-bots-in-the-fight-against-bad-bots-internet-archive/