#commoncrawl — Public Fediverse posts
Live and recent posts from across the Fediverse tagged #commoncrawl, aggregated by home.social.
-
Part of the trust model for AI, is users understanding that the data is a broad representation of the real world web.
Sharp, well considered insights from the #CommonCrawl team #POSAIS26 -
Part of the trust model for AI, is users understanding that the data is a broad representation of the real world web.
Sharp, well considered insights from the #CommonCrawl team #POSAIS26 -
Part of the trust model for AI, is users understanding that the data is a broad representation of the real world web.
Sharp, well considered insights from the #CommonCrawl team #POSAIS26 -
Part of the trust model for AI, is users understanding that the data is a broad representation of the real world web.
Sharp, well considered insights from the #CommonCrawl team #POSAIS26 -
Part of the trust model for AI, is users understanding that the data is a broad representation of the real world web.
Sharp, well considered insights from the #CommonCrawl team #POSAIS26 -
RT @glenngabe: „Stopp das Scraping“ – US-Verlage fordern Common Crawl auf, das Scraping einzustellen und Archive zu löschen. „Sie forderten Common Crawl auf, das „Scraping, die Speicherung oder die Weitergabe von urheberrechtlich geschützten, paywalled, nur für Abonnenten zugänglichen oder anderweitig geschützten Inhalten der DCN-Mitgliedsunternehmen in ihren Datensätzen“ unverzüglich zu stoppen.“ „Sie forderten zudem, dass bereits in den Common Crawl-Datensätzen enthaltene Verlagsinhalte entfernt werden.“
mehr auf Arint.info
#CommonCrawl #Datenschutz #Paywall #Scraping #Urheberrecht #USVerlage #arint_info
-
RT @glenngabe: „Stopp das Scraping“ – US-Verlage fordern Common Crawl auf, das Scraping einzustellen und Archive zu löschen. „Sie forderten Common Crawl auf, das „Scraping, die Speicherung oder die Weitergabe von urheberrechtlich geschützten, paywalled, nur für Abonnenten zugänglichen oder anderweitig geschützten Inhalten der DCN-Mitgliedsunternehmen in ihren Datensätzen“ unverzüglich zu stoppen.“ „Sie forderten zudem, dass bereits in den Common Crawl-Datensätzen enthaltene Verlagsinhalte entfernt werden.“
mehr auf Arint.info
#CommonCrawl #Datenschutz #Paywall #Scraping #Urheberrecht #USVerlage #arint_info
-
Search Engine Journal: US Publishers Demand Common Crawl Stop Scraping Their Content. “Digital Content Next, a trade body representing US digital publishers, has sent a cease and desist letter to the Common Crawl Foundation. The letter demands Common Crawl stop collecting publisher content and remove material already in its datasets.”
https://rbfirehose.com/2026/06/11/search-engine-journal-us-publishers-demand-common-crawl-stop-scraping-their-content/ -
Search Engine Journal: US Publishers Demand Common Crawl Stop Scraping Their Content. “Digital Content Next, a trade body representing US digital publishers, has sent a cease and desist letter to the Common Crawl Foundation. The letter demands Common Crawl stop collecting publisher content and remove material already in its datasets.”
https://rbfirehose.com/2026/06/11/search-engine-journal-us-publishers-demand-common-crawl-stop-scraping-their-content/ -
Search Engine Journal: US Publishers Demand Common Crawl Stop Scraping Their Content. “Digital Content Next, a trade body representing US digital publishers, has sent a cease and desist letter to the Common Crawl Foundation. The letter demands Common Crawl stop collecting publisher content and remove material already in its datasets.”
https://rbfirehose.com/2026/06/11/search-engine-journal-us-publishers-demand-common-crawl-stop-scraping-their-content/ -
Search Engine Journal: US Publishers Demand Common Crawl Stop Scraping Their Content. “Digital Content Next, a trade body representing US digital publishers, has sent a cease and desist letter to the Common Crawl Foundation. The letter demands Common Crawl stop collecting publisher content and remove material already in its datasets.”
https://rbfirehose.com/2026/06/11/search-engine-journal-us-publishers-demand-common-crawl-stop-scraping-their-content/ -
Search Engine Journal: US Publishers Demand Common Crawl Stop Scraping Their Content. “Digital Content Next, a trade body representing US digital publishers, has sent a cease and desist letter to the Common Crawl Foundation. The letter demands Common Crawl stop collecting publisher content and remove material already in its datasets.”
https://rbfirehose.com/2026/06/11/search-engine-journal-us-publishers-demand-common-crawl-stop-scraping-their-content/ -
RT @glenngabe: „Stopp das Scraping“ – US-Verlage fordern Common Crawl auf, das Scraping einzustellen und Archive zu löschen. Sie forderten Common Crawl auf, das „Scraping, die Speicherung oder die Weitergabe urheberrechtlich geschützter, paywall-geschützter, nur für Abonnenten zugänglicher oder anderweitig geschützter Inhalte von DCN-Mitgliedsunternehmen in ihren Datensätzen“ unverzüglich zu stoppen. Zudem wurde die Entfernung bereits in den Common-Crawl-Datensätzen vorhandener Verlagsinhalte gefordert.
mehr auf Arint.info
#CommonCrawl #Datenschutz #Paywall #Scraping #Urheberrecht #USVerlage #arint_info
-
RT @glenngabe: „Stopp das Scraping“ – US-Verlage fordern Common Crawl auf, das Scraping einzustellen und Archive zu löschen. Sie forderten Common Crawl auf, das „Scraping, die Speicherung oder die Weitergabe urheberrechtlich geschützter, paywall-geschützter, nur für Abonnenten zugänglicher oder anderweitig geschützter Inhalte von DCN-Mitgliedsunternehmen in ihren Datensätzen“ unverzüglich zu stoppen. Zudem wurde die Entfernung bereits in den Common-Crawl-Datensätzen vorhandener Verlagsinhalte gefordert.
mehr auf Arint.info
#CommonCrawl #Datenschutz #Paywall #Scraping #Urheberrecht #USVerlage #arint_info
-
https://winbuzzer.com/2026/06/05/microsoft-mai-data-promise-faces-common-crawl-test-xcxwbn/
Microsoft’s in-house MAI-Thinking-1 faces scrutiny over Common Crawl and public-web training data despite its pitch about clean, commercially licensed data.
#AI #CommonCrawl #MicrosoftMAI #MAIThinking1 #AITraining #Microsoft #MicrosoftAI #AIModels #EnterpriseAI
-
https://winbuzzer.com/2026/06/05/microsoft-mai-data-promise-faces-common-crawl-test-xcxwbn/
Microsoft’s in-house MAI-Thinking-1 faces scrutiny over Common Crawl and public-web training data despite its pitch about clean, commercially licensed data.
#AI #CommonCrawl #MicrosoftMAI #MAIThinking1 #AITraining #Microsoft #MicrosoftAI #AIModels #EnterpriseAI
-
https://winbuzzer.com/2026/06/05/microsoft-mai-data-promise-faces-common-crawl-test-xcxwbn/
Microsoft’s in-house MAI-Thinking-1 faces scrutiny over Common Crawl and public-web training data despite its pitch about clean, commercially licensed data.
#AI #CommonCrawl #MicrosoftMAI #MAIThinking1 #AITraining #Microsoft #MicrosoftAI #AIModels #EnterpriseAI
-
https://winbuzzer.com/2026/06/05/microsoft-mai-data-promise-faces-common-crawl-test-xcxwbn/
Microsoft’s in-house MAI-Thinking-1 faces scrutiny over Common Crawl and public-web training data despite its pitch about clean, commercially licensed data.
#AI #CommonCrawl #MicrosoftMAI #MAIThinking1 #AITraining #Microsoft #MicrosoftAI #AIModels #EnterpriseAI
-
https://winbuzzer.com/2026/06/05/microsoft-mai-data-promise-faces-common-crawl-test-xcxwbn/
Microsoft’s in-house MAI-Thinking-1 faces scrutiny over Common Crawl and public-web training data despite its pitch about clean, commercially licensed data.
#AI #CommonCrawl #MicrosoftMAI #MAIThinking1 #AITraining #Microsoft #MicrosoftAI #AIModels #EnterpriseAI
-
-
-
-
-
-
FYI: News publishers target Common Crawl, the AI training data backdoor: News/Media Alliance sent a formal letter to Common Crawl demanding it stop unauthorized scraping and block AI companies from using news content for training. https://ppc.land/news-publishers-target-common-crawl-the-ai-training-data-backdoor/ #News #Media #CommonCrawl #AITraining #DataPrivacy
-
FYI: News publishers target Common Crawl, the AI training data backdoor: News/Media Alliance sent a formal letter to Common Crawl demanding it stop unauthorized scraping and block AI companies from using news content for training. https://ppc.land/news-publishers-target-common-crawl-the-ai-training-data-backdoor/ #News #Media #CommonCrawl #AITraining #DataPrivacy
-
FYI: News publishers target Common Crawl, the AI training data backdoor: News/Media Alliance sent a formal letter to Common Crawl demanding it stop unauthorized scraping and block AI companies from using news content for training. https://ppc.land/news-publishers-target-common-crawl-the-ai-training-data-backdoor/ #News #Media #CommonCrawl #AITraining #DataPrivacy
-
ICYMI: News publishers target Common Crawl, the AI training data backdoor: News/Media Alliance sent a formal letter to Common Crawl demanding it stop unauthorized scraping and block AI companies from using news content for training. https://ppc.land/news-publishers-target-common-crawl-the-ai-training-data-backdoor/ #AI #NewsMedia #CommonCrawl #DataPrivacy #WebScraping
-
ICYMI: News publishers target Common Crawl, the AI training data backdoor: News/Media Alliance sent a formal letter to Common Crawl demanding it stop unauthorized scraping and block AI companies from using news content for training. https://ppc.land/news-publishers-target-common-crawl-the-ai-training-data-backdoor/ #AI #NewsMedia #CommonCrawl #DataPrivacy #WebScraping
-
ICYMI: News publishers target Common Crawl, the AI training data backdoor: News/Media Alliance sent a formal letter to Common Crawl demanding it stop unauthorized scraping and block AI companies from using news content for training. https://ppc.land/news-publishers-target-common-crawl-the-ai-training-data-backdoor/ #AI #NewsMedia #CommonCrawl #DataPrivacy #WebScraping
-
ICYMI: News publishers target Common Crawl, the AI training data backdoor: News/Media Alliance sent a formal letter to Common Crawl demanding it stop unauthorized scraping and block AI companies from using news content for training. https://ppc.land/news-publishers-target-common-crawl-the-ai-training-data-backdoor/ #AI #NewsMedia #CommonCrawl #DataPrivacy #WebScraping
-
ICYMI: News publishers target Common Crawl, the AI training data backdoor: News/Media Alliance sent a formal letter to Common Crawl demanding it stop unauthorized scraping and block AI companies from using news content for training. https://ppc.land/news-publishers-target-common-crawl-the-ai-training-data-backdoor/ #AI #NewsMedia #CommonCrawl #DataPrivacy #WebScraping
-
News publishers target Common Crawl, the AI training data backdoor: News/Media Alliance sent a formal letter to Common Crawl demanding it stop unauthorized scraping and block AI companies from using news content for training. https://ppc.land/news-publishers-target-common-crawl-the-ai-training-data-backdoor/ #CommonCrawl #AITechnology #NewsPublishers #MediaAlliance #DataPrivacy
-
News publishers target Common Crawl, the AI training data backdoor: News/Media Alliance sent a formal letter to Common Crawl demanding it stop unauthorized scraping and block AI companies from using news content for training. https://ppc.land/news-publishers-target-common-crawl-the-ai-training-data-backdoor/ #CommonCrawl #AITechnology #NewsPublishers #MediaAlliance #DataPrivacy
-
News publishers target Common Crawl, the AI training data backdoor: News/Media Alliance sent a formal letter to Common Crawl demanding it stop unauthorized scraping and block AI companies from using news content for training. https://ppc.land/news-publishers-target-common-crawl-the-ai-training-data-backdoor/ #CommonCrawl #AITechnology #NewsPublishers #MediaAlliance #DataPrivacy
-
Amazon.de hat HC Rank 201. Mediamarkt.de: 418.396 – 2.000x weiter außen im Web-Netzwerk. Bei fast identischer Backlink-Stärke.
Common Crawl nutzt Harmonic Centrality als Crawl-Priorität. 64% aller großen Sprachmodelle trainieren auf Common-Crawl-Daten (Mozilla 2024).
Wer nach Domain Authority optimiert, optimiert für Google. Für KI-Zitationen zählt die Position im Link-Graph.
-
Amazon.de hat HC Rank 201. Mediamarkt.de: 418.396 – 2.000x weiter außen im Web-Netzwerk. Bei fast identischer Backlink-Stärke.
Common Crawl nutzt Harmonic Centrality als Crawl-Priorität. 64% aller großen Sprachmodelle trainieren auf Common-Crawl-Daten (Mozilla 2024).
Wer nach Domain Authority optimiert, optimiert für Google. Für KI-Zitationen zählt die Position im Link-Graph.
-
Mediamarkt.de hat fast denselben PageRank wie Otto.de – aber einen HC Rank von 418.396 vs. 5.153. 📊
Harmonic Centrality misst, wo eine Domain im Web-Netzwerk sitzt. Nicht wer auf dich verlinkt, sondern wie zentral du bist. Common Crawl nutzt genau diesen Wert für die Crawl-Priorität – und 64 % aller LLMs trainieren auf Common-Crawl-Daten.
Backlink-Stärke und Netzwerkposition sind nicht dasselbe.
-
Mediamarkt.de hat fast denselben PageRank wie Otto.de – aber einen HC Rank von 418.396 vs. 5.153. 📊
Harmonic Centrality misst, wo eine Domain im Web-Netzwerk sitzt. Nicht wer auf dich verlinkt, sondern wie zentral du bist. Common Crawl nutzt genau diesen Wert für die Crawl-Priorität – und 64 % aller LLMs trainieren auf Common-Crawl-Daten.
Backlink-Stärke und Netzwerkposition sind nicht dasselbe.
-
Mediamarkt.de hat fast denselben PageRank wie Otto.de – aber einen HC Rank von 418.396 vs. 5.153. 📊
Harmonic Centrality misst, wo eine Domain im Web-Netzwerk sitzt. Nicht wer auf dich verlinkt, sondern wie zentral du bist. Common Crawl nutzt genau diesen Wert für die Crawl-Priorität – und 64 % aller LLMs trainieren auf Common-Crawl-Daten.
Backlink-Stärke und Netzwerkposition sind nicht dasselbe.
-
Mediamarkt.de hat fast denselben PageRank wie Otto.de – aber einen HC Rank von 418.396 vs. 5.153. 📊
Harmonic Centrality misst, wo eine Domain im Web-Netzwerk sitzt. Nicht wer auf dich verlinkt, sondern wie zentral du bist. Common Crawl nutzt genau diesen Wert für die Crawl-Priorität – und 64 % aller LLMs trainieren auf Common-Crawl-Daten.
Backlink-Stärke und Netzwerkposition sind nicht dasselbe.
-
https://winbuzzer.com/2026/02/15/publishers-block-internet-archive-ai-scraping-fears-xcxwbn/
Publishers Block Internet Archive Over AI Scraping Fears
#AI #WaybackMachine #InternetArchive #Google #Reddit #OpenAI #BigTech #TheNewYorkTimes #NewsPublishers #AIScraping #OpenWeb #CommonCrawl #PerplexityAI #Media
-
https://winbuzzer.com/2026/02/15/publishers-block-internet-archive-ai-scraping-fears-xcxwbn/
Publishers Block Internet Archive Over AI Scraping Fears
#AI #WaybackMachine #InternetArchive #Google #Reddit #OpenAI #BigTech #TheNewYorkTimes #NewsPublishers #AIScraping #OpenWeb #CommonCrawl #PerplexityAI #Media
-
https://winbuzzer.com/2026/02/15/publishers-block-internet-archive-ai-scraping-fears-xcxwbn/
Publishers Block Internet Archive Over AI Scraping Fears
#AI #WaybackMachine #InternetArchive #Google #Reddit #OpenAI #BigTech #TheNewYorkTimes #NewsPublishers #AIScraping #OpenWeb #CommonCrawl #PerplexityAI #Media
-
https://winbuzzer.com/2026/02/15/publishers-block-internet-archive-ai-scraping-fears-xcxwbn/
Publishers Block Internet Archive Over AI Scraping Fears
#AI #WaybackMachine #InternetArchive #Google #Reddit #OpenAI #BigTech #TheNewYorkTimes #NewsPublishers #AIScraping #OpenWeb #CommonCrawl #PerplexityAI #Media
-
https://winbuzzer.com/2026/02/15/publishers-block-internet-archive-ai-scraping-fears-xcxwbn/
Publishers Block Internet Archive Over AI Scraping Fears
#AI #WaybackMachine #InternetArchive #Google #Reddit #OpenAI #BigTech #TheNewYorkTimes #NewsPublishers #AIScraping #OpenWeb #CommonCrawl #PerplexityAI #Media
-
📸🤦♂️ Nathan Rooy discovers that flashy websites are like McDonald's cheeseburgers: popular for being just "good enough." Instead of a gourmet web experience, it's a buffet of #mediocrity sourced from Common Crawl's greatest hits. Web connoisseurs, prepare to feast on the bland! 🍔💻
https://nry.me/posts/2025-10-09/small-web-screenshots/ #flashywebsites #webdesign #cheeseburgers #CommonCrawl #HackerNews #ngated -
📸🤦♂️ Nathan Rooy discovers that flashy websites are like McDonald's cheeseburgers: popular for being just "good enough." Instead of a gourmet web experience, it's a buffet of #mediocrity sourced from Common Crawl's greatest hits. Web connoisseurs, prepare to feast on the bland! 🍔💻
https://nry.me/posts/2025-10-09/small-web-screenshots/ #flashywebsites #webdesign #cheeseburgers #CommonCrawl #HackerNews #ngated -
📸🤦♂️ Nathan Rooy discovers that flashy websites are like McDonald's cheeseburgers: popular for being just "good enough." Instead of a gourmet web experience, it's a buffet of #mediocrity sourced from Common Crawl's greatest hits. Web connoisseurs, prepare to feast on the bland! 🍔💻
https://nry.me/posts/2025-10-09/small-web-screenshots/ #flashywebsites #webdesign #cheeseburgers #CommonCrawl #HackerNews #ngated -
📸🤦♂️ Nathan Rooy discovers that flashy websites are like McDonald's cheeseburgers: popular for being just "good enough." Instead of a gourmet web experience, it's a buffet of #mediocrity sourced from Common Crawl's greatest hits. Web connoisseurs, prepare to feast on the bland! 🍔💻
https://nry.me/posts/2025-10-09/small-web-screenshots/ #flashywebsites #webdesign #cheeseburgers #CommonCrawl #HackerNews #ngated -
The Company Quietly Funneling #Paywalled Articles to #AI Developers
#CommonCrawl's website states that it scrapes the internet for "freely available content" without "going behind any '#paywall.'" Yet the organization has taken articles from major news websites that people normally have to pay for — allowing AI companies to train their #LLMs on high-quality journalism for free.
In #2020, #OpenAI used Common Crawl’s archives to train #GPT3.
https://www.msn.com/en-us/money/news/the-company-quietly-funneling-paywalled-articles-to-ai-developers/ar-AA1PMBHE -
The Company Quietly Funneling #Paywalled Articles to #AI Developers
#CommonCrawl's website states that it scrapes the internet for "freely available content" without "going behind any '#paywall.'" Yet the organization has taken articles from major news websites that people normally have to pay for — allowing AI companies to train their #LLMs on high-quality journalism for free.
In #2020, #OpenAI used Common Crawl’s archives to train #GPT3.
https://www.msn.com/en-us/money/news/the-company-quietly-funneling-paywalled-articles-to-ai-developers/ar-AA1PMBHE -
The Company Quietly Funneling #Paywalled Articles to #AI Developers
#CommonCrawl's website states that it scrapes the internet for "freely available content" without "going behind any '#paywall.'" Yet the organization has taken articles from major news websites that people normally have to pay for — allowing AI companies to train their #LLMs on high-quality journalism for free.
In #2020, #OpenAI used Common Crawl’s archives to train #GPT3.
https://www.msn.com/en-us/money/news/the-company-quietly-funneling-paywalled-articles-to-ai-developers/ar-AA1PMBHE -
The Company Quietly Funneling #Paywalled Articles to #AI Developers
#CommonCrawl's website states that it scrapes the internet for "freely available content" without "going behind any '#paywall.'" Yet the organization has taken articles from major news websites that people normally have to pay for — allowing AI companies to train their #LLMs on high-quality journalism for free.
In #2020, #OpenAI used Common Crawl’s archives to train #GPT3.
https://www.msn.com/en-us/money/news/the-company-quietly-funneling-paywalled-articles-to-ai-developers/ar-AA1PMBHE -
The Company Quietly Funneling #Paywalled Articles to #AI Developers
#CommonCrawl's website states that it scrapes the internet for "freely available content" without "going behind any '#paywall.'" Yet the organization has taken articles from major news websites that people normally have to pay for — allowing AI companies to train their #LLMs on high-quality journalism for free.
In #2020, #OpenAI used Common Crawl’s archives to train #GPT3.
https://www.msn.com/en-us/money/news/the-company-quietly-funneling-paywalled-articles-to-ai-developers/ar-AA1PMBHE -
Mashable: Common Crawl accused of feeding paywalled content to AI companies. “In a detailed investigation for The Atlantic, reporter Alex Reisner reveals that several major AI companies have quietly partnered with the Common Crawl Foundation — a nonprofit that scrapes the web to build a massive public archive of the internet for research purposes.”
-
Mashable: Common Crawl accused of feeding paywalled content to AI companies. “In a detailed investigation for The Atlantic, reporter Alex Reisner reveals that several major AI companies have quietly partnered with the Common Crawl Foundation — a nonprofit that scrapes the web to build a massive public archive of the internet for research purposes.”
-
Mashable: Common Crawl accused of feeding paywalled content to AI companies. “In a detailed investigation for The Atlantic, reporter Alex Reisner reveals that several major AI companies have quietly partnered with the Common Crawl Foundation — a nonprofit that scrapes the web to build a massive public archive of the internet for research purposes.”
-
Common Crawl - Setting the Record Straight: Common Crawl’s Commitment to Transparency, Fair Use, and the Public Good commoncrawl.org/blog/setting-t… #AI #CommonCrawl #data #WebArchiving (wow, that Atlantic piece was bad, needing this rebuttal)
-
Common Crawl - Setting the Record Straight: Common Crawl’s Commitment to Transparency, Fair Use, and the Public Good commoncrawl.org/blog/setting-t… #AI #CommonCrawl #data #WebArchiving (wow, that Atlantic piece was bad, needing this rebuttal)