home.social

#commoncrawl — Public Fediverse posts

Live and recent posts from across the Fediverse tagged #commoncrawl, aggregated by home.social.

fetched live
  1. Part of the trust model for AI, is users understanding that the data is a broad representation of the real world web.
    Sharp, well considered insights from the #CommonCrawl team #POSAIS26

  2. Part of the trust model for AI, is users understanding that the data is a broad representation of the real world web.
    Sharp, well considered insights from the #CommonCrawl team #POSAIS26

  3. Part of the trust model for AI, is users understanding that the data is a broad representation of the real world web.
    Sharp, well considered insights from the team

  4. Part of the trust model for AI, is users understanding that the data is a broad representation of the real world web.
    Sharp, well considered insights from the #CommonCrawl team #POSAIS26

  5. Part of the trust model for AI, is users understanding that the data is a broad representation of the real world web.
    Sharp, well considered insights from the #CommonCrawl team #POSAIS26

  6. RT @glenngabe: „Stopp das Scraping“ – US-Verlage fordern Common Crawl auf, das Scraping einzustellen und Archive zu löschen. „Sie forderten Common Crawl auf, das „Scraping, die Speicherung oder die Weitergabe von urheberrechtlich geschützten, paywalled, nur für Abonnenten zugänglichen oder anderweitig geschützten Inhalten der DCN-Mitgliedsunternehmen in ihren Datensätzen“ unverzüglich zu stoppen.“ „Sie forderten zudem, dass bereits in den Common Crawl-Datensätzen enthaltene Verlagsinhalte entfernt werden.“

    mehr auf Arint.info

    #CommonCrawl #Datenschutz #Paywall #Scraping #Urheberrecht #USVerlage #arint_info

    https://x.com/glenngabe/status/2064318799138918523#m

  7. RT @glenngabe: „Stopp das Scraping“ – US-Verlage fordern Common Crawl auf, das Scraping einzustellen und Archive zu löschen. „Sie forderten Common Crawl auf, das „Scraping, die Speicherung oder die Weitergabe von urheberrechtlich geschützten, paywalled, nur für Abonnenten zugänglichen oder anderweitig geschützten Inhalten der DCN-Mitgliedsunternehmen in ihren Datensätzen“ unverzüglich zu stoppen.“ „Sie forderten zudem, dass bereits in den Common Crawl-Datensätzen enthaltene Verlagsinhalte entfernt werden.“

    mehr auf Arint.info

    #CommonCrawl #Datenschutz #Paywall #Scraping #Urheberrecht #USVerlage #arint_info

    https://x.com/glenngabe/status/2064318799138918523#m

  8. Search Engine Journal: US Publishers Demand Common Crawl Stop Scraping Their Content. “Digital Content Next, a trade body representing US digital publishers, has sent a cease and desist letter to the Common Crawl Foundation. The letter demands Common Crawl stop collecting publisher content and remove material already in its datasets.”

    https://rbfirehose.com/2026/06/11/search-engine-journal-us-publishers-demand-common-crawl-stop-scraping-their-content/
  9. Search Engine Journal: US Publishers Demand Common Crawl Stop Scraping Their Content. “Digital Content Next, a trade body representing US digital publishers, has sent a cease and desist letter to the Common Crawl Foundation. The letter demands Common Crawl stop collecting publisher content and remove material already in its datasets.”

    https://rbfirehose.com/2026/06/11/search-engine-journal-us-publishers-demand-common-crawl-stop-scraping-their-content/
  10. Search Engine Journal: US Publishers Demand Common Crawl Stop Scraping Their Content. “Digital Content Next, a trade body representing US digital publishers, has sent a cease and desist letter to the Common Crawl Foundation. The letter demands Common Crawl stop collecting publisher content and remove material already in its datasets.”

    https://rbfirehose.com/2026/06/11/search-engine-journal-us-publishers-demand-common-crawl-stop-scraping-their-content/
  11. Search Engine Journal: US Publishers Demand Common Crawl Stop Scraping Their Content. “Digital Content Next, a trade body representing US digital publishers, has sent a cease and desist letter to the Common Crawl Foundation. The letter demands Common Crawl stop collecting publisher content and remove material already in its datasets.”

    https://rbfirehose.com/2026/06/11/search-engine-journal-us-publishers-demand-common-crawl-stop-scraping-their-content/
  12. Search Engine Journal: US Publishers Demand Common Crawl Stop Scraping Their Content. “Digital Content Next, a trade body representing US digital publishers, has sent a cease and desist letter to the Common Crawl Foundation. The letter demands Common Crawl stop collecting publisher content and remove material already in its datasets.”

    https://rbfirehose.com/2026/06/11/search-engine-journal-us-publishers-demand-common-crawl-stop-scraping-their-content/
  13. RT @glenngabe: „Stopp das Scraping“ – US-Verlage fordern Common Crawl auf, das Scraping einzustellen und Archive zu löschen. Sie forderten Common Crawl auf, das „Scraping, die Speicherung oder die Weitergabe urheberrechtlich geschützter, paywall-geschützter, nur für Abonnenten zugänglicher oder anderweitig geschützter Inhalte von DCN-Mitgliedsunternehmen in ihren Datensätzen“ unverzüglich zu stoppen. Zudem wurde die Entfernung bereits in den Common-Crawl-Datensätzen vorhandener Verlagsinhalte gefordert.

    mehr auf Arint.info

    #CommonCrawl #Datenschutz #Paywall #Scraping #Urheberrecht #USVerlage #arint_info

    https://x.com/glenngabe/status/2064318799138918523#m

  14. RT @glenngabe: „Stopp das Scraping“ – US-Verlage fordern Common Crawl auf, das Scraping einzustellen und Archive zu löschen. Sie forderten Common Crawl auf, das „Scraping, die Speicherung oder die Weitergabe urheberrechtlich geschützter, paywall-geschützter, nur für Abonnenten zugänglicher oder anderweitig geschützter Inhalte von DCN-Mitgliedsunternehmen in ihren Datensätzen“ unverzüglich zu stoppen. Zudem wurde die Entfernung bereits in den Common-Crawl-Datensätzen vorhandener Verlagsinhalte gefordert.

    mehr auf Arint.info

    #CommonCrawl #Datenschutz #Paywall #Scraping #Urheberrecht #USVerlage #arint_info

    https://x.com/glenngabe/status/2064318799138918523#m

  15. FYI: News publishers target Common Crawl, the AI training data backdoor: News/Media Alliance sent a formal letter to Common Crawl demanding it stop unauthorized scraping and block AI companies from using news content for training. ppc.land/news-publishers-targe #News #Media #CommonCrawl #AITraining #DataPrivacy

  16. FYI: News publishers target Common Crawl, the AI training data backdoor: News/Media Alliance sent a formal letter to Common Crawl demanding it stop unauthorized scraping and block AI companies from using news content for training. ppc.land/news-publishers-targe #News #Media #CommonCrawl #AITraining #DataPrivacy

  17. FYI: News publishers target Common Crawl, the AI training data backdoor: News/Media Alliance sent a formal letter to Common Crawl demanding it stop unauthorized scraping and block AI companies from using news content for training. ppc.land/news-publishers-targe #News #Media #CommonCrawl #AITraining #DataPrivacy

  18. ICYMI: News publishers target Common Crawl, the AI training data backdoor: News/Media Alliance sent a formal letter to Common Crawl demanding it stop unauthorized scraping and block AI companies from using news content for training. ppc.land/news-publishers-targe #AI #NewsMedia #CommonCrawl #DataPrivacy #WebScraping

  19. ICYMI: News publishers target Common Crawl, the AI training data backdoor: News/Media Alliance sent a formal letter to Common Crawl demanding it stop unauthorized scraping and block AI companies from using news content for training. ppc.land/news-publishers-targe #AI #NewsMedia #CommonCrawl #DataPrivacy #WebScraping

  20. ICYMI: News publishers target Common Crawl, the AI training data backdoor: News/Media Alliance sent a formal letter to Common Crawl demanding it stop unauthorized scraping and block AI companies from using news content for training. ppc.land/news-publishers-targe #AI #NewsMedia #CommonCrawl #DataPrivacy #WebScraping

  21. ICYMI: News publishers target Common Crawl, the AI training data backdoor: News/Media Alliance sent a formal letter to Common Crawl demanding it stop unauthorized scraping and block AI companies from using news content for training. ppc.land/news-publishers-targe #AI #NewsMedia #CommonCrawl #DataPrivacy #WebScraping

  22. ICYMI: News publishers target Common Crawl, the AI training data backdoor: News/Media Alliance sent a formal letter to Common Crawl demanding it stop unauthorized scraping and block AI companies from using news content for training. ppc.land/news-publishers-targe #AI #NewsMedia #CommonCrawl #DataPrivacy #WebScraping

  23. News publishers target Common Crawl, the AI training data backdoor: News/Media Alliance sent a formal letter to Common Crawl demanding it stop unauthorized scraping and block AI companies from using news content for training. ppc.land/news-publishers-targe #CommonCrawl #AITechnology #NewsPublishers #MediaAlliance #DataPrivacy

  24. News publishers target Common Crawl, the AI training data backdoor: News/Media Alliance sent a formal letter to Common Crawl demanding it stop unauthorized scraping and block AI companies from using news content for training. ppc.land/news-publishers-targe #CommonCrawl #AITechnology #NewsPublishers #MediaAlliance #DataPrivacy

  25. News publishers target Common Crawl, the AI training data backdoor: News/Media Alliance sent a formal letter to Common Crawl demanding it stop unauthorized scraping and block AI companies from using news content for training. ppc.land/news-publishers-targe #CommonCrawl #AITechnology #NewsPublishers #MediaAlliance #DataPrivacy

  26. Amazon.de hat HC Rank 201. Mediamarkt.de: 418.396 – 2.000x weiter außen im Web-Netzwerk. Bei fast identischer Backlink-Stärke.

    Common Crawl nutzt Harmonic Centrality als Crawl-Priorität. 64% aller großen Sprachmodelle trainieren auf Common-Crawl-Daten (Mozilla 2024).

    Wer nach Domain Authority optimiert, optimiert für Google. Für KI-Zitationen zählt die Position im Link-Graph.

    hechtinsgefecht.de/harmonic-ce

    #SEO #GEO #LLMSichtbarkeit #CommonCrawl

  27. Amazon.de hat HC Rank 201. Mediamarkt.de: 418.396 – 2.000x weiter außen im Web-Netzwerk. Bei fast identischer Backlink-Stärke.

    Common Crawl nutzt Harmonic Centrality als Crawl-Priorität. 64% aller großen Sprachmodelle trainieren auf Common-Crawl-Daten (Mozilla 2024).

    Wer nach Domain Authority optimiert, optimiert für Google. Für KI-Zitationen zählt die Position im Link-Graph.

    hechtinsgefecht.de/harmonic-ce

    #SEO #GEO #LLMSichtbarkeit #CommonCrawl

  28. Mediamarkt.de hat fast denselben PageRank wie Otto.de – aber einen HC Rank von 418.396 vs. 5.153. 📊

    Harmonic Centrality misst, wo eine Domain im Web-Netzwerk sitzt. Nicht wer auf dich verlinkt, sondern wie zentral du bist. Common Crawl nutzt genau diesen Wert für die Crawl-Priorität – und 64 % aller LLMs trainieren auf Common-Crawl-Daten.

    Backlink-Stärke und Netzwerkposition sind nicht dasselbe.

    hechtinsgefecht.de/harmonic-ce

    #SEO #HarmonicCentrality #GEO #KI #CommonCrawl

  29. Mediamarkt.de hat fast denselben PageRank wie Otto.de – aber einen HC Rank von 418.396 vs. 5.153. 📊

    Harmonic Centrality misst, wo eine Domain im Web-Netzwerk sitzt. Nicht wer auf dich verlinkt, sondern wie zentral du bist. Common Crawl nutzt genau diesen Wert für die Crawl-Priorität – und 64 % aller LLMs trainieren auf Common-Crawl-Daten.

    Backlink-Stärke und Netzwerkposition sind nicht dasselbe.

    hechtinsgefecht.de/harmonic-ce

    #SEO #HarmonicCentrality #GEO #KI #CommonCrawl

  30. Mediamarkt.de hat fast denselben PageRank wie Otto.de – aber einen HC Rank von 418.396 vs. 5.153. 📊

    Harmonic Centrality misst, wo eine Domain im Web-Netzwerk sitzt. Nicht wer auf dich verlinkt, sondern wie zentral du bist. Common Crawl nutzt genau diesen Wert für die Crawl-Priorität – und 64 % aller LLMs trainieren auf Common-Crawl-Daten.

    Backlink-Stärke und Netzwerkposition sind nicht dasselbe.

    hechtinsgefecht.de/harmonic-ce

    #SEO #HarmonicCentrality #GEO #KI #CommonCrawl

  31. Mediamarkt.de hat fast denselben PageRank wie Otto.de – aber einen HC Rank von 418.396 vs. 5.153. 📊

    Harmonic Centrality misst, wo eine Domain im Web-Netzwerk sitzt. Nicht wer auf dich verlinkt, sondern wie zentral du bist. Common Crawl nutzt genau diesen Wert für die Crawl-Priorität – und 64 % aller LLMs trainieren auf Common-Crawl-Daten.

    Backlink-Stärke und Netzwerkposition sind nicht dasselbe.

    hechtinsgefecht.de/harmonic-ce

    #SEO #HarmonicCentrality #GEO #KI #CommonCrawl

  32. 📸🤦‍♂️ Nathan Rooy discovers that flashy websites are like McDonald's cheeseburgers: popular for being just "good enough." Instead of a gourmet web experience, it's a buffet of #mediocrity sourced from Common Crawl's greatest hits. Web connoisseurs, prepare to feast on the bland! 🍔💻
    nry.me/posts/2025-10-09/small- #flashywebsites #webdesign #cheeseburgers #CommonCrawl #HackerNews #ngated

  33. 📸🤦‍♂️ Nathan Rooy discovers that flashy websites are like McDonald's cheeseburgers: popular for being just "good enough." Instead of a gourmet web experience, it's a buffet of #mediocrity sourced from Common Crawl's greatest hits. Web connoisseurs, prepare to feast on the bland! 🍔💻
    nry.me/posts/2025-10-09/small- #flashywebsites #webdesign #cheeseburgers #CommonCrawl #HackerNews #ngated

  34. 📸🤦‍♂️ Nathan Rooy discovers that flashy websites are like McDonald's cheeseburgers: popular for being just "good enough." Instead of a gourmet web experience, it's a buffet of #mediocrity sourced from Common Crawl's greatest hits. Web connoisseurs, prepare to feast on the bland! 🍔💻
    nry.me/posts/2025-10-09/small- #flashywebsites #webdesign #cheeseburgers #CommonCrawl #HackerNews #ngated

  35. 📸🤦‍♂️ Nathan Rooy discovers that flashy websites are like McDonald's cheeseburgers: popular for being just "good enough." Instead of a gourmet web experience, it's a buffet of #mediocrity sourced from Common Crawl's greatest hits. Web connoisseurs, prepare to feast on the bland! 🍔💻
    nry.me/posts/2025-10-09/small- #flashywebsites #webdesign #cheeseburgers #CommonCrawl #HackerNews #ngated

  36. The Company Quietly Funneling #Paywalled Articles to #AI Developers
    #CommonCrawl's website states that it scrapes the internet for "freely available content" without "going behind any '#paywall.'" Yet the organization has taken articles from major news websites that people normally have to pay for — allowing AI companies to train their #LLMs on high-quality journalism for free.
    In #2020, #OpenAI used Common Crawl’s archives to train #GPT3.
    msn.com/en-us/money/news/the-c

  37. The Company Quietly Funneling #Paywalled Articles to #AI Developers
    #CommonCrawl's website states that it scrapes the internet for "freely available content" without "going behind any '#paywall.'" Yet the organization has taken articles from major news websites that people normally have to pay for — allowing AI companies to train their #LLMs on high-quality journalism for free.
    In #2020, #OpenAI used Common Crawl’s archives to train #GPT3.
    msn.com/en-us/money/news/the-c

  38. The Company Quietly Funneling Articles to Developers
    's website states that it scrapes the internet for "freely available content" without "going behind any '#paywall.'" Yet the organization has taken articles from major news websites that people normally have to pay for — allowing AI companies to train their on high-quality journalism for free.
    In #2020, used Common Crawl’s archives to train .
    msn.com/en-us/money/news/the-c

  39. The Company Quietly Funneling #Paywalled Articles to #AI Developers
    #CommonCrawl's website states that it scrapes the internet for "freely available content" without "going behind any '#paywall.'" Yet the organization has taken articles from major news websites that people normally have to pay for — allowing AI companies to train their #LLMs on high-quality journalism for free.
    In #2020, #OpenAI used Common Crawl’s archives to train #GPT3.
    msn.com/en-us/money/news/the-c

  40. The Company Quietly Funneling #Paywalled Articles to #AI Developers
    #CommonCrawl's website states that it scrapes the internet for "freely available content" without "going behind any '#paywall.'" Yet the organization has taken articles from major news websites that people normally have to pay for — allowing AI companies to train their #LLMs on high-quality journalism for free.
    In #2020, #OpenAI used Common Crawl’s archives to train #GPT3.
    msn.com/en-us/money/news/the-c

  41. Mashable: Common Crawl accused of feeding paywalled content to AI companies. “In a detailed investigation for The Atlantic, reporter Alex Reisner reveals that several major AI companies have quietly partnered with the Common Crawl Foundation — a nonprofit that scrapes the web to build a massive public archive of the internet for research purposes.”

    https://rbfirehose.com/2025/11/09/mashable-common-crawl-accused-of-feeding-paywalled-content-to-ai-companies/

  42. Mashable: Common Crawl accused of feeding paywalled content to AI companies. “In a detailed investigation for The Atlantic, reporter Alex Reisner reveals that several major AI companies have quietly partnered with the Common Crawl Foundation — a nonprofit that scrapes the web to build a massive public archive of the internet for research purposes.”

    https://rbfirehose.com/2025/11/09/mashable-common-crawl-accused-of-feeding-paywalled-content-to-ai-companies/

  43. Mashable: Common Crawl accused of feeding paywalled content to AI companies. “In a detailed investigation for The Atlantic, reporter Alex Reisner reveals that several major AI companies have quietly partnered with the Common Crawl Foundation — a nonprofit that scrapes the web to build a massive public archive of the internet for research purposes.”

    https://rbfirehose.com/2025/11/09/mashable-common-crawl-accused-of-feeding-paywalled-content-to-ai-companies/

  44. Common Crawl - Setting the Record Straight: Common Crawl’s Commitment to Transparency, Fair Use, and the Public Good commoncrawl.org/blog/setting-t… #AI #CommonCrawl #data #WebArchiving (wow, that Atlantic piece was bad, needing this rebuttal)

  45. Common Crawl - Setting the Record Straight: Common Crawl’s Commitment to Transparency, Fair Use, and the Public Good commoncrawl.org/blog/setting-t… #AI #CommonCrawl #data #WebArchiving (wow, that Atlantic piece was bad, needing this rebuttal)