home.social

#contentscraping — Public Fediverse posts

Live and recent posts from across the Fediverse tagged #contentscraping, aggregated by home.social.

fetched live
  1. alojapan.com/1356281/two-major Two major Japanese publishers have sued Perplexity AI, here’s why #AICopyrightInfringement #ContentScraping #GenerativeAILegalIssues #Japan #JapanNews #Japanese #JapaneseNews #JournalismAndAI #news #NikkeiAndAsahiShimbun #PerplexityAI #PerplexityAILawsuit Two of Japan’s most prominent newspapers, Nikkei and the Asahi Shimbun, have filed a lawsuit against artificial intelligence (AI) startup Perplexity AI. The legal action, announced in a join

  2. alojapan.com/1356281/two-major Two major Japanese publishers have sued Perplexity AI, here’s why #AICopyrightInfringement #ContentScraping #GenerativeAILegalIssues #Japan #JapanNews #Japanese #JapaneseNews #JournalismAndAI #news #NikkeiAndAsahiShimbun #PerplexityAI #PerplexityAILawsuit Two of Japan’s most prominent newspapers, Nikkei and the Asahi Shimbun, have filed a lawsuit against artificial intelligence (AI) startup Perplexity AI. The legal action, announced in a join

  3. Podcast: What the First AI Copyright Ruling Means for Authors

    ALLi Director Orna Ross unpacks the first court ruling on AI’s use of copyrighted books and why it matters to indie authors. She explains how the judge balanced “fair use” with pirate copying, then walks through ALLi’s Four Cs—consent, compensation, clarity, and…
    selfpublishingadvice.org/podca

    #Podcast #AIcopyrightruling #authorrights #contentscraping #creativepublishing
    @indieauthors

  4. Podcast: What the First AI Copyright Ruling Means for Authors

    ALLi Director Orna Ross unpacks the first court ruling on AI’s use of copyrighted books and why it matters to indie authors. She explains how the judge balanced “fair use” with pirate copying, then walks through ALLi’s Four Cs—consent, compensation, clarity, and…
    selfpublishingadvice.org/podca

    #Podcast #AIcopyrightruling #authorrights #contentscraping #creativepublishing
    @indieauthors

  5. Podcast: What the First AI Copyright Ruling Means for Authors

    ALLi Director Orna Ross unpacks the first court ruling on AI’s use of copyrighted books and why it matters to indie authors. She explains how the judge balanced “fair use” with pirate copying, then walks through ALLi’s Four Cs—consent, compensation, clarity, and…
    selfpublishingadvice.org/podca

    #Podcast #AIcopyrightruling #authorrights #contentscraping #creativepublishing
    @indieauthors

  6. Podcast: What the First AI Copyright Ruling Means for Authors

    ALLi Director Orna Ross unpacks the first court ruling on AI’s use of copyrighted books and why it matters to indie authors. She explains how the judge balanced “fair use” with pirate copying, then walks through ALLi’s Four Cs—consent, compensation, clarity, and…
    selfpublishingadvice.org/podca

    #Podcast #AIcopyrightruling #authorrights #contentscraping #creativepublishing
    @indieauthors

  7. This is sparking interesting discussions. wired.com/story/perplexity-is- Hallucinations and “bullshitting” are definitely an AI thing, I’d probably say it’s their best feature… But Wired’s article focuses on an important topic, scraping content without permission, ignoring robot.txt among other things. Perplexity is not the first and won’t be the last doing this and it definitely causes harm to publishers. The question is: “Why is this happening”? It’s not just because AIs need more accurate sources (instead of making stuff up), but imho it’s because finding the right content has become increasingly challenging, search engines are dominated by SEO practices and search results are disappointing at best. News sites, obviously in need of getting some revenues, are paywalling everything. In many fora, sites like archive.is and unpaywall extensions are often praised under the “free the information” slogan, RSS feeds kind of play a role there too, because in many cases they don’t drive people to visit the original websites. I think this is not much different than what AIs are doing now, and I’m not saying this is legal or ethical, it’s just a fact.
    So, my question is: is the ball on the court of AIs, needing to be regulated, or is it on the publishers’, to identify other ways of getting revenues out of this?
    #AI #Hallucinations #ContentScraping #SEO #Paywalls #AIRegulation #News #Publishers #Ethics #searchengine #perplexity #chatgpt

  8. This is sparking interesting discussions. wired.com/story/perplexity-is- Hallucinations and “bullshitting” are definitely an AI thing, I’d probably say it’s their best feature… But Wired’s article focuses on an important topic, scraping content without permission, ignoring robot.txt among other things. Perplexity is not the first and won’t be the last doing this and it definitely causes harm to publishers. The question is: “Why is this happening”? It’s not just because AIs need more accurate sources (instead of making stuff up), but imho it’s because finding the right content has become increasingly challenging, search engines are dominated by SEO practices and search results are disappointing at best. News sites, obviously in need of getting some revenues, are paywalling everything. In many fora, sites like archive.is and unpaywall extensions are often praised under the “free the information” slogan, RSS feeds kind of play a role there too, because in many cases they don’t drive people to visit the original websites. I think this is not much different than what AIs are doing now, and I’m not saying this is legal or ethical, it’s just a fact.
    So, my question is: is the ball on the court of AIs, needing to be regulated, or is it on the publishers’, to identify other ways of getting revenues out of this?
    #AI #Hallucinations #ContentScraping #SEO #Paywalls #AIRegulation #News #Publishers #Ethics #searchengine #perplexity #chatgpt

  9. This is sparking interesting discussions. wired.com/story/perplexity-is- Hallucinations and “bullshitting” are definitely an AI thing, I’d probably say it’s their best feature… But Wired’s article focuses on an important topic, scraping content without permission, ignoring robot.txt among other things. Perplexity is not the first and won’t be the last doing this and it definitely causes harm to publishers. The question is: “Why is this happening”? It’s not just because AIs need more accurate sources (instead of making stuff up), but imho it’s because finding the right content has become increasingly challenging, search engines are dominated by SEO practices and search results are disappointing at best. News sites, obviously in need of getting some revenues, are paywalling everything. In many fora, sites like archive.is and unpaywall extensions are often praised under the “free the information” slogan, RSS feeds kind of play a role there too, because in many cases they don’t drive people to visit the original websites. I think this is not much different than what AIs are doing now, and I’m not saying this is legal or ethical, it’s just a fact.
    So, my question is: is the ball on the court of AIs, needing to be regulated, or is it on the publishers’, to identify other ways of getting revenues out of this?
    #AI #Hallucinations #ContentScraping #SEO #Paywalls #AIRegulation #News #Publishers #Ethics #searchengine #perplexity #chatgpt

  10. This is sparking interesting discussions. wired.com/story/perplexity-is- Hallucinations and “bullshitting” are definitely an AI thing, I’d probably say it’s their best feature… But Wired’s article focuses on an important topic, scraping content without permission, ignoring robot.txt among other things. Perplexity is not the first and won’t be the last doing this and it definitely causes harm to publishers. The question is: “Why is this happening”? It’s not just because AIs need more accurate sources (instead of making stuff up), but imho it’s because finding the right content has become increasingly challenging, search engines are dominated by SEO practices and search results are disappointing at best. News sites, obviously in need of getting some revenues, are paywalling everything. In many fora, sites like archive.is and unpaywall extensions are often praised under the “free the information” slogan, RSS feeds kind of play a role there too, because in many cases they don’t drive people to visit the original websites. I think this is not much different than what AIs are doing now, and I’m not saying this is legal or ethical, it’s just a fact.
    So, my question is: is the ball on the court of AIs, needing to be regulated, or is it on the publishers’, to identify other ways of getting revenues out of this?
    #AI #Hallucinations #ContentScraping #SEO #Paywalls #AIRegulation #News #Publishers #Ethics #searchengine #perplexity #chatgpt

  11. 😤 #Scraperbots are automating data theft, extracting your website's content without permission! 🌐

    💣 Learn about the impact of scraper bots and how to prevent them: bit.ly/3RiXgya

    #contentscraping #bots #webscrapers #webcrawlers #scraping #waf #botmanagement #waap #scrapingbots #apptrana #indusface

  12. 😤 #Scraperbots are automating data theft, extracting your website's content without permission! 🌐

    💣 Learn about the impact of scraper bots and how to prevent them: bit.ly/3RiXgya

    #contentscraping #bots #webscrapers #webcrawlers #scraping #waf #botmanagement #waap #scrapingbots #apptrana #indusface

  13. @jackwilliambell

    I've had newsmast.social domain-blocked for a good while

    Last night I saw a rather personal, "sorry I've been gone for so long here's why" post that someone made that -- hashtags used aside -- was clearly not public

    Newmast scraped it because of the hashtag and broadcast it out to all subscribers to that hashtag

    I have little doubt that the person making the original post was aware of any of this or would have wanted their post broadcast by a content scraper

    And note that on their web site they claim to be run by a "charitable organization" and to request donations

    "Help us to remain ad-free and non-profit, whilst amplifying impactful, unheard voices on matters of global interest"

    Here: newsmastfoundation.org/donate/

    NOTE: Firefox throws a security warning; I continued because I've been on that web site before...

    cc @seb This (Newsmast) and its ilk is really an issue that needs to be addressed at a larger scale

    #Newsmast #Fediblock #ContentScraping

  14. @jackwilliambell

    I've had newsmast.social domain-blocked for a good while

    Last night I saw a rather personal, "sorry I've been gone for so long here's why" post that someone made that -- hashtags used aside -- was clearly not public

    Newmast scraped it because of the hashtag and broadcast it out to all subscribers to that hashtag

    I have little doubt that the person making the original post was aware of any of this or would have wanted their post broadcast by a content scraper

    And note that on their web site they claim to be run by a "charitable organization" and to request donations

    "Help us to remain ad-free and non-profit, whilst amplifying impactful, unheard voices on matters of global interest"

    Here: newsmastfoundation.org/donate/

    NOTE: Firefox throws a security warning; I continued because I've been on that web site before...

    cc @seb This (Newsmast) and its ilk is really an issue that needs to be addressed at a larger scale

    #Newsmast #Fediblock #ContentScraping

  15. @jackwilliambell

    I've had newsmast.social domain-blocked for a good while

    Last night I saw a rather personal, "sorry I've been gone for so long here's why" post that someone made that -- hashtags used aside -- was clearly not public

    Newmast scraped it because of the hashtag and broadcast it out to all subscribers to that hashtag

    I have little doubt that the person making the original post was aware of any of this or would have wanted their post broadcast by a content scraper

    And note that on their web site they claim to be run by a "charitable organization" and to request donations

    "Help us to remain ad-free and non-profit, whilst amplifying impactful, unheard voices on matters of global interest"

    Here: newsmastfoundation.org/donate/

    NOTE: Firefox throws a security warning; I continued because I've been on that web site before...

    cc @seb This (Newsmast) and its ilk is really an issue that needs to be addressed at a larger scale

    #Newsmast #Fediblock #ContentScraping

  16. @jackwilliambell

    I've had newsmast.social domain-blocked for a good while

    Last night I saw a rather personal, "sorry I've been gone for so long here's why" post that someone made that -- hashtags used aside -- was clearly not public

    Newmast scraped it because of the hashtag and broadcast it out to all subscribers to that hashtag

    I have little doubt that the person making the original post was aware of any of this or would have wanted their post broadcast by a content scraper

    And note that on their web site they claim to be run by a "charitable organization" and to request donations

    "Help us to remain ad-free and non-profit, whilst amplifying impactful, unheard voices on matters of global interest"

    Here: newsmastfoundation.org/donate/

    NOTE: Firefox throws a security warning; I continued because I've been on that web site before...

    cc @seb This (Newsmast) and its ilk is really an issue that needs to be addressed at a larger scale

    #Newsmast #Fediblock #ContentScraping

  17. @jackwilliambell

    I've had newsmast.social domain-blocked for a good while

    Last night I saw a rather personal, "sorry I've been gone for so long here's why" post that someone made that -- hashtags used aside -- was clearly not public

    Newmast scraped it because of the hashtag and broadcast it out to all subscribers to that hashtag

    I have little doubt that the person making the original post was aware of any of this or would have wanted their post broadcast by a content scraper

    And note that on their web site they claim to be run by a "charitable organization" and to request donations

    "Help us to remain ad-free and non-profit, whilst amplifying impactful, unheard voices on matters of global interest"

    Here: newsmastfoundation.org/donate/

    NOTE: Firefox throws a security warning; I continued because I've been on that web site before...

    cc @seb This (Newsmast) and its ilk is really an issue that needs to be addressed at a larger scale

    #Newsmast #Fediblock #ContentScraping