home.social

#ai-training — Public Fediverse posts

Live and recent posts from across the Fediverse tagged #ai-training, aggregated by home.social.

fetched live
  1. US National Science Foundation: New NSF initiative aims to unlock dataset value for AI-enabled scientific discovery. “The U.S. National Science Foundation announced the NSF Unlocking Dataset Value for AI-Enabled Scientific Discovery program, a new investment to advance scientific community datasets and enable discovery and innovation using artificial intelligence and other methods in national […]

    https://rbfirehose.com/2026/07/24/us-national-science-foundation-new-nsf-initiative-aims-to-unlock-dataset-value-for-ai-enabled-scientific-discovery/
  2. US National Science Foundation: New NSF initiative aims to unlock dataset value for AI-enabled scientific discovery. “The U.S. National Science Foundation announced the NSF Unlocking Dataset Value for AI-Enabled Scientific Discovery program, a new investment to advance scientific community datasets and enable discovery and innovation using artificial intelligence and other methods in national […]

    https://rbfirehose.com/2026/07/24/us-national-science-foundation-new-nsf-initiative-aims-to-unlock-dataset-value-for-ai-enabled-scientific-discovery/
  3. #Reddit shares fell 9% after the Wall Street Journal reported the company is considering ending its #contentdeal with #Google for #AItraining. The deal, worth $60 million annually, is expiring soon, and Reddit is reevaluating its benefits as Google’s #AIsummaries impact #searchtraffic. This comes amid a broader trend of media companies reassessing their relationships with AI giants. cnbc.com/2026/07/22/reddit-sto #tech #media #news

  4. #Reddit shares fell 9% after the Wall Street Journal reported the company is considering ending its #contentdeal with #Google for #AItraining. The deal, worth $60 million annually, is expiring soon, and Reddit is reevaluating its benefits as Google’s #AIsummaries impact #searchtraffic. This comes amid a broader trend of media companies reassessing their relationships with AI giants. cnbc.com/2026/07/22/reddit-sto #tech #media #news

  5. Georgia Tech: Researchers Use GeoGuessr Champion to Test Geolocation Accuracy in VLMs. “GeoGuessr is a geography browser game launched in 2013 that invites players to guess the location of random Google Street View images. [Radu] Casapu was already known as one of the top players in the world before he won the third annual GeoGussr World Championship in September. At the beginning of the spring […]

    https://rbfirehose.com/2026/07/23/georgia-tech-researchers-use-geoguessr-champion-to-test-geolocation-accuracy-in-vlms/
  6. Georgia Tech: Researchers Use GeoGuessr Champion to Test Geolocation Accuracy in VLMs. “GeoGuessr is a geography browser game launched in 2013 that invites players to guess the location of random Google Street View images. [Radu] Casapu was already known as one of the top players in the world before he won the third annual GeoGussr World Championship in September. At the beginning of the spring […]

    https://rbfirehose.com/2026/07/23/georgia-tech-researchers-use-geoguessr-champion-to-test-geolocation-accuracy-in-vlms/
  7. Reuters: News Corp countersues Brave for allegedly ‘scraping’ articles for AI . “News Corp, facing a lawsuit by search engine Brave Software, has filed a countersuit accusing it ‌of “flagrant theft” in distributing and selling versions of articles from the Wall Street Journal and New York Post to AI companies.”

    https://rbfirehose.com/2026/07/23/reuters-news-corp-countersues-brave-for-allegedly-scraping-articles-for-ai/
  8. Reuters: News Corp countersues Brave for allegedly ‘scraping’ articles for AI . “News Corp, facing a lawsuit by search engine Brave Software, has filed a countersuit accusing it ‌of “flagrant theft” in distributing and selling versions of articles from the Wall Street Journal and New York Post to AI companies.”

    https://rbfirehose.com/2026/07/23/reuters-news-corp-countersues-brave-for-allegedly-scraping-articles-for-ai/
  9. winbuzzer.com/2026/07/23/sony-

    Sony has filed a separate lawsuit against AI music generator Udio identifying 30,117 recordings used for AI training after a previous ruling kept its larger catalog out of its first case.

    #AI #Udio #Sony #AIMusic #GenAI #AITraining #AIAudio #AudioGeneration #Copyright #FairUse #MusicIndustry #YouTube

  10. winbuzzer.com/2026/07/23/sony-

    Sony has filed a separate lawsuit against AI music generator Udio identifying 30,117 recordings used for AI training after a previous ruling kept its larger catalog out of its first case.

    #AI #Udio #Sony #AIMusic #GenAI #AITraining #AIAudio #AudioGeneration #Copyright #FairUse #MusicIndustry #YouTube

  11. “The payout will deliver $3,000 per work across an estimated 500,000 works, shared among the authors and publishers who hold rights to them. While the settlement is believed to be the largest in the history of U.S. copyright law, many authors and creators still don’t view it as a win.

    That’s because of how the legal question was resolved. Alsup sided with Anthropic on the core issue. He ruled that training an AI model on copyrighted text counts as fair use — a decision widely seen as a turning point for the AI industry. But the ruling didn’t excuse how Anthropic obtained the books in the first place. Anthropic had built its training library from two sources: books it purchased and scanned (fine), and books it downloaded from pirate sites like Library Genesis and Pirate Library Mirror. Alsup found the second method illegal on its own terms and said that piracy question could go to trial; Anthropic agreed to a settlement soon after to avoid a trial and whatever damages a jury might have awarded.

    While the final approval closes out this case, it doesn’t settle the legal question industrywide because Alsup’s ruling was a single district court decision, and Anthropic’s decision to settle means the case will never reach an appeals court to become binding precedent.”

    techcrunch.com/2026/07/20/anth

    #AI #GenerativeAI #AITraining #Copyright #FairUse #Piracy #Anthropic #Claude

  12. “The payout will deliver $3,000 per work across an estimated 500,000 works, shared among the authors and publishers who hold rights to them. While the settlement is believed to be the largest in the history of U.S. copyright law, many authors and creators still don’t view it as a win.

    That’s because of how the legal question was resolved. Alsup sided with Anthropic on the core issue. He ruled that training an AI model on copyrighted text counts as fair use — a decision widely seen as a turning point for the AI industry. But the ruling didn’t excuse how Anthropic obtained the books in the first place. Anthropic had built its training library from two sources: books it purchased and scanned (fine), and books it downloaded from pirate sites like Library Genesis and Pirate Library Mirror. Alsup found the second method illegal on its own terms and said that piracy question could go to trial; Anthropic agreed to a settlement soon after to avoid a trial and whatever damages a jury might have awarded.

    While the final approval closes out this case, it doesn’t settle the legal question industrywide because Alsup’s ruling was a single district court decision, and Anthropic’s decision to settle means the case will never reach an appeals court to become binding precedent.”

    techcrunch.com/2026/07/20/anth

    #AI #GenerativeAI #AITraining #Copyright #FairUse #Piracy #Anthropic #Claude

  13. Zhinit: How To Train a Generative Kick Drum Model on Your Old Linux Desktop With 6GB of VRAM. “This article goes through how I trained and deployed a generative latent diffusion kick drum model from scratch on more than 13,000 kick drums from my personal sample library on a local Linux machine with a 7-year-old NVIDIA GeForce GTX 1660 SUPER (6GB of VRAM).”

    https://rbfirehose.com/2026/07/21/zhinit-how-to-train-a-generative-kick-drum-model-on-your-old-linux-desktop-with-6gb-of-vram/
  14. Zhinit: How To Train a Generative Kick Drum Model on Your Old Linux Desktop With 6GB of VRAM. “This article goes through how I trained and deployed a generative latent diffusion kick drum model from scratch on more than 13,000 kick drums from my personal sample library on a local Linux machine with a 7-year-old NVIDIA GeForce GTX 1660 SUPER (6GB of VRAM).”

    https://rbfirehose.com/2026/07/21/zhinit-how-to-train-a-generative-kick-drum-model-on-your-old-linux-desktop-with-6gb-of-vram/
  15. Reuters: US judge approves Anthropic’s $1.5 billion settlement of copyright lawsuit. “A federal judge in San Francisco on Monday signed off on ‌artificial intelligence company Anthropic’s landmark $1.5 billion settlement of a class action lawsuit brought by a group of authors who accused it of misusing their books to train its AI chatbot Claude.”

    https://rbfirehose.com/2026/07/21/reuters-us-judge-approves-anthropics-1-5-billion-settlement-of-copyright-lawsuit/
  16. Reuters: US judge approves Anthropic’s $1.5 billion settlement of copyright lawsuit. “A federal judge in San Francisco on Monday signed off on ‌artificial intelligence company Anthropic’s landmark $1.5 billion settlement of a class action lawsuit brought by a group of authors who accused it of misusing their books to train its AI chatbot Claude.”

    https://rbfirehose.com/2026/07/21/reuters-us-judge-approves-anthropics-1-5-billion-settlement-of-copyright-lawsuit/
  17. Music Business Worldwide: Sony Music sues Udio again, asserting over 30,000 recordings a judge barred the major from adding to its original case. “Sony Music Entertainment has filed a second copyright infringement lawsuit against Udio, asserting 30,117 sound recordings it says the AI music company copied without permission to train its generative AI models. The complaint, obtained and first […]

    https://rbfirehose.com/2026/07/21/music-business-worldwide-sony-music-sues-udio-again-asserting-over-30000-recordings-a-judge-barred-the-major-from-adding-to-its-original-case/
  18. Music Business Worldwide: Sony Music sues Udio again, asserting over 30,000 recordings a judge barred the major from adding to its original case. “Sony Music Entertainment has filed a second copyright infringement lawsuit against Udio, asserting 30,117 sound recordings it says the AI music company copied without permission to train its generative AI models. The complaint, obtained and first […]

    https://rbfirehose.com/2026/07/21/music-business-worldwide-sony-music-sues-udio-again-asserting-over-30000-recordings-a-judge-barred-the-major-from-adding-to-its-original-case/
  19. 🐢 Oh, joy! #Perforce has graced us with their AI-narrated "training" videos for a mere $500! Because nothing says value like paying a small fortune to hear a robot explain—well, who knows what, since it's just perpetually "Loading." 🙄
    training.perforce.com/learn/co #AItraining #TechHumor #OverpricedTech #AutomationFails #HackerNews #ngated

  20. 🐢 Oh, joy! #Perforce has graced us with their AI-narrated "training" videos for a mere $500! Because nothing says value like paying a small fortune to hear a robot explain—well, who knows what, since it's just perpetually "Loading." 🙄
    training.perforce.com/learn/co #AItraining #TechHumor #OverpricedTech #AutomationFails #HackerNews #ngated

  21. Did you miss a previous AI webinar? Would you like to review previous lessons in our AI series? Visit our Learn AI page to  access lessons, practice exercises, and additional AI resources. freedomscientific.com/training

    #AITraining #VisperoTraining

  22. Did you miss a previous AI webinar? Would you like to review previous lessons in our AI series? Visit our Learn AI page to  access lessons, practice exercises, and additional AI resources. freedomscientific.com/training

    #AITraining #VisperoTraining

  23. Engadget: A hacker accessed Suno source code that reportedly details how the company scraped millions of songs . “Suno — an app that vomits out soulless audio in the form of AI-generated ‘music’ — has been hacked. According to 404 Media, the hacker accessed data related to Suno’s training practices, as well as details on its customers.”

    https://rbfirehose.com/2026/07/16/engadget-a-hacker-accessed-suno-source-code-that-reportedly-details-how-the-company-scraped-millions-of-songs/
  24. Engadget: A hacker accessed Suno source code that reportedly details how the company scraped millions of songs . “Suno — an app that vomits out soulless audio in the form of AI-generated ‘music’ — has been hacked. According to 404 Media, the hacker accessed data related to Suno’s training practices, as well as details on its customers.”

    https://rbfirehose.com/2026/07/16/engadget-a-hacker-accessed-suno-source-code-that-reportedly-details-how-the-company-scraped-millions-of-songs/
  25. PetaPixel: Patreon Blocks AI Crawlers From Copying Content: ‘Creators Deserve Compenstion’. “Patreon has announced a new update that blocks AI crawlers from hoovering up content on the platform for training purposes.”

    https://rbfirehose.com/2026/07/16/patreon-blocks-ai-crawlers-from-copying-content-creators-deserve-compenstion-petapixel/
  26. PetaPixel: Patreon Blocks AI Crawlers From Copying Content: ‘Creators Deserve Compenstion’. “Patreon has announced a new update that blocks AI crawlers from hoovering up content on the platform for training purposes.”

    https://rbfirehose.com/2026/07/16/patreon-blocks-ai-crawlers-from-copying-content-creators-deserve-compenstion-petapixel/
  27. 9to5 Google: Samsung Health won’t delete all of your data if you opt out of AI training after all [U]. “In a statement to 9to5Google, Samsung says that it won’t delete all of your data if you opt-out of AI training. Rather, it will only delete ‘data collected for AI development.’ In other words, it sounds like it is deleting data from Samsung’s end rather than the user’s.”

    https://rbfirehose.com/2026/07/16/9to5-google-samsung-health-wont-delete-all-of-your-data-if-you-opt-out-of-ai-training-after-all-u/
  28. Enterprise AI training vs. #AIInference: two fundamentally different workloads.

    #AITraining: intensive GPU compute over days/weeks
    Inference: fast, continuous production responses

    Conflate them and you overpay for infrastructure, choose the wrong hardware, and you miss compliance. Both introduce distinct security risks.

    The solution?
    Private, sovereign AI infrastructure = full data control + compliance.

    amazee.ai/blog/ai-training-vs-

  29. "The hacked data is a rare look at exactly how AI models and tools are built. Suno is one of the largest AI music generation tools on the internet, and has been the subject of several major lawsuits from the record industry, which accused the company of training on millions of copyrighted songs. As part of these legal proceedings, Suno previously admitted that it was trained on “essentially all music files of reasonable quality that are accessible on the open internet,” which included a total of “tens of millions of recordings.” Suno has been making the argument that it is allowed to train on copyrighted works as fair use in those cases, one of which has been settled.

    The lawsuits have made clear that Suno did train on huge amounts of copyrighted works, but the hacked data shared with 404 Media sheds more light on how Suno scraped songs from the internet and where it took them from. The Recording Industry Association of America accused Suno of ripping songs directly from YouTube; the hacked data seen by 404 Media confirms this.

    The hacked material includes source code that appears to be from 2023 and 2024 that includes scraping instructions and details about the scope of at least some of the scraping. For example, the comments in one file note that they will pull from “genius_hq, youtube_music, freesound, jamendo, imp, deezer, ytm_tagged,” and that “non-music will be filtered out.” A file called “youtube_music” notes that at the time the file was last updated, it had ingested “2,013,545 music clips.” Another file contains comments about different datasets Suno had created, which included “113,879 hours of youtube_music,” “17,615 hours of genius_hq,” “410 hours of free sound,” “19,514 hours of imslp,” “3,726 hours of jamendo,” “62,117 hours of pond5_music,” “12,287 hours of deezer,” “152,162 hours of ytm_tagged,” and “103 hours of musescore_lyrics.”"

    404media.co/hack-reveals-suno-

    #AI #GenerativeAI #Suno #AITraining #CyberSecurity #Copyright #IP

  30. Axios: Help wanted: Anthropic hires to head off catastrophe. “Anthropic has 32 very scary job openings for roles designed to prevent people from using AI to build everything from human-made explosives to nuclear weapons. Catch up quick: Anthropic is hiring analysts focused on chemicals and explosives, nuclear weapons, financial scams, cybercrime and more.”

    https://rbfirehose.com/2026/07/15/help-wanted-anthropic-hires-to-head-off-catastrophe-axios/
  31. Reuters: SoftBank’s Son says AI will need $5 trillion per year by 2040, dismisses bubble talk. “The development of AI will require investment of $5 trillion each year by 2040, and any talk ‌of a bubble forming around the technology is “absurd”, SoftBank Group CEO Masayoshi Son said on Tuesday.”

    https://rbfirehose.com/2026/07/14/reuters-softbanks-son-says-ai-will-need-5-trillion-per-year-by-2040-dismisses-bubble-talk/
  32. 9to5 Google: Samsung Health will delete your data unless you hand it over for AI training. “As its new redesign rolls out, Samsung Health is also forcing users into a choice – let their data be used for AI training, or make it effectively unusable in the long term.”

    https://rbfirehose.com/2026/07/14/9to5-google-samsung-health-will-delete-your-data-unless-you-hand-it-over-for-ai-training/
  33. Ah, the infinite #recursion of the tech world's buzzword bingo! 🤖💸 Train an AI to train AIs with RL, and somehow end up $1.3k poorer. Because who needs money when you can have infinite loops of machine learning pretense? 🙃
    github.com/Danau5tin/ai-trains #buzzwordbingo #AItraining #infinitelearning #techhumor #HackerNews #ngated