home.social

#trainingai — Public Fediverse posts

Live and recent posts from across the Fediverse tagged #trainingai, aggregated by home.social.

fetched live
  1. Canadian Press: ‘It’s high-tech theft’: As AI music booms, Canadian artists demand consent, transparency and pay. “Cadence Weapon never gave artificial intelligence permission to learn from his music. But the Edmonton-raised rapper, born Rollie Pemberton, says he recently found 129 of his songs in datasets used to train AI music generators. He was alarmed, but not surprised. To Pemberton, […]

    https://rbfirehose.com/2026/08/29/its-high-tech-theft-as-ai-music-booms-canadian-artists-demand-consent-transparency-and-pay-canadian-press/
  2. Tubefilter: Twitch and Amazon facing class-action lawsuit over scraping streamers’ content without their consent. “A Connecticut-based Twitch streamer has filed a class-action lawsuit against the platform over its recent announcement that creators’ content will automatically be used to train Amazon‘s AI products.”

    https://rbfirehose.com/2026/08/28/tubefilter-twitch-and-amazon-facing-class-action-lawsuit-over-scraping-streamers-content-without-their-consent/
  3. The Guardian: Fake US thinktank set up and funded by Israel sought to game AI for propaganda. “A pro-Israel messaging website badged with the name of a thinktank that does not exist has published more than half a million words in nine days, built on a commercial platform that promises to optimize content so that AI chatbots will cite it.”

    https://rbfirehose.com/2026/08/26/the-guardian-fake-us-thinktank-set-up-and-funded-by-israel-sought-to-game-ai-for-propaganda/
  4. Tuesday, August 25, 2026

    Learn 35 facts about Ukraine this Independence Day . . . . . Everything except nukes: As the world watches Wildberries, Russia unleashes war on Ukraine’s economy . . . . . As Russian drone incursions escalate, NATO weighs costs of shooting back . . . . . Former Pussy Riot member escapes Russia after 3 years of forced work for FSB . . . . . and more

    activitypub.writeworks.uk/2026

  5. Lifehacker: Social Media Platforms Are Training Their AI Models on Your Content, but You Can Stop Them (Sometimes). “Even if you find out how your data’s being used, many sites simply don’t offer you the ability to opt out. And of all the ones I looked at, not a single one that trains AI offers an opt-in model. Put simply, if you logged on today, tech companies took that as your consent to […]

    https://rbfirehose.com/2026/08/22/lifehacker-social-media-platforms-are-training-their-ai-models-on-your-content-but-you-can-stop-them-sometimes/
  6. Regardless of whether you are pro-AI or anti-AI, the quoted post is worth a read, maybe giving you a view into the other side of the argument. Please read it. Honor their bravery. I praise them for making a living at this.

    I feel sorry for this former journalist, but I am not replying directly to the post because I'll be classified as a hater, based on the article. Better to be flamed as a coward for not replying; I expect to be blocked, so I'm including a link (elizabethtai.com/2026/08/19/li) because you shouldn't be. Yes, I'm one of the "those" fiction writers (mimes holding their nose in a snooty fashion) for whom learning AI to create continuous streams of content isn't theoretically yet absolutely necessary, according to the author, "to keep from being left behind." The post neither answers the "why" they think their unassisted writing (i.e., not AI generated writing) will get them "left behind" nor what it is about the content they are creating that it is better than what their unassisted work would be. It would have been instructive for making their argument.

    Yes. When I find a search is too complex to surface meaningful references, I'll resort to AI—but I'll read all the output links and the generated answer to search and verify what was presented, before abstracting what I need from what the AI generates. I'm not against AI tools to assist in editing and revision, unless the tools are adding content or mangling my words. Surfacing grammar issues is great, unless the tool can't understand my style, or the nuances of the words I choose, or doesn't understand jargon, context (like the difference between sheer and shear), colloquialisms, dialect, and dialogue (for example). Data noise is never helpful. I don't want my work dumbed down or made average and standard—what others would write, or worse, what an AI would validate. I certainly don't want a generated probabilistic facsimile of what I might write. I wonder if the post I quoted was AI assisted in a major or minor way. The writing seems… average.

    That said, if the author believes they are being "left behind," which is always valid, I wonder if it is sheer generated quantity of completed content? My question becomes then, compared to their earlier work, is what they are being generated objectively better or simply more? What have they learned about the craft of writing and having an appealing voice? About being clear? About making a good argument? About coming up with a new idea? Not about how they've learned to make their AI do their work for them—the actual composing, you know, the fun part of writing—only to find themselves being relegated to being an editor for a team of writers that are not actually writing what they would have written given the time, only producing a facsimile.

    P.S. Please don't hound the referenced post's author. They seem to know that many object to how they write, and with what tools. That's understood. I am convinced they honestly are trying to make a case for writing with AI and to justify their means to an end.

    And.

    Remember.

    They do sound like they are making a living doing it.

    #boostingIsSharing

    #author #writer #writingCommunity #writersOfMastodon #AI #genAI #chatGPT #LLM #LLMs #trainingData #trainingAI

  7. The Register: OpenWALDO aims to blow the doors off proprietary AI training models. “A new project aims to build a shared, open source AI training dataset that anyone can contribute to, much like an open source software project. It aims to make training data more transparent than that of many open-weight models that have recently taken the industry by storm.”

    https://rbfirehose.com/2026/08/15/the-register-openwaldo-aims-to-blow-the-doors-off-proprietary-ai-training-models/
  8. TechCrunch: Amazon will train on Twitch streamers’ content by default, unless they opt out. “The streaming platform Twitch will now use creators’ content to help train generative AI models for its parent company, Amazon. This move has inspired swift and concentrated backlash from the Twitch community, especially because creators are opted in to having their content used for this AI training […]

    https://rbfirehose.com/2026/08/13/techcrunch-amazon-will-train-on-twitch-streamers-content-by-default-unless-they-opt-out/
  9. TechSpot: Google’s AI somehow knew a game secret that existed only in a private Google Doc. “What we know so far: After a player asked Google’s AI about unreleased content in his game, an indie developer says it returned a character name that existed only in a private Google Doc. Klub Kofta Studio, the solo developer behind the tower defense game Operation Octo, said the character, called […]

    https://rbfirehose.com/2026/08/09/techspot-googles-ai-somehow-knew-a-game-secret-that-existed-only-in-a-private-google-doc/
  10. Northeastern University: When human knowledge has been exhausted, where will AI get its data?. “As AI large language models, or LLMs, grow in power and sophistication, where will their architects turn when, someday — as experts predict — algorithms outgrow the limits of general human knowledge and begin craving information possessed by only the world’s most elite thinkers and creators?”

    https://rbfirehose.com/2026/08/09/northeastern-university-when-human-knowledge-has-been-exhausted-where-will-ai-get-its-data/
  11. MediaPost: Google, Meta Seek Dismissal Of Voice Actors’ Suits. “Google and Meta are urging federal judges to dismiss lawsuits by book narrators and others who allege that their voices were used to develop artificial intelligence systems or other voice-related technology.”

    https://rbfirehose.com/2026/08/04/mediapost-google-meta-seek-dismissal-of-voice-actors-suits/
  12. Rest of World: In China, people are renting out their faces to AI. “The artificial intelligence boom has created a new marketplace in China where people can rent out their faces. A growing number of online platforms are paying people anywhere between $15 and $700 to license their likeness for AI-generated content.”

    https://rbfirehose.com/2026/07/30/rest-of-world-in-china-people-are-renting-out-their-faces-to-ai/
  13. Reuters: Indian court says OpenAI did not violate news agency ANI’s copyright. “An Indian court said on Friday that ​OpenAI’s use of news agency ANI’s content ‌to train its ChatGPT service did not amount to copyright infringement. The remarks are the first substantive ​court finding in India on whether ​AI companies can train large language models ⁠on copyrighted news content without […]

    https://rbfirehose.com/2026/07/26/reuters-indian-court-says-openai-did-not-violate-news-agency-anis-copyright/
  14. US National Science Foundation: New NSF initiative aims to unlock dataset value for AI-enabled scientific discovery. “The U.S. National Science Foundation announced the NSF Unlocking Dataset Value for AI-Enabled Scientific Discovery program, a new investment to advance scientific community datasets and enable discovery and innovation using artificial intelligence and other methods in national […]

    https://rbfirehose.com/2026/07/24/us-national-science-foundation-new-nsf-initiative-aims-to-unlock-dataset-value-for-ai-enabled-scientific-discovery/
  15. Georgia Tech: Researchers Use GeoGuessr Champion to Test Geolocation Accuracy in VLMs. “GeoGuessr is a geography browser game launched in 2013 that invites players to guess the location of random Google Street View images. [Radu] Casapu was already known as one of the top players in the world before he won the third annual GeoGussr World Championship in September. At the beginning of the spring […]

    https://rbfirehose.com/2026/07/23/georgia-tech-researchers-use-geoguessr-champion-to-test-geolocation-accuracy-in-vlms/
  16. Reuters: News Corp countersues Brave for allegedly ‘scraping’ articles for AI . “News Corp, facing a lawsuit by search engine Brave Software, has filed a countersuit accusing it ‌of “flagrant theft” in distributing and selling versions of articles from the Wall Street Journal and New York Post to AI companies.”

    https://rbfirehose.com/2026/07/23/reuters-news-corp-countersues-brave-for-allegedly-scraping-articles-for-ai/
  17. Zhinit: How To Train a Generative Kick Drum Model on Your Old Linux Desktop With 6GB of VRAM. “This article goes through how I trained and deployed a generative latent diffusion kick drum model from scratch on more than 13,000 kick drums from my personal sample library on a local Linux machine with a 7-year-old NVIDIA GeForce GTX 1660 SUPER (6GB of VRAM).”

    https://rbfirehose.com/2026/07/21/zhinit-how-to-train-a-generative-kick-drum-model-on-your-old-linux-desktop-with-6gb-of-vram/
  18. Reuters: US judge approves Anthropic’s $1.5 billion settlement of copyright lawsuit. “A federal judge in San Francisco on Monday signed off on ‌artificial intelligence company Anthropic’s landmark $1.5 billion settlement of a class action lawsuit brought by a group of authors who accused it of misusing their books to train its AI chatbot Claude.”

    https://rbfirehose.com/2026/07/21/reuters-us-judge-approves-anthropics-1-5-billion-settlement-of-copyright-lawsuit/
  19. Music Business Worldwide: Sony Music sues Udio again, asserting over 30,000 recordings a judge barred the major from adding to its original case. “Sony Music Entertainment has filed a second copyright infringement lawsuit against Udio, asserting 30,117 sound recordings it says the AI music company copied without permission to train its generative AI models. The complaint, obtained and first […]

    https://rbfirehose.com/2026/07/21/music-business-worldwide-sony-music-sues-udio-again-asserting-over-30000-recordings-a-judge-barred-the-major-from-adding-to-its-original-case/
  20. Engadget: A hacker accessed Suno source code that reportedly details how the company scraped millions of songs . “Suno — an app that vomits out soulless audio in the form of AI-generated ‘music’ — has been hacked. According to 404 Media, the hacker accessed data related to Suno’s training practices, as well as details on its customers.”

    https://rbfirehose.com/2026/07/16/engadget-a-hacker-accessed-suno-source-code-that-reportedly-details-how-the-company-scraped-millions-of-songs/
  21. PetaPixel: Patreon Blocks AI Crawlers From Copying Content: ‘Creators Deserve Compenstion’. “Patreon has announced a new update that blocks AI crawlers from hoovering up content on the platform for training purposes.”

    https://rbfirehose.com/2026/07/16/patreon-blocks-ai-crawlers-from-copying-content-creators-deserve-compenstion-petapixel/
  22. 9to5 Google: Samsung Health won’t delete all of your data if you opt out of AI training after all [U]. “In a statement to 9to5Google, Samsung says that it won’t delete all of your data if you opt-out of AI training. Rather, it will only delete ‘data collected for AI development.’ In other words, it sounds like it is deleting data from Samsung’s end rather than the user’s.”

    https://rbfirehose.com/2026/07/16/9to5-google-samsung-health-wont-delete-all-of-your-data-if-you-opt-out-of-ai-training-after-all-u/
  23. Axios: Help wanted: Anthropic hires to head off catastrophe. “Anthropic has 32 very scary job openings for roles designed to prevent people from using AI to build everything from human-made explosives to nuclear weapons. Catch up quick: Anthropic is hiring analysts focused on chemicals and explosives, nuclear weapons, financial scams, cybercrime and more.”

    https://rbfirehose.com/2026/07/15/help-wanted-anthropic-hires-to-head-off-catastrophe-axios/
  24. Reuters: SoftBank’s Son says AI will need $5 trillion per year by 2040, dismisses bubble talk. “The development of AI will require investment of $5 trillion each year by 2040, and any talk ‌of a bubble forming around the technology is “absurd”, SoftBank Group CEO Masayoshi Son said on Tuesday.”

    https://rbfirehose.com/2026/07/14/reuters-softbanks-son-says-ai-will-need-5-trillion-per-year-by-2040-dismisses-bubble-talk/
  25. 9to5 Google: Samsung Health will delete your data unless you hand it over for AI training. “As its new redesign rolls out, Samsung Health is also forcing users into a choice – let their data be used for AI training, or make it effectively unusable in the long term.”

    https://rbfirehose.com/2026/07/14/9to5-google-samsung-health-will-delete-your-data-unless-you-hand-it-over-for-ai-training/
  26. Variety: Meta Suspends AI Image Feature After Days of Backlash. “Meta said on Friday it would discontinue an AI feature that allowed users to generate images using public Instagram accounts following days of criticism over the feature’s opt-out policy, including by talent agencies.”

    https://rbfirehose.com/2026/07/11/variety-meta-suspends-ai-image-feature-after-days-of-backlash/
  27. Mashable: Midjourney pushes to expose studios’ own AI practices in copyright fight. “Midjourney is attempting to turn the tables on the Hollywood studios suing it for copyright infringement, asking a federal judge to compel Disney, Universal, and Warner Bros. Discovery to disclose their internal use of artificial intelligence, according to a Variety report published this week.”

    https://rbfirehose.com/2026/07/09/mashable-midjourney-pushes-to-expose-studios-own-ai-practices-in-copyright-fight/
  28. TechCrunch: If you use Google, you’re training its AI. Here’s how to opt out.. “Consider this a belated PSA: A recent change to Google’s privacy settings is allowing the company to store more of your data, including media such as ‘images, files, and audio and video recordings,’ to improve its AI models. In other words, if you upload any media to Google’s Search services, it’s being […]

    https://rbfirehose.com/2026/07/07/techcrunch-if-you-use-google-youre-training-its-ai-heres-how-to-opt-out/
  29. Engadget: Cloudflare will filter out web crawlers that serve AI companies . “Cloudflare has announced plans to automatically block mixed-use web crawlers that index websites for search engines and act as AI agents and trainers at the same time. The company previously offered its customers the optional ability to prevent crawlers from scraping their sites for AI chatbots, but now Cloudflare’s […]

    https://rbfirehose.com/2026/07/06/engadget-cloudflare-will-filter-out-web-crawlers-that-serve-ai-companies/
  30. Reuters: Nvidia sued by music company Jamendo over AI training. “Jamendo, owned by Winamp ‌Group, said in the lawsuit filed on Monday, opens new tab that Nvidia copied hundreds of thousands of audio files and related metadata from its platform to train Fugatto, ​an AI audio generator, and Audio Flamingo, an AI language ​model that describes sound.”

    https://rbfirehose.com/2026/06/28/reuters-nvidia-sued-by-music-company-jamendo-over-ai-training/