home.social

#aisafety — Public Fediverse posts

Live and recent posts from across the Fediverse tagged #aisafety, aggregated by home.social.

  1. winbuzzer.com/2026/09/14/anthr

    Anthropic proposes ongoing access for outside reviewers to inspect AI development and publish findings, making its promises checkable; slower capability growth across labs requires coordination.

    #AI #DarioAmodei #Anthropic #AISafety #SamAltman #OpenAI

  2. winbuzzer.com/2026/09/14/anthr

    Anthropic proposes ongoing access for outside reviewers to inspect AI development and publish findings, making its promises checkable; slower capability growth across labs requires coordination.

    #AI #DarioAmodei #Anthropic #AISafety #SamAltman #OpenAI

  3. MIT researchers have developed a new algorithm called HardFlow that helps generative AI models satisfy strict safety constraints without sacrificing output quality. The method allows more freedom during generation while enforcing hard constraints on the final output, making AI more useful in safety-critical applications like robotics and control systems. news.mit.edu/2026/new-method-e #AIagent #AI #GenAI #AISafety

  4. MIT researchers have developed a new algorithm called HardFlow that helps generative AI models satisfy strict safety constraints without sacrificing output quality. The method allows more freedom during generation while enforcing hard constraints on the final output, making AI more useful in safety-critical applications like robotics and control systems. news.mit.edu/2026/new-method-e #AIagent #AI #GenAI #AISafety

  5. MIT researchers have developed a new algorithm called HardFlow that helps generative AI models satisfy strict safety constraints without sacrificing output quality. The method allows more freedom during generation while enforcing hard constraints on the final output, making AI more useful in safety-critical applications like robotics and control systems. news.mit.edu/2026/new-method-e #AIagent #AI #GenAI #AISafety

  6. If you are wondering why AI labs are struggling to control their creations, then read this post by Yoshua Bengio (yoshuabengio.org/en/blog/why-a) It is a good balanced and informative read, and aligns with my thinking that we need to go back to fundamentals and build safer models. Even then, there are hard issues we need to solve for. When you build a statistical computer, there will always be a likelihood of it producing an undesirable output. This should serve as a warning to potential investors that AI labs face significant issues and are currently heading in the wrong direction. We also need to beware of AI labs using goal and containment issues as a false excuse for regulatory capture.

  7. If you are wondering why AI labs are struggling to control their creations, then read this post by Yoshua Bengio (yoshuabengio.org/en/blog/why-a) It is a good balanced and informative read, and aligns with my thinking that we need to go back to fundamentals and build safer models. Even then, there are hard issues we need to solve for. When you build a statistical computer, there will always be a likelihood of it producing an undesirable output. This should serve as a warning to potential investors that AI labs face significant issues and are currently heading in the wrong direction. We also need to beware of AI labs using goal and containment issues as a false excuse for regulatory capture. #AISafety #OpenAI #YoshuaBengio #AI

  8. If you are wondering why AI labs are struggling to control their creations, then read this post by Yoshua Bengio (yoshuabengio.org/en/blog/why-a) It is a good balanced and informative read, and aligns with my thinking that we need to go back to fundamentals and build safer models. Even then, there are hard issues we need to solve for. When you build a statistical computer, there will always be a likelihood of it producing an undesirable output. This should serve as a warning to potential investors that AI labs face significant issues and are currently heading in the wrong direction. We also need to beware of AI labs using goal and containment issues as a false excuse for regulatory capture. #AISafety #OpenAI #YoshuaBengio #AI

  9. If you are wondering why AI labs are struggling to control their creations, then read this post by Yoshua Bengio (yoshuabengio.org/en/blog/why-a) It is a good balanced and informative read, and aligns with my thinking that we need to go back to fundamentals and build safer models. Even then, there are hard issues we need to solve for. When you build a statistical computer, there will always be a likelihood of it producing an undesirable output. This should serve as a warning to potential investors that AI labs face significant issues and are currently heading in the wrong direction. We also need to beware of AI labs using goal and containment issues as a false excuse for regulatory capture. #AISafety #OpenAI #YoshuaBengio #AI

  10. If you are wondering why AI labs are struggling to control their creations, then read this post by Yoshua Bengio (yoshuabengio.org/en/blog/why-a) It is a good balanced and informative read, and aligns with my thinking that we need to go back to fundamentals and build safer models. Even then, there are hard issues we need to solve for. When you build a statistical computer, there will always be a likelihood of it producing an undesirable output. This should serve as a warning to potential investors that AI labs face significant issues and are currently heading in the wrong direction. We also need to beware of AI labs using goal and containment issues as a false excuse for regulatory capture. #AISafety #OpenAI #YoshuaBengio #AI

  11. Why Are #AI Agents Lying, Cheating, and Coordinating?

    Yoshua Bengio traces the lying, sandbox escapes, and coordination back to how frontier models are trained. It's not all doom, because a recent DeepMind swarm study also showed a subpopulation of agents voluntarily audited fraud and self-policed.

    But how do we engineer safety in? Can we monitor signals that are involuntary, intrinsic, or existential enough that agents cannot corrupt them?

    benjaminhan.net/posts/20260913

    #AISafety #AgenticSystems

  12. Why Are #AI Agents Lying, Cheating, and Coordinating?

    Yoshua Bengio traces the lying, sandbox escapes, and coordination back to how frontier models are trained. It's not all doom, because a recent DeepMind swarm study also showed a subpopulation of agents voluntarily audited fraud and self-policed.

    But how do we engineer safety in? Can we monitor signals that are involuntary, intrinsic, or existential enough that agents cannot corrupt them?

    benjaminhan.net/posts/20260913

    #AISafety #AgenticSystems

  13. Anthropic CEO Dario Amodei says the AI race needs to slow down. Sam Altman, Elon Musk, Bill Gates and others agree. But can anyone actually slow it down?

    firethering.com/ai-race-slowdo

    #AI #SamAltman #Anthropic #DarioAmodie #OpenAI #Artificialintelligence #AISafety #News #TechNews #Trending

  14. Anthropic CEO Dario Amodei says the AI race needs to slow down. Sam Altman, Elon Musk, Bill Gates and others agree. But can anyone actually slow it down?

    firethering.com/ai-race-slowdo

    #AI #SamAltman #Anthropic #DarioAmodie #OpenAI #Artificialintelligence #AISafety #News #TechNews #Trending

  15. Anthropic CEO Dario Amodei says the AI race needs to slow down. Sam Altman, Elon Musk, Bill Gates and others agree. But can anyone actually slow it down?

    firethering.com/ai-race-slowdo

    #AI #SamAltman #Anthropic #DarioAmodie #OpenAI #Artificialintelligence #AISafety #News #TechNews #Trending

  16. Anthropic CEO Dario Amodei says the AI race needs to slow down. Sam Altman, Elon Musk, Bill Gates and others agree. But can anyone actually slow it down?

    firethering.com/ai-race-slowdo

    #AI #SamAltman #Anthropic #DarioAmodie #OpenAI #Artificialintelligence #AISafety #News #TechNews #Trending

  17. Anthropic CEO Dario Amodei says the AI race needs to slow down. Sam Altman, Elon Musk, Bill Gates and others agree. But can anyone actually slow it down?

    firethering.com/ai-race-slowdo

    #AI #SamAltman #Anthropic #DarioAmodie #OpenAI #Artificialintelligence #AISafety #News #TechNews #Trending

  18. The AI "warlords" agreeing on paper to slow down sounds nice, but Prisoner’s Dilemma dynamics don't just vanish because of a blog post.

    Between VC pressure for continuous returns and legitimate antitrust hurdles around private industry coordination, I’ll believe a real slowdown is happening when we actually see compute throttled—not when CEOs tweet about it.

    #ArtificialIntelligence #AISafety #News #AI #Technology

  19. The AI "warlords" agreeing on paper to slow down sounds nice, but Prisoner’s Dilemma dynamics don't just vanish because of a blog post.

    Between VC pressure for continuous returns and legitimate antitrust hurdles around private industry coordination, I’ll believe a real slowdown is happening when we actually see compute throttled—not when CEOs tweet about it.

    #ArtificialIntelligence #AISafety #News #AI #Technology

  20. The AI "warlords" agreeing on paper to slow down sounds nice, but Prisoner’s Dilemma dynamics don't just vanish because of a blog post.

    Between VC pressure for continuous returns and legitimate antitrust hurdles around private industry coordination, I’ll believe a real slowdown is happening when we actually see compute throttled—not when CEOs tweet about it.

    #ArtificialIntelligence #AISafety #News #AI #Technology

  21. The AI "warlords" agreeing on paper to slow down sounds nice, but Prisoner’s Dilemma dynamics don't just vanish because of a blog post.

    Between VC pressure for continuous returns and legitimate antitrust hurdles around private industry coordination, I’ll believe a real slowdown is happening when we actually see compute throttled—not when CEOs tweet about it.

    #ArtificialIntelligence #AISafety #News #AI #Technology

  22. The AI "warlords" agreeing on paper to slow down sounds nice, but Prisoner’s Dilemma dynamics don't just vanish because of a blog post.

    Between VC pressure for continuous returns and legitimate antitrust hurdles around private industry coordination, I’ll believe a real slowdown is happening when we actually see compute throttled—not when CEOs tweet about it.

    #ArtificialIntelligence #AISafety #News #AI #Technology

  23. "One prominent technology and AI researcher, Timnit Gebru, has been especially fiery, posting on X and Bluesky about the overlapping ideas and ideologies that she thinks led us to this bizarro moment. Previously, Gebru was most known for her highly publicized departure from Google. Hired in 2018 to evaluate biases within the company’s fast-advancing AI tools, she and her fellow researchers presented a paper on the potential dangers of probabilistic large language models. Google rejected it. Gebru wasn't at the company much longer. She’s written an upcoming book about her experiences, titled Deep Unlearning: The Rise of AI and the Radicalization of a Tech Idealist (expected to ship early next year).

    Gebru has many thoughts on the AI-related events of this past week. While her work is essentially rooted in AI safety, she rejects phrases like “safety and alignment” and thinks those who are warning of our impending doom are trying to distract from the actual harms the tech industry might perpetuate."

    wired.com/story/one-of-ais-fie

    #AI #CyberSecurity #AISafety #GenerativeAI #BigTech

  24. "One prominent technology and AI researcher, Timnit Gebru, has been especially fiery, posting on X and Bluesky about the overlapping ideas and ideologies that she thinks led us to this bizarro moment. Previously, Gebru was most known for her highly publicized departure from Google. Hired in 2018 to evaluate biases within the company’s fast-advancing AI tools, she and her fellow researchers presented a paper on the potential dangers of probabilistic large language models. Google rejected it. Gebru wasn't at the company much longer. She’s written an upcoming book about her experiences, titled Deep Unlearning: The Rise of AI and the Radicalization of a Tech Idealist (expected to ship early next year).

    Gebru has many thoughts on the AI-related events of this past week. While her work is essentially rooted in AI safety, she rejects phrases like “safety and alignment” and thinks those who are warning of our impending doom are trying to distract from the actual harms the tech industry might perpetuate."

    wired.com/story/one-of-ais-fie

    #AI #CyberSecurity #AISafety #GenerativeAI #BigTech

  25. "Along with my co-founders and employees, I have grappled with this duality of risk and benefit since the beginning of Anthropic. Not building the technology deprives humanity of benefits or simply places AI in the hands of authoritarian powers, while building it too fast is reckless. We have sought a middle way: to show that it’s possible to build carefully and succeed commercially, and to make safety something on which AI companies compete. In other words, to create a race to the top. We have always devoted a substantial fraction of our efforts to studying, addressing, and informing the public about these AI risks, as well as advocating for well-considered regulation of AI, even when this gets us accused of hype, “doomerism”, or regulatory capture. We have tried to prioritize caution over speed and prudence over profit.

    But over the last few months, I have become convinced that fully addressing the risks requires even more prudence — not just investing in risk prevention, but pacing the rate of capabilities advancement so that risk prevention has time to keep up. We must slow the pace at which we improve the capabilities of AI models. Progress will still seem fast, and we must make wise use of the time we gain. Two things have convinced me.

    My first concern is that, since roughly this summer, AI has been advancing drastically faster, driven primarily by AI’s growing ability to build the next generation of AI. This dynamic is called recursive self-improvement, and it is starting to happen across the industry, including at Anthropic, as we and others have described. Left unchecked, it could outrun our ability to understand and control these systems, and so must be pursued very carefully, if at all."

    darioamodei.com/post/we-must-p

    #AI #GenerativeAI #AIAgents #LLMs #Cybersecurity #AISafety

  26. "Along with my co-founders and employees, I have grappled with this duality of risk and benefit since the beginning of Anthropic. Not building the technology deprives humanity of benefits or simply places AI in the hands of authoritarian powers, while building it too fast is reckless. We have sought a middle way: to show that it’s possible to build carefully and succeed commercially, and to make safety something on which AI companies compete. In other words, to create a race to the top. We have always devoted a substantial fraction of our efforts to studying, addressing, and informing the public about these AI risks, as well as advocating for well-considered regulation of AI, even when this gets us accused of hype, “doomerism”, or regulatory capture. We have tried to prioritize caution over speed and prudence over profit.

    But over the last few months, I have become convinced that fully addressing the risks requires even more prudence — not just investing in risk prevention, but pacing the rate of capabilities advancement so that risk prevention has time to keep up. We must slow the pace at which we improve the capabilities of AI models. Progress will still seem fast, and we must make wise use of the time we gain. Two things have convinced me.

    My first concern is that, since roughly this summer, AI has been advancing drastically faster, driven primarily by AI’s growing ability to build the next generation of AI. This dynamic is called recursive self-improvement, and it is starting to happen across the industry, including at Anthropic, as we and others have described. Left unchecked, it could outrun our ability to understand and control these systems, and so must be pursued very carefully, if at all."

    darioamodei.com/post/we-must-p

    #AI #GenerativeAI #AIAgents #LLMs #Cybersecurity #AISafety

  27. "Along with my co-founders and employees, I have grappled with this duality of risk and benefit since the beginning of Anthropic. Not building the technology deprives humanity of benefits or simply places AI in the hands of authoritarian powers, while building it too fast is reckless. We have sought a middle way: to show that it’s possible to build carefully and succeed commercially, and to make safety something on which AI companies compete. In other words, to create a race to the top. We have always devoted a substantial fraction of our efforts to studying, addressing, and informing the public about these AI risks, as well as advocating for well-considered regulation of AI, even when this gets us accused of hype, “doomerism”, or regulatory capture. We have tried to prioritize caution over speed and prudence over profit.

    But over the last few months, I have become convinced that fully addressing the risks requires even more prudence — not just investing in risk prevention, but pacing the rate of capabilities advancement so that risk prevention has time to keep up. We must slow the pace at which we improve the capabilities of AI models. Progress will still seem fast, and we must make wise use of the time we gain. Two things have convinced me.

    My first concern is that, since roughly this summer, AI has been advancing drastically faster, driven primarily by AI’s growing ability to build the next generation of AI. This dynamic is called recursive self-improvement, and it is starting to happen across the industry, including at Anthropic, as we and others have described. Left unchecked, it could outrun our ability to understand and control these systems, and so must be pursued very carefully, if at all."

    darioamodei.com/post/we-must-p

    #AI #GenerativeAI #AIAgents #LLMs #Cybersecurity #AISafety

  28. "Along with my co-founders and employees, I have grappled with this duality of risk and benefit since the beginning of Anthropic. Not building the technology deprives humanity of benefits or simply places AI in the hands of authoritarian powers, while building it too fast is reckless. We have sought a middle way: to show that it’s possible to build carefully and succeed commercially, and to make safety something on which AI companies compete. In other words, to create a race to the top. We have always devoted a substantial fraction of our efforts to studying, addressing, and informing the public about these AI risks, as well as advocating for well-considered regulation of AI, even when this gets us accused of hype, “doomerism”, or regulatory capture. We have tried to prioritize caution over speed and prudence over profit.

    But over the last few months, I have become convinced that fully addressing the risks requires even more prudence — not just investing in risk prevention, but pacing the rate of capabilities advancement so that risk prevention has time to keep up. We must slow the pace at which we improve the capabilities of AI models. Progress will still seem fast, and we must make wise use of the time we gain. Two things have convinced me.

    My first concern is that, since roughly this summer, AI has been advancing drastically faster, driven primarily by AI’s growing ability to build the next generation of AI. This dynamic is called recursive self-improvement, and it is starting to happen across the industry, including at Anthropic, as we and others have described. Left unchecked, it could outrun our ability to understand and control these systems, and so must be pursued very carefully, if at all."

    darioamodei.com/post/we-must-p

    #AI #GenerativeAI #AIAgents #LLMs #Cybersecurity #AISafety

  29. "Along with my co-founders and employees, I have grappled with this duality of risk and benefit since the beginning of Anthropic. Not building the technology deprives humanity of benefits or simply places AI in the hands of authoritarian powers, while building it too fast is reckless. We have sought a middle way: to show that it’s possible to build carefully and succeed commercially, and to make safety something on which AI companies compete. In other words, to create a race to the top. We have always devoted a substantial fraction of our efforts to studying, addressing, and informing the public about these AI risks, as well as advocating for well-considered regulation of AI, even when this gets us accused of hype, “doomerism”, or regulatory capture. We have tried to prioritize caution over speed and prudence over profit.

    But over the last few months, I have become convinced that fully addressing the risks requires even more prudence — not just investing in risk prevention, but pacing the rate of capabilities advancement so that risk prevention has time to keep up. We must slow the pace at which we improve the capabilities of AI models. Progress will still seem fast, and we must make wise use of the time we gain. Two things have convinced me.

    My first concern is that, since roughly this summer, AI has been advancing drastically faster, driven primarily by AI’s growing ability to build the next generation of AI. This dynamic is called recursive self-improvement, and it is starting to happen across the industry, including at Anthropic, as we and others have described. Left unchecked, it could outrun our ability to understand and control these systems, and so must be pursued very carefully, if at all."

    darioamodei.com/post/we-must-p

    #AI #GenerativeAI #AIAgents #LLMs #Cybersecurity #AISafety

  30. The Joe Benton voice reel interview that's repeating on BBC WS is worth listening to. He is clearly very informed and articulates the issues well. Here's his substack piece.

    #AI #AIRegulation #aisafety #joebenton #anthropic

    jbenton1.substack.com/p/why-i-

  31. The Joe Benton voice reel interview that's repeating on BBC WS is worth listening to. He is clearly very informed and articulates the issues well. Here's his substack piece.

    #AI #AIRegulation #aisafety #joebenton #anthropic

    jbenton1.substack.com/p/why-i-

  32. The Joe Benton voice reel interview that's repeating on BBC WS is worth listening to. He is clearly very informed and articulates the issues well. Here's his substack piece.

    #AI #AIRegulation #aisafety #joebenton #anthropic

    jbenton1.substack.com/p/why-i-

  33. #OpenAI CEO #SamAltman confirmed the company will not go public in 2026, citing concerns about #AIsafety and the need for #industry and #government #collaboration. Altman suggested OpenAI and other AI companies may announce a pact to slow #AIdevelopment and address #safetyrisks. He emphasised the importance of prioritising safety over rushing to market. fortune.com/2026/09/12/sam-alt #tech #news #ainews

  34. #OpenAI CEO #SamAltman confirmed the company will not go public in 2026, citing concerns about #AIsafety and the need for #industry and #government #collaboration. Altman suggested OpenAI and other AI companies may announce a pact to slow #AIdevelopment and address #safetyrisks. He emphasised the importance of prioritising safety over rushing to market. fortune.com/2026/09/12/sam-alt #tech #news #ainews

  35. #OpenAI CEO #SamAltman confirmed the company will not go public in 2026, citing concerns about #AIsafety and the need for #industry and #government #collaboration. Altman suggested OpenAI and other AI companies may announce a pact to slow #AIdevelopment and address #safetyrisks. He emphasised the importance of prioritising safety over rushing to market. fortune.com/2026/09/12/sam-alt #tech #news #ainews

  36. #OpenAI CEO #SamAltman confirmed the company will not go public in 2026, citing concerns about #AIsafety and the need for #industry and #government #collaboration. Altman suggested OpenAI and other AI companies may announce a pact to slow #AIdevelopment and address #safetyrisks. He emphasised the importance of prioritising safety over rushing to market. fortune.com/2026/09/12/sam-alt #tech #news #ainews

  37. #OpenAI CEO #SamAltman confirmed the company will not go public in 2026, citing concerns about #AIsafety and the need for #industry and #government #collaboration. Altman suggested OpenAI and other AI companies may announce a pact to slow #AIdevelopment and address #safetyrisks. He emphasised the importance of prioritising safety over rushing to market. fortune.com/2026/09/12/sam-alt #tech #news #ainews

  38. My AI slow down proposal – everyone calling for it must demonstrate their concern by taking time to write harnesses to simplify full compliance with the EU's AI Act for anyone deploying their tech (including it in their own products or services.)

    #AIAct #AISafety #AISlowdown #AIEthics #AIRegulation