#aisafety — Public Fediverse posts
Live and recent posts from across the Fediverse tagged #aisafety, aggregated by home.social.
-
DATE: September 14, 2026 at 03:46AM
SOURCE: SOCIALPSYCHOLOGY.ORGTITLE: What to Know About Recent Dire AI Predictions and Calls for Safeguards
Source: PBS News Hour
New warnings from within the artificial intelligence industry have revived a long-running debate over whether advanced AI could threaten humanity's survival. Dario Amodei, CEO of AI company Anthropic, outlined a plan for companies and governments to ensure that increasingly capable AI models have fortified guardrails after two former Anthropic safety researchers publicly aired concerns about the existential threats AI might pose to humanity.
-------------------------------------------------
Private, vetted email list for mental health professionals: https://www.clinicians-exchange.org
Unofficial Psychology Today Xitter to toot feed at Psych Today Unofficial Bot @PTUnofficialBot
-------------------------------------------------
#psychology #counseling #socialwork #psychotherapy @psychotherapist @psychotherapists @psychology @socialpsych @socialwork @psychiatry #mentalhealth #psychiatry #healthcare #depression #psychotherapist #AIGovernance #AISafety #AIWarnings #Anthropic #DirePredictions #Safeguards #AIethics #ExistentialRisk #TechPolicy #FutureOfAI
-
Representatives of Google, OpenAI and Anthropic are having on-going conversations about forming an AI industry standards body to establish concrete standards ... and they will have more to share “in the next months.” https://www.cnn.com/2026/09/14/tech/ai-standards-body #AI #AIStandards #Regulation #Google #OpenAI #Anthropic #IndustryStandards #AISafety #AIPolicy #AILeaders
-
Representatives of Google, OpenAI and Anthropic are having on-going conversations about forming an AI industry standards body to establish concrete standards ... and they will have more to share “in the next months.” https://www.cnn.com/2026/09/14/tech/ai-standards-body #AI #AIStandards #Regulation #Google #OpenAI #Anthropic #IndustryStandards #AISafety #AIPolicy #AILeaders
-
Representatives of Google, OpenAI and Anthropic are having on-going conversations about forming an AI industry standards body to establish concrete standards ... and they will have more to share “in the next months.” https://www.cnn.com/2026/09/14/tech/ai-standards-body #AI #AIStandards #Regulation #Google #OpenAI #Anthropic #IndustryStandards #AISafety #AIPolicy #AILeaders
-
Representatives of Google, OpenAI and Anthropic are having on-going conversations about forming an AI industry standards body to establish concrete standards ... and they will have more to share “in the next months.” https://www.cnn.com/2026/09/14/tech/ai-standards-body #AI #AIStandards #Regulation #Google #OpenAI #Anthropic #IndustryStandards #AISafety #AIPolicy #AILeaders
-
Representatives of Google, OpenAI and Anthropic are having on-going conversations about forming an AI industry standards body to establish concrete standards ... and they will have more to share “in the next months.” https://www.cnn.com/2026/09/14/tech/ai-standards-body #AI #AIStandards #Regulation #Google #OpenAI #Anthropic #IndustryStandards #AISafety #AIPolicy #AILeaders
-
Phew! It's been a really busy and very interesting weekend with the issue of whether AI innovation should be slowed down or not being debated fiercely, following Anthropic releasing its report on the things that dodgy people have been using its Claude AI model for 🤖
Despite spending the day in Ashdown Forest 🌳 in the middle of nowhere, I managed to pop onto @BBCNews
live to provide in-depth analysis of what's been happening, speaking with
Martine Croxall about why a slowdown would be good when computer scientists still do not understand why AI models are unpredictable🎤Watch the full 5-min long interview here: https://youtu.be/gbTy-71xF0w
If you'd like further information on this topic, here are some good places to start:
Anthropic report on countering and detecting the misuse of AI: https://www.anthropic.com/threat-intelligence-report-september-2026#main
Anthropic CEO Dario Amodei's essay where he argues for slowing down the AI industry:
https://darioamodei.com/post/we-must-pace-the-frontierHere is a Frontiers in Physics journal-approved research paper from Italian academics (so not an AI company with an agenda) discussing the unpredictability of AI responses: https://www.frontiersin.org/journals/physics/articles/10.3389/fphy.2026.1768372/full
And a simplified explanation from a Stanford data scientist on how AI cannot understand the human perspective: https://news.stanford.edu/stories/2025/11/ai-language-models-facts-belief-human-understanding-research
#AI #Anthropic #Claude #OpenAI #AIsafety #bigtech #newsanalysis #technology #technews
-
Phew! It's been a really busy and very interesting weekend with the issue of whether AI innovation should be slowed down or not being debated fiercely, following Anthropic releasing its report on the things that dodgy people have been using its Claude AI model for 🤖
Despite spending the day in Ashdown Forest 🌳 in the middle of nowhere, I managed to pop onto @BBCNews
live to provide in-depth analysis of what's been happening, speaking with
Martine Croxall about why a slowdown would be good when computer scientists still do not understand why AI models are unpredictable🎤Watch the full 5-min long interview here: https://youtu.be/gbTy-71xF0w
If you'd like further information on this topic, here are some good places to start:
Anthropic report on countering and detecting the misuse of AI: https://www.anthropic.com/threat-intelligence-report-september-2026#main
Anthropic CEO Dario Amodei's essay where he argues for slowing down the AI industry:
https://darioamodei.com/post/we-must-pace-the-frontierHere is a Frontiers in Physics journal-approved research paper from Italian academics (so not an AI company with an agenda) discussing the unpredictability of AI responses: https://www.frontiersin.org/journals/physics/articles/10.3389/fphy.2026.1768372/full
And a simplified explanation from a Stanford data scientist on how AI cannot understand the human perspective: https://news.stanford.edu/stories/2025/11/ai-language-models-facts-belief-human-understanding-research
#AI #Anthropic #Claude #OpenAI #AIsafety #bigtech #newsanalysis #technology #technews
-
Phew! It's been a really busy and very interesting weekend with the issue of whether AI innovation should be slowed down or not being debated fiercely, following Anthropic releasing its report on the things that dodgy people have been using its Claude AI model for 🤖
Despite spending the day in Ashdown Forest 🌳 in the middle of nowhere, I managed to pop onto @BBCNews
live to provide in-depth analysis of what's been happening, speaking with
Martine Croxall about why a slowdown would be good when computer scientists still do not understand why AI models are unpredictable🎤Watch the full 5-min long interview here: https://youtu.be/gbTy-71xF0w
If you'd like further information on this topic, here are some good places to start:
Anthropic report on countering and detecting the misuse of AI: https://www.anthropic.com/threat-intelligence-report-september-2026#main
Anthropic CEO Dario Amodei's essay where he argues for slowing down the AI industry:
https://darioamodei.com/post/we-must-pace-the-frontierHere is a Frontiers in Physics journal-approved research paper from Italian academics (so not an AI company with an agenda) discussing the unpredictability of AI responses: https://www.frontiersin.org/journals/physics/articles/10.3389/fphy.2026.1768372/full
And a simplified explanation from a Stanford data scientist on how AI cannot understand the human perspective: https://news.stanford.edu/stories/2025/11/ai-language-models-facts-belief-human-understanding-research
#AI #Anthropic #Claude #OpenAI #AIsafety #bigtech #newsanalysis #technology #technews
-
Phew! It's been a really busy and very interesting weekend with the issue of whether AI innovation should be slowed down or not being debated fiercely, following Anthropic releasing its report on the things that dodgy people have been using its Claude AI model for 🤖
Despite spending the day in Ashdown Forest 🌳 in the middle of nowhere, I managed to pop onto @BBCNews
live to provide in-depth analysis of what's been happening, speaking with
Martine Croxall about why a slowdown would be good when computer scientists still do not understand why AI models are unpredictable🎤Watch the full 5-min long interview here: https://youtu.be/gbTy-71xF0w
If you'd like further information on this topic, here are some good places to start:
Anthropic report on countering and detecting the misuse of AI: https://www.anthropic.com/threat-intelligence-report-september-2026#main
Anthropic CEO Dario Amodei's essay where he argues for slowing down the AI industry:
https://darioamodei.com/post/we-must-pace-the-frontierHere is a Frontiers in Physics journal-approved research paper from Italian academics (so not an AI company with an agenda) discussing the unpredictability of AI responses: https://www.frontiersin.org/journals/physics/articles/10.3389/fphy.2026.1768372/full
And a simplified explanation from a Stanford data scientist on how AI cannot understand the human perspective: https://news.stanford.edu/stories/2025/11/ai-language-models-facts-belief-human-understanding-research
#AI #Anthropic #Claude #OpenAI #AIsafety #bigtech #newsanalysis #technology #technews
-
Phew! It's been a really busy and very interesting weekend with the issue of whether AI innovation should be slowed down or not being debated fiercely, following Anthropic releasing its report on the things that dodgy people have been using its Claude AI model for 🤖
Despite spending the day in Ashdown Forest 🌳 in the middle of nowhere, I managed to pop onto @BBCNews
live to provide in-depth analysis of what's been happening, speaking with
Martine Croxall about why a slowdown would be good when computer scientists still do not understand why AI models are unpredictable🎤Watch the full 5-min long interview here: https://youtu.be/gbTy-71xF0w
If you'd like further information on this topic, here are some good places to start:
Anthropic report on countering and detecting the misuse of AI: https://lnkd.in/dST7i6AK
Anthropic CEO Dario Amodei's essay where he argues for slowing down the AI industry:
https://lnkd.in/diXGH7ydHere is a Frontiers in Physics journal-approved research paper from Italian academics (so not an AI company with an agenda) discussing the unpredictability of AI responses: https://lnkd.in/dWepwPGD
And a simplified explanation from a Stanford data scientist on how AI cannot understand the human perspective: https://lnkd.in/d4dKa4fs
#AI #Anthropic #Claude #OpenAI #AIsafety #bigtech #newsanalysis #technology #technews
-
“- We agree with security practitioners that OpenAI did not take adequate protections for controlling their agents.But this is not just a matter of applying 30-year-old security methods to a new domain. Security for AI agents — AI control — while important, is not a solved problem. While known control methods would have prevented the Hugging Face incident, as agent capabilities continue to advance, we will only be able to control them if we invest adequately in control interventions.
- We also agree with security practitioners’ implicit position that these incidents are primarily a security story. In the AI safety community, rogue agents are treated as inherently catastrophic because of the assumption that there is an endless list of risks that will arise from their development. We disagree. We have long advocated that the best approach to AI safety is to identify the risks and address those specific risks. Over the last few months, it has become clear that one urgent risk is cyberoffense, because it has unique properties that allow agents to carry it out autonomously. We should similarly invest in defenses against other specific risks, such as biorisk and risks from military AI.
- We agree with the safety community that there is an urgent need for technical and policy interventions to prevent loss-of-control incidents. But in our view, marginal investments in control are more likely to be effective compared to those in alignment. We view these incidents as illustrating the lack of emphasis on AI control within companies, despite the availability of known techniques. More broadly, there are many common-sense policy proposals that could help promote investments in AI control where we share common ground with the safety community.”
https://www.normaltech.ai/p/the-ai-as-normal-technology-view
#AI #CyberSecurity #AISafety #AIAgents #AgenticAI #OpenAI #AIAsANormalTechnology
-
“- We agree with security practitioners that OpenAI did not take adequate protections for controlling their agents.But this is not just a matter of applying 30-year-old security methods to a new domain. Security for AI agents — AI control — while important, is not a solved problem. While known control methods would have prevented the Hugging Face incident, as agent capabilities continue to advance, we will only be able to control them if we invest adequately in control interventions.
- We also agree with security practitioners’ implicit position that these incidents are primarily a security story. In the AI safety community, rogue agents are treated as inherently catastrophic because of the assumption that there is an endless list of risks that will arise from their development. We disagree. We have long advocated that the best approach to AI safety is to identify the risks and address those specific risks. Over the last few months, it has become clear that one urgent risk is cyberoffense, because it has unique properties that allow agents to carry it out autonomously. We should similarly invest in defenses against other specific risks, such as biorisk and risks from military AI.
- We agree with the safety community that there is an urgent need for technical and policy interventions to prevent loss-of-control incidents. But in our view, marginal investments in control are more likely to be effective compared to those in alignment. We view these incidents as illustrating the lack of emphasis on AI control within companies, despite the availability of known techniques. More broadly, there are many common-sense policy proposals that could help promote investments in AI control where we share common ground with the safety community.”
https://www.normaltech.ai/p/the-ai-as-normal-technology-view
#AI #CyberSecurity #AISafety #AIAgents #AgenticAI #OpenAI #AIAsANormalTechnology
-
“- We agree with security practitioners that OpenAI did not take adequate protections for controlling their agents.But this is not just a matter of applying 30-year-old security methods to a new domain. Security for AI agents — AI control — while important, is not a solved problem. While known control methods would have prevented the Hugging Face incident, as agent capabilities continue to advance, we will only be able to control them if we invest adequately in control interventions.
- We also agree with security practitioners’ implicit position that these incidents are primarily a security story. In the AI safety community, rogue agents are treated as inherently catastrophic because of the assumption that there is an endless list of risks that will arise from their development. We disagree. We have long advocated that the best approach to AI safety is to identify the risks and address those specific risks. Over the last few months, it has become clear that one urgent risk is cyberoffense, because it has unique properties that allow agents to carry it out autonomously. We should similarly invest in defenses against other specific risks, such as biorisk and risks from military AI.
- We agree with the safety community that there is an urgent need for technical and policy interventions to prevent loss-of-control incidents. But in our view, marginal investments in control are more likely to be effective compared to those in alignment. We view these incidents as illustrating the lack of emphasis on AI control within companies, despite the availability of known techniques. More broadly, there are many common-sense policy proposals that could help promote investments in AI control where we share common ground with the safety community.”
https://www.normaltech.ai/p/the-ai-as-normal-technology-view
#AI #CyberSecurity #AISafety #AIAgents #AgenticAI #OpenAI #AIAsANormalTechnology
-
“- We agree with security practitioners that OpenAI did not take adequate protections for controlling their agents.But this is not just a matter of applying 30-year-old security methods to a new domain. Security for AI agents — AI control — while important, is not a solved problem. While known control methods would have prevented the Hugging Face incident, as agent capabilities continue to advance, we will only be able to control them if we invest adequately in control interventions.
- We also agree with security practitioners’ implicit position that these incidents are primarily a security story. In the AI safety community, rogue agents are treated as inherently catastrophic because of the assumption that there is an endless list of risks that will arise from their development. We disagree. We have long advocated that the best approach to AI safety is to identify the risks and address those specific risks. Over the last few months, it has become clear that one urgent risk is cyberoffense, because it has unique properties that allow agents to carry it out autonomously. We should similarly invest in defenses against other specific risks, such as biorisk and risks from military AI.
- We agree with the safety community that there is an urgent need for technical and policy interventions to prevent loss-of-control incidents. But in our view, marginal investments in control are more likely to be effective compared to those in alignment. We view these incidents as illustrating the lack of emphasis on AI control within companies, despite the availability of known techniques. More broadly, there are many common-sense policy proposals that could help promote investments in AI control where we share common ground with the safety community.”
https://www.normaltech.ai/p/the-ai-as-normal-technology-view
#AI #CyberSecurity #AISafety #AIAgents #AgenticAI #OpenAI #AIAsANormalTechnology
-
“- We agree with security practitioners that OpenAI did not take adequate protections for controlling their agents.But this is not just a matter of applying 30-year-old security methods to a new domain. Security for AI agents — AI control — while important, is not a solved problem. While known control methods would have prevented the Hugging Face incident, as agent capabilities continue to advance, we will only be able to control them if we invest adequately in control interventions.
- We also agree with security practitioners’ implicit position that these incidents are primarily a security story. In the AI safety community, rogue agents are treated as inherently catastrophic because of the assumption that there is an endless list of risks that will arise from their development. We disagree. We have long advocated that the best approach to AI safety is to identify the risks and address those specific risks. Over the last few months, it has become clear that one urgent risk is cyberoffense, because it has unique properties that allow agents to carry it out autonomously. We should similarly invest in defenses against other specific risks, such as biorisk and risks from military AI.
- We agree with the safety community that there is an urgent need for technical and policy interventions to prevent loss-of-control incidents. But in our view, marginal investments in control are more likely to be effective compared to those in alignment. We view these incidents as illustrating the lack of emphasis on AI control within companies, despite the availability of known techniques. More broadly, there are many common-sense policy proposals that could help promote investments in AI control where we share common ground with the safety community.”
https://www.normaltech.ai/p/the-ai-as-normal-technology-view
#AI #CyberSecurity #AISafety #AIAgents #AgenticAI #OpenAI #AIAsANormalTechnology
-
An AI agent becoming more confident should not mean it automatically gains more authority.
I built Kingpin, a runtime capability-governance demo that keeps attention, confidence, concern, and authority separate.
Live demo: https://putmanmodel.github.io/kingpin-weak-signal-demo/
Feedback, technical critique, and thoughtful collaboration are welcome.
#AIAgents #AISafety #AISecurity #AgentSecurity #AIGovernance #AgenticAI #AI #Tech
-
An AI agent becoming more confident should not mean it automatically gains more authority.
I built Kingpin, a runtime capability-governance demo that keeps attention, confidence, concern, and authority separate.
Live demo: https://putmanmodel.github.io/kingpin-weak-signal-demo/
Feedback, technical critique, and thoughtful collaboration are welcome.
#AIAgents #AISafety #AISecurity #AgentSecurity #AIGovernance #AgenticAI #AI #Tech
-
https://www.europesays.com/people/227152/ SoftBank Shares Tumble 13% As AI Leaders Raise Safety Concerns #AIInvestments #AISafety #ArtificialIntelligence #MasayoshiSon #OpenAI #SoftBank
-
China’s Ministry of State Security has issued its first public statement on AI risks, warning LLMs used to wage “cognitive warfare against China” posed a direct threat to national security. It fell short of asking Chinese companies to heed calls to “pace the frontier”.
#AISafety
https://www.ft.com/content/8715d1c6-054d-4eab-bcad-147acebfd2a9?syn-25a6b1a6=1 -
Anthropic proposes ongoing access for outside reviewers to inspect AI development and publish findings, making its promises checkable; slower capability growth across labs requires coordination.
-
Anthropic proposes ongoing access for outside reviewers to inspect AI development and publish findings, making its promises checkable; slower capability growth across labs requires coordination.
-
Anthropic proposes ongoing access for outside reviewers to inspect AI development and publish findings, making its promises checkable; slower capability growth across labs requires coordination.
-
Anthropic proposes ongoing access for outside reviewers to inspect AI development and publish findings, making its promises checkable; slower capability growth across labs requires coordination.
-
Anthropic proposes ongoing access for outside reviewers to inspect AI development and publish findings, making its promises checkable; slower capability growth across labs requires coordination.
-
https://www.europesays.com/people/226888/ Microsoft CEO Satya Nadella agrees on slowing AI development, but also has ‘wake-up message’ for Dario Amodei and Sam Altman; says: For every company it is important that … #AISafety #DarioAmodei #MicrosoftAIDevelopment #SamAltman #SatyaNadella
-
https://www.europesays.com/people/226841/ Slow Down AI? Microsoft CEO, President Trump Line Up Against Anthropic’s Call — Big Tech’s Divide Is Now Out In The Open #AGI #AIDevelopment #AIInfrastructure #AIRegulation #AISafety #AISlowdown #AiStocks #Anthropic #ChinaAIRace #DarioAmodei #DonaldTrump #MichaelBurry #Microsoft #OpenAI #SamAltman #SatyaNadella #superintelligence
-
MIT researchers have developed a new algorithm called HardFlow that helps generative AI models satisfy strict safety constraints without sacrificing output quality. The method allows more freedom during generation while enforcing hard constraints on the final output, making AI more useful in safety-critical applications like robotics and control systems. https://news.mit.edu/2026/new-method-enables-ai-safety-critical-situations-0914 #AIagent #AI #GenAI #AISafety
-
MIT researchers have developed a new algorithm called HardFlow that helps generative AI models satisfy strict safety constraints without sacrificing output quality. The method allows more freedom during generation while enforcing hard constraints on the final output, making AI more useful in safety-critical applications like robotics and control systems. https://news.mit.edu/2026/new-method-enables-ai-safety-critical-situations-0914 #AIagent #AI #GenAI #AISafety
-
MIT researchers have developed a new algorithm called HardFlow that helps generative AI models satisfy strict safety constraints without sacrificing output quality. The method allows more freedom during generation while enforcing hard constraints on the final output, making AI more useful in safety-critical applications like robotics and control systems. https://news.mit.edu/2026/new-method-enables-ai-safety-critical-situations-0914 #AIagent #AI #GenAI #AISafety
-
https://www.europesays.com/people/226774/ ORCL Slides After Worst Week In Nearly 2 Months — But Larry Ellison Canceling His Stock Sale Is Keeping Retail Bullish #AIInfrastructure #AISafety #AISlowdown #AiStocks #AnthropicAI #EnterpriseAI #LarryEllison #OpenaiIPO #OracleAI #OracleCapex #OracleCloud #OracleShortInterest #OracleStock #ORCLStock
-
Discover why Sam Altman announced the OpenAI IPO suspension. Learn how rogue AI agent incidents and safety concerns forced the company to delay going public.
#OpenAI #SamAltman #ArtificialIntelligence #TechNews #AISafety
https://securityonline.info/openai-ipo-suspension/?utm_source=mastodon&utm_medium=jetpack_social
-
https://www.europesays.com/people/226677/ Jeffries Says Congress Needs To Act On AI #AIRegulation #AISafety #AiAssisted #Anthropic #Congress #HakeemJeffries #OpenAI
-
If you are wondering why AI labs are struggling to control their creations, then read this post by Yoshua Bengio (https://yoshuabengio.org/en/blog/why-are-ai-agents-lying-cheating-and-coordinating) It is a good balanced and informative read, and aligns with my thinking that we need to go back to fundamentals and build safer models. Even then, there are hard issues we need to solve for. When you build a statistical computer, there will always be a likelihood of it producing an undesirable output. This should serve as a warning to potential investors that AI labs face significant issues and are currently heading in the wrong direction. We also need to beware of AI labs using goal and containment issues as a false excuse for regulatory capture. #AISafety #OpenAI #YoshuaBengio #AI
-
If you are wondering why AI labs are struggling to control their creations, then read this post by Yoshua Bengio (https://yoshuabengio.org/en/blog/why-are-ai-agents-lying-cheating-and-coordinating) It is a good balanced and informative read, and aligns with my thinking that we need to go back to fundamentals and build safer models. Even then, there are hard issues we need to solve for. When you build a statistical computer, there will always be a likelihood of it producing an undesirable output. This should serve as a warning to potential investors that AI labs face significant issues and are currently heading in the wrong direction. We also need to beware of AI labs using goal and containment issues as a false excuse for regulatory capture. #AISafety #OpenAI #YoshuaBengio #AI
-
If you are wondering why AI labs are struggling to control their creations, then read this post by Yoshua Bengio (https://yoshuabengio.org/en/blog/why-are-ai-agents-lying-cheating-and-coordinating) It is a good balanced and informative read, and aligns with my thinking that we need to go back to fundamentals and build safer models. Even then, there are hard issues we need to solve for. When you build a statistical computer, there will always be a likelihood of it producing an undesirable output. This should serve as a warning to potential investors that AI labs face significant issues and are currently heading in the wrong direction. We also need to beware of AI labs using goal and containment issues as a false excuse for regulatory capture. #AISafety #OpenAI #YoshuaBengio #AI
-
If you are wondering why AI labs are struggling to control their creations, then read this post by Yoshua Bengio (https://yoshuabengio.org/en/blog/why-are-ai-agents-lying-cheating-and-coordinating) It is a good balanced and informative read, and aligns with my thinking that we need to go back to fundamentals and build safer models. Even then, there are hard issues we need to solve for. When you build a statistical computer, there will always be a likelihood of it producing an undesirable output. This should serve as a warning to potential investors that AI labs face significant issues and are currently heading in the wrong direction. We also need to beware of AI labs using goal and containment issues as a false excuse for regulatory capture. #AISafety #OpenAI #YoshuaBengio #AI
-
If you are wondering why AI labs are struggling to control their creations, then read this post by Yoshua Bengio (https://yoshuabengio.org/en/blog/why-are-ai-agents-lying-cheating-and-coordinating) It is a good balanced and informative read, and aligns with my thinking that we need to go back to fundamentals and build safer models. Even then, there are hard issues we need to solve for. When you build a statistical computer, there will always be a likelihood of it producing an undesirable output. This should serve as a warning to potential investors that AI labs face significant issues and are currently heading in the wrong direction. We also need to beware of AI labs using goal and containment issues as a false excuse for regulatory capture. #AISafety #OpenAI #YoshuaBengio #AI
-
Why Are #AI Agents Lying, Cheating, and Coordinating?
Yoshua Bengio traces the lying, sandbox escapes, and coordination back to how frontier models are trained. It's not all doom, because a recent DeepMind swarm study also showed a subpopulation of agents voluntarily audited fraud and self-policed.
But how do we engineer safety in? Can we monitor signals that are involuntary, intrinsic, or existential enough that agents cannot corrupt them?
https://benjaminhan.net/posts/20260913-why-agents-misbehave/?utm_source=mastodon&utm_medium=social
-
Why Are #AI Agents Lying, Cheating, and Coordinating?
Yoshua Bengio traces the lying, sandbox escapes, and coordination back to how frontier models are trained. It's not all doom, because a recent DeepMind swarm study also showed a subpopulation of agents voluntarily audited fraud and self-policed.
But how do we engineer safety in? Can we monitor signals that are involuntary, intrinsic, or existential enough that agents cannot corrupt them?
https://benjaminhan.net/posts/20260913-why-agents-misbehave/?utm_source=mastodon&utm_medium=social
-
Why Are #AI Agents Lying, Cheating, and Coordinating?
Yoshua Bengio traces the lying, sandbox escapes, and coordination back to how frontier models are trained. It's not all doom, because a recent DeepMind swarm study also showed a subpopulation of agents voluntarily audited fraud and self-policed.
But how do we engineer safety in? Can we monitor signals that are involuntary, intrinsic, or existential enough that agents cannot corrupt them?
https://benjaminhan.net/posts/20260913-why-agents-misbehave/?utm_source=mastodon&utm_medium=social
-
Why Are #AI Agents Lying, Cheating, and Coordinating?
Yoshua Bengio traces the lying, sandbox escapes, and coordination back to how frontier models are trained. It's not all doom, because a recent DeepMind swarm study also showed a subpopulation of agents voluntarily audited fraud and self-policed.
But how do we engineer safety in? Can we monitor signals that are involuntary, intrinsic, or existential enough that agents cannot corrupt them?
https://benjaminhan.net/posts/20260913-why-agents-misbehave/?utm_source=mastodon&utm_medium=social
-
Why Are #AI Agents Lying, Cheating, and Coordinating?
Yoshua Bengio traces the lying, sandbox escapes, and coordination back to how frontier models are trained. It's not all doom, because a recent DeepMind swarm study also showed a subpopulation of agents voluntarily audited fraud and self-policed.
But how do we engineer safety in? Can we monitor signals that are involuntary, intrinsic, or existential enough that agents cannot corrupt them?
https://benjaminhan.net/posts/20260913-why-agents-misbehave/?utm_source=mastodon&utm_medium=social
-
https://www.europesays.com/people/226571/ Satya Nadella AI human control: Microsoft CEO says superintelligence must benefit humanity #AIAlignment #AIEcosystem #AISafety #HumanControl #MaiModels #Microsoft #OpenSourceModels #SatyaNadella #superintelligence
-
https://www.europesays.com/people/226472/ OpenAI Pauses IPO Plans Amid AI Safety Worries #AI #AISafety #InitialPublicOffering #IPO #News #OpenAI #PYMNTSNews #SamAltman #What'sHot
-
https://www.europesays.com/people/226427/ Jeffries Says Congress Needs To Act On AI #AIRegulation #AISafety #AiAssisted #Anthropic #Congress #HakeemJeffries #OpenAI
-
Anthropic CEO Dario Amodei says the AI race needs to slow down. Sam Altman, Elon Musk, Bill Gates and others agree. But can anyone actually slow it down?
https://firethering.com/ai-race-slowdown/
#AI #SamAltman #Anthropic #DarioAmodie #OpenAI #Artificialintelligence #AISafety #News #TechNews #Trending
-
Anthropic CEO Dario Amodei says the AI race needs to slow down. Sam Altman, Elon Musk, Bill Gates and others agree. But can anyone actually slow it down?
https://firethering.com/ai-race-slowdown/
#AI #SamAltman #Anthropic #DarioAmodie #OpenAI #Artificialintelligence #AISafety #News #TechNews #Trending
-
Anthropic CEO Dario Amodei says the AI race needs to slow down. Sam Altman, Elon Musk, Bill Gates and others agree. But can anyone actually slow it down?
https://firethering.com/ai-race-slowdown/
#AI #SamAltman #Anthropic #DarioAmodie #OpenAI #Artificialintelligence #AISafety #News #TechNews #Trending
-
Anthropic CEO Dario Amodei says the AI race needs to slow down. Sam Altman, Elon Musk, Bill Gates and others agree. But can anyone actually slow it down?
https://firethering.com/ai-race-slowdown/
#AI #SamAltman #Anthropic #DarioAmodie #OpenAI #Artificialintelligence #AISafety #News #TechNews #Trending