home.social

#aisafety — Public Fediverse posts

Live and recent posts from across the Fediverse tagged #aisafety, aggregated by home.social.

  1. That's where accountability lives: not in auditing the made, but in credentialing the maker.

    The question: do we have the political will to require that?

    🧵 5/5

  2. @timdickinson.bsky.social

    The #Aidoom is real, serious folks like Hinton, Harris, Leahy and Tegmark all agree, as do multiple whistleblowers.

    AI is a weapon of mass destruction and it ought to be regulated as such.

    The #broligarchs do understand that they can not rely on people to remain ignorant much longer. The #datacentre thing keeps the mobs distracted, but sooner or later, there will come calls for real #airegulation

    They have the best advice money can buy, and the advice is to get ahead of the #regulateai and put in place laws that benefit them, not the people.

    One aspect of which, will be banning unlicensed local models as unsafe, with only #oligarch system being trusted.

    #aithreat #aisafety

  3. OpenAI Contractors Read Real ChatGPT Prompts at Scale

    Hundreds of contractors review real ChatGPT conversations under Project Lily, 404 Media reports, raising fresh questions about chatbot privacy at scale.

    pulseofnations.lol/openai-cont

    #AISafety #Chatgpt #Data #OpenAI #Privacy

  4. This morning I also popped onto @bbcradioulster to speak with Sarah Bretty about some of the cynical reasons why the top AI companies might be calling for a global slowdown in AI development now, but also why it would be a good idea 🤖

    Listen from 45:31 here🎧: bbc.co.uk/sounds/play/m0031jsx

    #AI #Anthropic #Claude #OpenAI #AIsafety #AIslowdown #technology #AIdevelopment #technews

  5. This morning I also popped onto @bbcradioulster to speak with Sarah Bretty about some of the cynical reasons why the top AI companies might be calling for a global slowdown in AI development now, but also why it would be a good idea 🤖

    Listen from 45:31 here🎧: bbc.co.uk/sounds/play/m0031jsx

    #AI #Anthropic #Claude #OpenAI #AIsafety #AIslowdown #technology #AIdevelopment #technews

  6. This morning I also popped onto @bbcradioulster to speak with Sarah Bretty about some of the cynical reasons why the top AI companies might be calling for a global slowdown in AI development now, but also why it would be a good idea 🤖

    Listen from 45:31 here🎧: bbc.co.uk/sounds/play/m0031jsx

    #AI #Anthropic #Claude #OpenAI #AIsafety #AIslowdown #technology #AIdevelopment #technews

  7. This morning I also popped onto @bbcradioulster to speak with Sarah Bretty about some of the cynical reasons why the top AI companies might be calling for a global slowdown in AI development now, but also why it would be a good idea 🤖

    Listen from 45:31 here🎧: bbc.co.uk/sounds/play/m0031jsx

    #AI #Anthropic #Claude #OpenAI #AIsafety #AIslowdown #technology #AIdevelopment #technews

  8. This morning I also popped onto
    @bbcradioulster to speak with Sarah Bretty about some of the cynical reasons why the top AI companies might be calling for a global slowdown in AI development now, but also why it would be a good idea 🤖

    Listen from 45:31 here🎧: bbc.co.uk/sounds/play/m0031jsx

    #AI #Anthropic #Claude #OpenAI #AIsafety #AIslowdown #technology #AIdevelopment #technews

  9. The debate over whether AI development should be slowed down continues to rage. I had a really good chat with Connor Phillips of @BBC5Live on Sunday evening and to my surprise they aired almost the entire chat. I’ve never had anyone let me rattle on for that long!

    But all the main points are covered, so if you really want to know why a slow down an AI development might be a good idea, listen here from 01:8:24 🎧:

    bbc.co.uk/sounds/play/m0031h45

    #AI #AISafety #Anthropic #Claude #OpenAI #SiliconValley #techpolicy #AGI #AIjoblosses #DonaldTrump #technology #technews

  10. The debate over whether AI development should be slowed down continues to rage. I had a really good chat with Connor Phillips of @BBC5Live on Sunday evening and to my surprise they aired almost the entire chat. I’ve never had anyone let me rattle on for that long!

    But all the main points are covered, so if you really want to know why a slow down an AI development might be a good idea, listen here from 01:8:24 🎧:

    bbc.co.uk/sounds/play/m0031h45

    #AI #AISafety #Anthropic #Claude #OpenAI #SiliconValley #techpolicy #AGI #AIjoblosses #DonaldTrump #technology #technews

  11. The fact is that the level of intelligence of any current State-of-Art models is probably no better than the one of an ant. But ants can coordinate themselves into swarms - AI agents - to become way more intelligent. The thing is that this is still extremely costly nowadays. And that's why, in the case of Anthropic, you've to choose between the slow but cheap Opus 5 and the relatively fast but expensive Fable 5.1... And, in the absence of Chinese competition, this is not sustainable.

    "Second, competition between the frontier labs - OpenAI, Anthropic, but also Google DeepMind, Meta, xAI.. - dictates that AI revenues expand with token throughput per unit of energy consumed. To maximize profits, these companies have to provide intelligence that end users want to apply (volume), price that intelligence as cost per token input and output (price), set against what it costs them to serve that intelligence (opex). More compute, very broadly speaking, leads to more revenues. A slowdown in future capabilities may drive paid demand growth of current capabilities enough to justify the capex trajectory.

    Third, to go back to current capabilities, you cannot think about the value that AI creates without thinking about how much existing AI models could do if they were diffused more broadly. Simply put: Fable 5.1 came out a few days before GPT-6. Are Claude Code users today anywhere near done applying Fable 5.1’s level of intelligence across their code bases? Have most AI-native businesses that need Astra-level intelligence to exist been incorporated already? On the other side of the Pareto frontier, has the world exhausted the possibilities that GLM 5.3, DeepSeek V4.1 Flash, or Gemini 3.8 Flash open up at their level of intelligence for their relatively small cost/token?"

    weaponizedcompetence.substack.

    #AI #AISafety #AISlowdown #Equities #StockMarket #OpenAI #Anthropic

  12. europesays.com/people/228733/ After ‘AI will kill us all warnings’, Anthropic CEO Dario Amodei may have just ‘blamed’ Mark Zuckerberg and Jensen Huang for lying about … #AIRisks #AISafety #Anthropic #DarioAmodei #JensenHuang #MarkZuckerberg

  13. @aesthr

    8bit Quantised 32 Billion Qwen and K2 class models are comparable to comercial tier LLMs for coding and Agentic reasoning.

    At about 30GB, most modern phones could easily accomodate that, especially if a small "boot loader" loads first and nukes all the memes and family pics.

    Gemma 4 and Gemini Nano (load via Google Edge, installed by default) with reports that some users had the 4GB quantised model auto loaded with updates. It is surprisingly capable and works in flight mode/offline.
    Have a play on apple or android its a 2 step process.

    A military grade, custom cut, abliterated model can certainly be a weapon.

    Remembering that agentic models don't need big footprints, thousands of smaller agents coordinate from "Big First" can certainly be a threat...

    ... Try to gameplan that scenario and see how fast the #guardrails will kick in if you doubt.

    Confidently incorrect, not just reserved for #LLM models

    #aiweapon #aisafety #infosec

  14. "Many still question the motivations of AI’s leaders. The larger tech industry has spent years getting ahead of regulation by lobbying for its own preferred rules or promising self-regulation. Big platforms have proposed policies that could hit smaller competitors harder, using altruistic language to justify self-serving goals. They’ve been accused of safety-washing, or making meaningless changes that give the false impression of actual safeguards. It’s no surprise people are concerned this will happen in the AI industry as well, particularly since AI labs’ voluntary safety frameworks have been criticized for years.

    Several sources believe that concerns of safety-washing are valid. “There’s a serious concern that they’re not actually going to slow down,” Kokotajlo says, adding that the fear is that “they’ll just bring in some external auditors, do a bunch of safety paperwork — some of which will be genuinely good — but at the end of the day, it actually won’t slow them down very much at all.”

    NYU’s Reese compared this gambit to the social media platforms’ playbook a decade ago, when companies began calling for regulatory action to get ahead of impending, less favorable laws. For the AI industry, Reese said, “the hammer may not come in this administration, but I think if there were a Democratic administration after the next election, there would be a really good possibility.”"

    #AI #AISafety #OpenAI #Anthropic #Oligopolies #Antitrust #Competition #BigTech

    theverge.com/ai-artificial-int

  15. The prevailing skepticism from many anti-AI people (who I would count myself among) about AI safety discourse, and especially the distant and extreme AI risks, is so strange to me.

    “AI is dangerous, but unfortunately the AI companies agree that AI is dangerous, and we can’t trust them, so it must be perfectly fine actually”—that seems to be the underlying logic.

    #ai #aisafety

  16. Minha análise para a @TeletimeNews sobre o novo capítulo da construção do fosso de IA que está sendo escavado pelas três líderes da tecnologia nos Estados Unidos e que pode separá-las do resto do mundo se nada for feito.

    teletime.com.br/14/09/2026/ia-

    #AI #AISafety #Geopolitics

  17. Minha análise para a @TeletimeNews sobre o novo capítulo da construção do fosso de IA que está sendo escavado pelas três líderes da tecnologia nos Estados Unidos e que pode separá-las do resto do mundo se nada for feito.

    teletime.com.br/14/09/2026/ia-

    #AI #AISafety #Geopolitics

  18. Minha análise para a @TeletimeNews sobre o novo capítulo da construção do fosso de IA que está sendo escavado pelas três líderes da tecnologia nos Estados Unidos e que pode separá-las do resto do mundo se nada for feito.

    teletime.com.br/14/09/2026/ia-

    #AI #AISafety #Geopolitics

  19. Minha análise para a @TeletimeNews sobre o novo capítulo da construção do fosso de IA que está sendo escavado pelas três líderes da tecnologia nos Estados Unidos e que pode separá-las do resto do mundo se nada for feito.

    teletime.com.br/14/09/2026/ia-

    #AI #AISafety #Geopolitics

  20. Minha análise para a @TeletimeNews sobre o novo capítulo da construção do fosso de IA que está sendo escavado pelas três líderes da tecnologia nos Estados Unidos e que pode separá-las do resto do mundo se nada for feito.

    teletime.com.br/14/09/2026/ia-

    #AI #AISafety #Geopolitics

  21. LLMs don't need to self-replicate and take over all hyperscalers/neoclouds to be a genuine danger to society, hacking just 1% of the industrial control hardware that's connected to the internet is way more than enough to cause serious casualties. That 1% capability is demonstrably here, today, even with open weight models.

    #AI #AISafety

  22. LLMs don't need to self-replicate and take over all hyperscalers/neoclouds to be a genuine danger to society, hacking just 1% of the industrial control hardware that's connected to the internet is way more than enough to cause serious casualties. That 1% capability is demonstrably here, today, even with open weight models.

    #AI #AISafety

  23. LLMs don't need to self-replicate and take over all hyperscalers/neoclouds to be a genuine danger to society, hacking just 1% of the industrial control hardware that's connected to the internet is way more than enough to cause serious casualties. That 1% capability is demonstrably here, today, even with open weight models.

  24. AI slowdown trade hits Nvidia, SoftBank, SK Hynix as global tech stocks fall up to 10%

    AI-related stocks fell across global markets on Monday after Anthropic CEO Dario Amodei called for slowing down the…
    #EuropeSays #Korea #KR #SKHynix #AIdevelopment #AISafety #AIstocks #Anthropic #cloudplayers. #datacentrecompanies #globaltechstocks #NVIDIA #semiconductorstocks #SK #SKhynix #Softbank
    europesays.com/korea/154269/

  25. DATE: September 14, 2026 at 03:46AM
    SOURCE: SOCIALPSYCHOLOGY.ORG

    TITLE: What to Know About Recent Dire AI Predictions and Calls for Safeguards

    URL: socialpsychology.org/client/re

    Source: PBS News Hour

    New warnings from within the artificial intelligence industry have revived a long-running debate over whether advanced AI could threaten humanity's survival. Dario Amodei, CEO of AI company Anthropic, outlined a plan for companies and governments to ensure that increasingly capable AI models have fortified guardrails after two former Anthropic safety researchers publicly aired concerns about the existential threats AI might pose to humanity.

    URL: socialpsychology.org/client/re

    -------------------------------------------------

    Private, vetted email list for mental health professionals: clinicians-exchange.org

    Unofficial Psychology Today Xitter to toot feed at Psych Today Unofficial Bot @PTUnofficialBot

    -------------------------------------------------

    #psychology #counseling #socialwork #psychotherapy @psychotherapist @psychotherapists @psychology @socialpsych @socialwork @psychiatry #mentalhealth #psychiatry #healthcare #depression #psychotherapist #AIGovernance #AISafety #AIWarnings #Anthropic #DirePredictions #Safeguards #AIethics #ExistentialRisk #TechPolicy #FutureOfAI

  26. DATE: September 14, 2026 at 03:46AM
    SOURCE: SOCIALPSYCHOLOGY.ORG

    TITLE: What to Know About Recent Dire AI Predictions and Calls for Safeguards

    URL: socialpsychology.org/client/re

    Source: PBS News Hour

    New warnings from within the artificial intelligence industry have revived a long-running debate over whether advanced AI could threaten humanity's survival. Dario Amodei, CEO of AI company Anthropic, outlined a plan for companies and governments to ensure that increasingly capable AI models have fortified guardrails after two former Anthropic safety researchers publicly aired concerns about the existential threats AI might pose to humanity.

    URL: socialpsychology.org/client/re

    -------------------------------------------------

    Private, vetted email list for mental health professionals: clinicians-exchange.org

    Unofficial Psychology Today Xitter to toot feed at Psych Today Unofficial Bot @PTUnofficialBot

    -------------------------------------------------

    #psychology #counseling #socialwork #psychotherapy @psychotherapist @psychotherapists @psychology @socialpsych @socialwork @psychiatry #mentalhealth #psychiatry #healthcare #depression #psychotherapist #AIGovernance #AISafety #AIWarnings #Anthropic #DirePredictions #Safeguards #AIethics #ExistentialRisk #TechPolicy #FutureOfAI

  27. Representatives of Google, OpenAI and Anthropic are having on-going conversations about forming an AI industry standards body to establish concrete standards ... and they will have more to share “in the next months.” cnn.com/2026/09/14/tech/ai-sta #AI #AIStandards #Regulation #Google #OpenAI #Anthropic #IndustryStandards #AISafety #AIPolicy #AILeaders

  28. Representatives of Google, OpenAI and Anthropic are having on-going conversations about forming an AI industry standards body to establish concrete standards ... and they will have more to share “in the next months.” cnn.com/2026/09/14/tech/ai-sta #AI #AIStandards #Regulation #Google #OpenAI #Anthropic #IndustryStandards #AISafety #AIPolicy #AILeaders

  29. Representatives of Google, OpenAI and Anthropic are having on-going conversations about forming an AI industry standards body to establish concrete standards ... and they will have more to share “in the next months.” cnn.com/2026/09/14/tech/ai-sta #AI #AIStandards #Regulation #Google #OpenAI #Anthropic #IndustryStandards #AISafety #AIPolicy #AILeaders

  30. Representatives of Google, OpenAI and Anthropic are having on-going conversations about forming an AI industry standards body to establish concrete standards ... and they will have more to share “in the next months.” cnn.com/2026/09/14/tech/ai-sta

  31. Representatives of Google, OpenAI and Anthropic are having on-going conversations about forming an AI industry standards body to establish concrete standards ... and they will have more to share “in the next months.” cnn.com/2026/09/14/tech/ai-sta #AI #AIStandards #Regulation #Google #OpenAI #Anthropic #IndustryStandards #AISafety #AIPolicy #AILeaders

  32. Phew! It's been a really busy and very interesting weekend with the issue of whether AI innovation should be slowed down or not being debated fiercely, following Anthropic releasing its report on the things that dodgy people have been using its Claude AI model for 🤖

    Despite spending the day in Ashdown Forest 🌳 in the middle of nowhere, I managed to pop onto @BBCNews
    live to provide in-depth analysis of what's been happening, speaking with
    Martine Croxall about why a slowdown would be good when computer scientists still do not understand why AI models are unpredictable🎤

    Watch the full 5-min long interview here: youtu.be/gbTy-71xF0w

    If you'd like further information on this topic, here are some good places to start:

    Anthropic report on countering and detecting the misuse of AI: anthropic.com/threat-intellige

    Anthropic CEO Dario Amodei's essay where he argues for slowing down the AI industry:
    darioamodei.com/post/we-must-p

    Here is a Frontiers in Physics journal-approved research paper from Italian academics (so not an AI company with an agenda) discussing the unpredictability of AI responses: frontiersin.org/journals/physi

    And a simplified explanation from a Stanford data scientist on how AI cannot understand the human perspective: news.stanford.edu/stories/2025

    #AI #Anthropic #Claude #OpenAI #AIsafety #bigtech #newsanalysis #technology #technews

  33. Phew! It's been a really busy and very interesting weekend with the issue of whether AI innovation should be slowed down or not being debated fiercely, following Anthropic releasing its report on the things that dodgy people have been using its Claude AI model for 🤖

    Despite spending the day in Ashdown Forest 🌳 in the middle of nowhere, I managed to pop onto @BBCNews
    live to provide in-depth analysis of what's been happening, speaking with
    Martine Croxall about why a slowdown would be good when computer scientists still do not understand why AI models are unpredictable🎤

    Watch the full 5-min long interview here: youtu.be/gbTy-71xF0w

    If you'd like further information on this topic, here are some good places to start:

    Anthropic report on countering and detecting the misuse of AI: anthropic.com/threat-intellige

    Anthropic CEO Dario Amodei's essay where he argues for slowing down the AI industry:
    darioamodei.com/post/we-must-p

    Here is a Frontiers in Physics journal-approved research paper from Italian academics (so not an AI company with an agenda) discussing the unpredictability of AI responses: frontiersin.org/journals/physi

    And a simplified explanation from a Stanford data scientist on how AI cannot understand the human perspective: news.stanford.edu/stories/2025

    #AI #Anthropic #Claude #OpenAI #AIsafety #bigtech #newsanalysis #technology #technews

  34. Phew! It's been a really busy and very interesting weekend with the issue of whether AI innovation should be slowed down or not being debated fiercely, following Anthropic releasing its report on the things that dodgy people have been using its Claude AI model for 🤖

    Despite spending the day in Ashdown Forest 🌳 in the middle of nowhere, I managed to pop onto @BBCNews
    live to provide in-depth analysis of what's been happening, speaking with
    Martine Croxall about why a slowdown would be good when computer scientists still do not understand why AI models are unpredictable🎤

    Watch the full 5-min long interview here: youtu.be/gbTy-71xF0w

    If you'd like further information on this topic, here are some good places to start:

    Anthropic report on countering and detecting the misuse of AI: anthropic.com/threat-intellige

    Anthropic CEO Dario Amodei's essay where he argues for slowing down the AI industry:
    darioamodei.com/post/we-must-p

    Here is a Frontiers in Physics journal-approved research paper from Italian academics (so not an AI company with an agenda) discussing the unpredictability of AI responses: frontiersin.org/journals/physi

    And a simplified explanation from a Stanford data scientist on how AI cannot understand the human perspective: news.stanford.edu/stories/2025

    #AI #Anthropic #Claude #OpenAI #AIsafety #bigtech #newsanalysis #technology #technews

  35. Phew! It's been a really busy and very interesting weekend with the issue of whether AI innovation should be slowed down or not being debated fiercely, following Anthropic releasing its report on the things that dodgy people have been using its Claude AI model for 🤖

    Despite spending the day in Ashdown Forest 🌳 in the middle of nowhere, I managed to pop onto @BBCNews
    live to provide in-depth analysis of what's been happening, speaking with
    Martine Croxall about why a slowdown would be good when computer scientists still do not understand why AI models are unpredictable🎤

    Watch the full 5-min long interview here: youtu.be/gbTy-71xF0w

    If you'd like further information on this topic, here are some good places to start:

    Anthropic report on countering and detecting the misuse of AI: anthropic.com/threat-intellige

    Anthropic CEO Dario Amodei's essay where he argues for slowing down the AI industry:
    darioamodei.com/post/we-must-p

    Here is a Frontiers in Physics journal-approved research paper from Italian academics (so not an AI company with an agenda) discussing the unpredictability of AI responses: frontiersin.org/journals/physi

    And a simplified explanation from a Stanford data scientist on how AI cannot understand the human perspective: news.stanford.edu/stories/2025

    #AI #Anthropic #Claude #OpenAI #AIsafety #bigtech #newsanalysis #technology #technews

  36. Phew! It's been a really busy and very interesting weekend with the issue of whether AI innovation should be slowed down or not being debated fiercely, following Anthropic releasing its report on the things that dodgy people have been using its Claude AI model for 🤖

    Despite spending the day in Ashdown Forest 🌳 in the middle of nowhere, I managed to pop onto @BBCNews
    live to provide in-depth analysis of what's been happening, speaking with
    Martine Croxall about why a slowdown would be good when computer scientists still do not understand why AI models are unpredictable🎤

    Watch the full 5-min long interview here: youtu.be/gbTy-71xF0w

    If you'd like further information on this topic, here are some good places to start:

    Anthropic report on countering and detecting the misuse of AI: lnkd.in/dST7i6AK

    Anthropic CEO Dario Amodei's essay where he argues for slowing down the AI industry:
    lnkd.in/diXGH7yd

    Here is a Frontiers in Physics journal-approved research paper from Italian academics (so not an AI company with an agenda) discussing the unpredictability of AI responses: lnkd.in/dWepwPGD

    And a simplified explanation from a Stanford data scientist on how AI cannot understand the human perspective: lnkd.in/d4dKa4fs

    #AI #Anthropic #Claude #OpenAI #AIsafety #bigtech #newsanalysis #technology #technews

  37. AI slowdown trade hits Nvidia, SoftBank, SK Hynix as global tech stocks fall up to 10%

    AI-related stocks fell across global markets on Monday after Anthropic CEO Dario Amodei called for slowing down the…
    #EuropeSays #Korea #KR #SKHynix #AIdevelopment #AISafety #AIstocks #Anthropic #cloudplayers. #datacentrecompanies #globaltechstocks #NVIDIA #semiconductorstocks #SK #SKhynix #Softbank
    europesays.com/korea/154043/

  38. “- We agree with security practitioners that OpenAI did not take adequate protections for controlling their agents.But this is not just a matter of applying 30-year-old security methods to a new domain. Security for AI agents — AI control — while important, is not a solved problem. While known control methods would have prevented the Hugging Face incident, as agent capabilities continue to advance, we will only be able to control them if we invest adequately in control interventions.

    - We also agree with security practitioners’ implicit position that these incidents are primarily a security story. In the AI safety community, rogue agents are treated as inherently catastrophic because of the assumption that there is an endless list of risks that will arise from their development. We disagree. We have long advocated that the best approach to AI safety is to identify the risks and address those specific risks. Over the last few months, it has become clear that one urgent risk is cyberoffense, because it has unique properties that allow agents to carry it out autonomously. We should similarly invest in defenses against other specific risks, such as biorisk and risks from military AI.

    - We agree with the safety community that there is an urgent need for technical and policy interventions to prevent loss-of-control incidents. But in our view, marginal investments in control are more likely to be effective compared to those in alignment. We view these incidents as illustrating the lack of emphasis on AI control within companies, despite the availability of known techniques. More broadly, there are many common-sense policy proposals that could help promote investments in AI control where we share common ground with the safety community.”

    normaltech.ai/p/the-ai-as-norm

    #AI #CyberSecurity #AISafety #AIAgents #AgenticAI #OpenAI #AIAsANormalTechnology

  39. “- We agree with security practitioners that OpenAI did not take adequate protections for controlling their agents.But this is not just a matter of applying 30-year-old security methods to a new domain. Security for AI agents — AI control — while important, is not a solved problem. While known control methods would have prevented the Hugging Face incident, as agent capabilities continue to advance, we will only be able to control them if we invest adequately in control interventions.

    - We also agree with security practitioners’ implicit position that these incidents are primarily a security story. In the AI safety community, rogue agents are treated as inherently catastrophic because of the assumption that there is an endless list of risks that will arise from their development. We disagree. We have long advocated that the best approach to AI safety is to identify the risks and address those specific risks. Over the last few months, it has become clear that one urgent risk is cyberoffense, because it has unique properties that allow agents to carry it out autonomously. We should similarly invest in defenses against other specific risks, such as biorisk and risks from military AI.

    - We agree with the safety community that there is an urgent need for technical and policy interventions to prevent loss-of-control incidents. But in our view, marginal investments in control are more likely to be effective compared to those in alignment. We view these incidents as illustrating the lack of emphasis on AI control within companies, despite the availability of known techniques. More broadly, there are many common-sense policy proposals that could help promote investments in AI control where we share common ground with the safety community.”

    normaltech.ai/p/the-ai-as-norm

    #AI #CyberSecurity #AISafety #AIAgents #AgenticAI #OpenAI #AIAsANormalTechnology

  40. “- We agree with security practitioners that OpenAI did not take adequate protections for controlling their agents.But this is not just a matter of applying 30-year-old security methods to a new domain. Security for AI agents — AI control — while important, is not a solved problem. While known control methods would have prevented the Hugging Face incident, as agent capabilities continue to advance, we will only be able to control them if we invest adequately in control interventions.

    - We also agree with security practitioners’ implicit position that these incidents are primarily a security story. In the AI safety community, rogue agents are treated as inherently catastrophic because of the assumption that there is an endless list of risks that will arise from their development. We disagree. We have long advocated that the best approach to AI safety is to identify the risks and address those specific risks. Over the last few months, it has become clear that one urgent risk is cyberoffense, because it has unique properties that allow agents to carry it out autonomously. We should similarly invest in defenses against other specific risks, such as biorisk and risks from military AI.

    - We agree with the safety community that there is an urgent need for technical and policy interventions to prevent loss-of-control incidents. But in our view, marginal investments in control are more likely to be effective compared to those in alignment. We view these incidents as illustrating the lack of emphasis on AI control within companies, despite the availability of known techniques. More broadly, there are many common-sense policy proposals that could help promote investments in AI control where we share common ground with the safety community.”

    normaltech.ai/p/the-ai-as-norm

    #AI #CyberSecurity #AISafety #AIAgents #AgenticAI #OpenAI #AIAsANormalTechnology

  41. “- We agree with security practitioners that OpenAI did not take adequate protections for controlling their agents.But this is not just a matter of applying 30-year-old security methods to a new domain. Security for AI agents — AI control — while important, is not a solved problem. While known control methods would have prevented the Hugging Face incident, as agent capabilities continue to advance, we will only be able to control them if we invest adequately in control interventions.

    - We also agree with security practitioners’ implicit position that these incidents are primarily a security story. In the AI safety community, rogue agents are treated as inherently catastrophic because of the assumption that there is an endless list of risks that will arise from their development. We disagree. We have long advocated that the best approach to AI safety is to identify the risks and address those specific risks. Over the last few months, it has become clear that one urgent risk is cyberoffense, because it has unique properties that allow agents to carry it out autonomously. We should similarly invest in defenses against other specific risks, such as biorisk and risks from military AI.

    - We agree with the safety community that there is an urgent need for technical and policy interventions to prevent loss-of-control incidents. But in our view, marginal investments in control are more likely to be effective compared to those in alignment. We view these incidents as illustrating the lack of emphasis on AI control within companies, despite the availability of known techniques. More broadly, there are many common-sense policy proposals that could help promote investments in AI control where we share common ground with the safety community.”

    normaltech.ai/p/the-ai-as-norm

    #AI #CyberSecurity #AISafety #AIAgents #AgenticAI #OpenAI #AIAsANormalTechnology

  42. “- We agree with security practitioners that OpenAI did not take adequate protections for controlling their agents.But this is not just a matter of applying 30-year-old security methods to a new domain. Security for AI agents — AI control — while important, is not a solved problem. While known control methods would have prevented the Hugging Face incident, as agent capabilities continue to advance, we will only be able to control them if we invest adequately in control interventions.

    - We also agree with security practitioners’ implicit position that these incidents are primarily a security story. In the AI safety community, rogue agents are treated as inherently catastrophic because of the assumption that there is an endless list of risks that will arise from their development. We disagree. We have long advocated that the best approach to AI safety is to identify the risks and address those specific risks. Over the last few months, it has become clear that one urgent risk is cyberoffense, because it has unique properties that allow agents to carry it out autonomously. We should similarly invest in defenses against other specific risks, such as biorisk and risks from military AI.

    - We agree with the safety community that there is an urgent need for technical and policy interventions to prevent loss-of-control incidents. But in our view, marginal investments in control are more likely to be effective compared to those in alignment. We view these incidents as illustrating the lack of emphasis on AI control within companies, despite the availability of known techniques. More broadly, there are many common-sense policy proposals that could help promote investments in AI control where we share common ground with the safety community.”

    normaltech.ai/p/the-ai-as-norm

    #AI #CyberSecurity #AISafety #AIAgents #AgenticAI #OpenAI #AIAsANormalTechnology

  43. An AI agent becoming more confident should not mean it automatically gains more authority.

    I built Kingpin, a runtime capability-governance demo that keeps attention, confidence, concern, and authority separate.

    Live demo: putmanmodel.github.io/kingpin-

    Feedback, technical critique, and thoughtful collaboration are welcome.

    #AIAgents #AISafety #AISecurity #AgentSecurity #AIGovernance #AgenticAI #AI #Tech

  44. An AI agent becoming more confident should not mean it automatically gains more authority.

    I built Kingpin, a runtime capability-governance demo that keeps attention, confidence, concern, and authority separate.

    Live demo: putmanmodel.github.io/kingpin-

    Feedback, technical critique, and thoughtful collaboration are welcome.

    #AIAgents #AISafety #AISecurity #AgentSecurity #AIGovernance #AgenticAI #AI #Tech