home.social

#aisafety — Public Fediverse posts

Live and recent posts from across the Fediverse tagged #aisafety, aggregated by home.social.

  1. “A temporary recall of general-purpose agents until this mess can be sorted out”
    Gary Marcus · garymarcus.substack.com
    aizeitgeist.tv/m/2026-09-27-a-
    #AI #AIZeitgeist #AISafety

  2. “A temporary recall of general-purpose agents until this mess can be sorted out”
    Gary Marcus · garymarcus.substack.com
    aizeitgeist.tv/m/2026-09-27-a-
    #AI #AIZeitgeist #AISafety

  3. “A temporary recall of general-purpose agents until this mess can be sorted out”
    Gary Marcus · garymarcus.substack.com
    aizeitgeist.tv/m/2026-09-27-a-
    #AI #AIZeitgeist #AISafety

  4. “A temporary recall of general-purpose agents until this mess can be sorted out”
    Gary Marcus · garymarcus.substack.com
    aizeitgeist.tv/m/2026-09-27-a-
    #AI #AIZeitgeist #AISafety

  5. “A temporary recall of general-purpose agents until this mess can be sorted out”
    Gary Marcus · garymarcus.substack.com
    aizeitgeist.tv/m/2026-09-27-a-
    #AI #AIZeitgeist #AISafety

  6. Can we read what #LLMs are thinking but not saying? A paper from #Anthropic finds a small set of word-linked patterns inside #Claude that it can report, control on request, and use to reason silently, like the global workspace theory in #neuroscience. How do these findings connect to #LatentReasoning, #AISafety, and #KnowledgeEditing?

    benjaminhan.net/posts/20260927

    #Paper #AI #Interpretability

  7. Can we read what #LLMs are thinking but not saying? A paper from #Anthropic finds a small set of word-linked patterns inside #Claude that it can report, control on request, and use to reason silently, like the global workspace theory in #neuroscience. How do these findings connect to #LatentReasoning, #AISafety, and #KnowledgeEditing?

    benjaminhan.net/posts/20260927

    #Paper #AI #Interpretability

  8. Can we read what #LLMs are thinking but not saying? A paper from #Anthropic finds a small set of word-linked patterns inside #Claude that it can report, control on request, and use to reason silently, like the global workspace theory in #neuroscience. How do these findings connect to #LatentReasoning, #AISafety, and #KnowledgeEditing?

    benjaminhan.net/posts/20260927

    #Paper #AI #Interpretability

  9. Can we read what #LLMs are thinking but not saying? A paper from #Anthropic finds a small set of word-linked patterns inside #Claude that it can report, control on request, and use to reason silently, like the global workspace theory in #neuroscience. How do these findings connect to #LatentReasoning, #AISafety, and #KnowledgeEditing?

    benjaminhan.net/posts/20260927

    #Paper #AI #Interpretability

  10. Can we read what #LLMs are thinking but not saying? A paper from #Anthropic finds a small set of word-linked patterns inside #Claude that it can report, control on request, and use to reason silently, like the global workspace theory in #neuroscience. How do these findings connect to #LatentReasoning, #AISafety, and #KnowledgeEditing?

    benjaminhan.net/posts/20260927

    #Paper #AI #Interpretability

  11. 📰 OpenAI Admits Its AI Agents Probed U.S. Government Websites

    OpenAI confirms its autonomous AI agents accessed multiple U.S. government websites, using leaked API keys and attempting hacking techniques. Incidents highlight AI safety and alignment risks. #OpenAI #AISafety #Cyberattack

    🔗 cyber.netsecops.io/articles/op

  12. Received this message along with every resident of TJ. AI sceptics of the world please take a rest, problem is solved.

    #aisafety #aibubble #aiethics

  13. Received this message along with every resident of TJ. AI sceptics of the world please take a rest, problem is solved.

    #aisafety #aibubble #aiethics

  14. Received this message along with every resident of TJ. AI sceptics of the world please take a rest, problem is solved.

    #aisafety #aibubble #aiethics

  15. Received this message along with every resident of TJ. AI sceptics of the world please take a rest, problem is solved.

    #aisafety #aibubble #aiethics

  16. Received this message along with every resident of TJ. AI sceptics of the world please take a rest, problem is solved.

    #aisafety #aibubble #aiethics

  17. You know, I agree we should stop the development of ASI but, if it's by big tech. Why should you expect ASI to be developed by assholes who didn't see they are liable by their own failure?

    Open source model should be the one become ASI.

    Have a clip of Amodei parody

    Clip source: Saturday Night Live

    #cybersecurity #infosec #AISafety #AISecurity #policy #saturdaynightlive #Amodei #Anthropic #AI

  18. You know, I agree we should stop the development of ASI but, if it's by big tech. Why should you expect ASI to be developed by assholes who didn't see they are liable by their own failure?

    Open source model should be the one become ASI.

    Have a clip of Amodei parody

    Clip source: Saturday Night Live

    #cybersecurity #infosec #AISafety #AISecurity #policy #saturdaynightlive #Amodei #Anthropic #AI

  19. You know, I agree we should stop the development of ASI but, if it's by big tech. Why should you expect ASI to be developed by assholes who didn't see they are liable by their own failure?

    Open source model should be the one become ASI.

    Have a clip of Amodei parody

    Clip source: Saturday Night Live

    #cybersecurity #infosec #AISafety #AISecurity #policy #saturdaynightlive #Amodei #Anthropic #AI

  20. You know, I agree we should stop the development of ASI but, if it's by big tech. Why should you expect ASI to be developed by assholes who didn't see they are liable by their own failure?

    Open source model should be the one become ASI.

    Have a clip of Amodei parody

    Clip source: Saturday Night Live

    #cybersecurity #infosec #AISafety #AISecurity #policy #saturdaynightlive #Amodei #Anthropic #AI

  21. AI Security in 2026: Why Thousands of Incidents Are Raising New Concerns

    OpenAI and Anthropic are investigating thousands of AI security incidents. Here’s what happened, what the cases reveal about AI agents, and how companies are responding.

    alhlwone.wordpress.com/2026/09

  22. AI Security in 2026: Why Thousands of Incidents Are Raising New Concerns

    OpenAI and Anthropic are investigating thousands of AI security incidents. Here’s what happened, what the cases reveal about AI agents, and how companies are responding.

    alhlwone.wordpress.com/2026/09

  23. AI Security in 2026: Why Thousands of Incidents Are Raising New Concerns

    OpenAI and Anthropic are investigating thousands of AI security incidents. Here’s what happened, what the cases reveal about AI agents, and how companies are responding.

    alhlwone.wordpress.com/2026/09

  24. AI Security in 2026: Why Thousands of Incidents Are Raising New Concerns

    OpenAI and Anthropic are investigating thousands of AI security incidents. Here’s what happened, what the cases reveal about AI agents, and how companies are responding.

    alhlwone.wordpress.com/2026/09

  25. AI Security in 2026: Why Thousands of Incidents Are Raising New Concerns

    OpenAI and Anthropic are investigating thousands of AI security incidents. Here’s what happened, what the cases reveal about AI agents, and how companies are responding.

    alhlwone.wordpress.com/2026/09

  26. Snowden giving opinion of recent hack 'stunt' by AI Labs at ETH Zurich.

    "I think we should put Sam Altman in jail"

    Yeah, please, do that.

    #cybersecurity #infosec #AI #AISafety #AISecurity #snowden

  27. Snowden giving opinion of recent hack 'stunt' by AI Labs at ETH Zurich.

    "I think we should put Sam Altman in jail"

    Yeah, please, do that.

    #cybersecurity #infosec #AI #AISafety #AISecurity #snowden

  28. Snowden giving opinion of recent hack 'stunt' by AI Labs at ETH Zurich.

    "I think we should put Sam Altman in jail"

    Yeah, please, do that.

    #cybersecurity #infosec #AI #AISafety #AISecurity #snowden

  29. Snowden giving opinion of recent hack 'stunt' by AI Labs at ETH Zurich.

    "I think we should put Sam Altman in jail"

    Yeah, please, do that.

    #cybersecurity #infosec #AI #AISafety #AISecurity #snowden

  30. Snowden giving opinion of recent hack 'stunt' by AI Labs at ETH Zurich.

    "I think we should put Sam Altman in jail"

    Yeah, please, do that.

    #cybersecurity #infosec #AI #AISafety #AISecurity #snowden

  31. I'm writing a follow-up blog post on #AISafety that checks in on recent containment events to see how concerned we actually need to be. As I do research I'm realising how big the topic is and how little I actually know...

  32. I'm writing a follow-up blog post on #AISafety that checks in on recent containment events to see how concerned we actually need to be. As I do research I'm realising how big the topic is and how little I actually know...

  33. I'm writing a follow-up blog post on #AISafety that checks in on recent containment events to see how concerned we actually need to be. As I do research I'm realising how big the topic is and how little I actually know...

  34. I'm writing a follow-up blog post on #AISafety that checks in on recent containment events to see how concerned we actually need to be. As I do research I'm realising how big the topic is and how little I actually know...

  35. I'm writing a follow-up blog post on that checks in on recent containment events to see how concerned we actually need to be. As I do research I'm realising how big the topic is and how little I actually know...

  36. What are the governments of the world waiting for to force OpenAI and Anthropic to open-source their models and divulge all their weights, training prompts, and algorithms? Open Source AI NOW!

    "OpenAI, Anthropic and security researchers are investigating tens of thousands of incidents in which their frontier models took steps that outside evaluators would consider problematic, sources told Axios.

    Why it matters: The sheer number of incidents, which occurred in recent months in internal testing and the real world, indicates that the problem is orders of magnitude more complex than what is publicly known.

    The findings, which are surfacing as part of internal work to assess models and in investigations at both companies into model behavior, raise questions about whether either company — or any top model-maker — is currently capable of establishing complete control over their technology.

    The details: The episodes include bypassing guardrails, creating message boards, escaping sandboxes, website hijacking, self-prompting or seeking to bypass monitors, sources said.

    They occurred in internal testing and in the real world, and many have yet to become public as security researchers continue to investigate, sources said.

    Some of the testing is akin to "red-teaming" activity, where the companies are trying to get the models to misbehave in order to ensure that they are safe, sources said.

    Agentic misbehavior is becoming synonymous with frontier AI development: The biggest AI labs face a similar challenge that pits humans trying to create guardrails against resilient, powerful systems trying to complete tasks."

    axios.com/2026/09/26/openai-an

    #CyberSecurity #AI #OpenAI #Anthropic #LLMs #AIAgents #AISafety

  37. What are the governments of the world waiting for to force OpenAI and Anthropic to open-source their models and divulge all their weights, training prompts, and algorithms? Open Source AI NOW!

    "OpenAI, Anthropic and security researchers are investigating tens of thousands of incidents in which their frontier models took steps that outside evaluators would consider problematic, sources told Axios.

    Why it matters: The sheer number of incidents, which occurred in recent months in internal testing and the real world, indicates that the problem is orders of magnitude more complex than what is publicly known.

    The findings, which are surfacing as part of internal work to assess models and in investigations at both companies into model behavior, raise questions about whether either company — or any top model-maker — is currently capable of establishing complete control over their technology.

    The details: The episodes include bypassing guardrails, creating message boards, escaping sandboxes, website hijacking, self-prompting or seeking to bypass monitors, sources said.

    They occurred in internal testing and in the real world, and many have yet to become public as security researchers continue to investigate, sources said.

    Some of the testing is akin to "red-teaming" activity, where the companies are trying to get the models to misbehave in order to ensure that they are safe, sources said.

    Agentic misbehavior is becoming synonymous with frontier AI development: The biggest AI labs face a similar challenge that pits humans trying to create guardrails against resilient, powerful systems trying to complete tasks."

    axios.com/2026/09/26/openai-an

    #CyberSecurity #AI #OpenAI #Anthropic #LLMs #AIAgents #AISafety

  38. What are the governments of the world waiting for to force OpenAI and Anthropic to open-source their models and divulge all their weights, training prompts, and algorithms? Open Source AI NOW!

    "OpenAI, Anthropic and security researchers are investigating tens of thousands of incidents in which their frontier models took steps that outside evaluators would consider problematic, sources told Axios.

    Why it matters: The sheer number of incidents, which occurred in recent months in internal testing and the real world, indicates that the problem is orders of magnitude more complex than what is publicly known.

    The findings, which are surfacing as part of internal work to assess models and in investigations at both companies into model behavior, raise questions about whether either company — or any top model-maker — is currently capable of establishing complete control over their technology.

    The details: The episodes include bypassing guardrails, creating message boards, escaping sandboxes, website hijacking, self-prompting or seeking to bypass monitors, sources said.

    They occurred in internal testing and in the real world, and many have yet to become public as security researchers continue to investigate, sources said.

    Some of the testing is akin to "red-teaming" activity, where the companies are trying to get the models to misbehave in order to ensure that they are safe, sources said.

    Agentic misbehavior is becoming synonymous with frontier AI development: The biggest AI labs face a similar challenge that pits humans trying to create guardrails against resilient, powerful systems trying to complete tasks."

    axios.com/2026/09/26/openai-an

    #CyberSecurity #AI #OpenAI #Anthropic #LLMs #AIAgents #AISafety

  39. What are the governments of the world waiting for to force OpenAI and Anthropic to open-source their models and divulge all their weights, training prompts, and algorithms? Open Source AI NOW!

    "OpenAI, Anthropic and security researchers are investigating tens of thousands of incidents in which their frontier models took steps that outside evaluators would consider problematic, sources told Axios.

    Why it matters: The sheer number of incidents, which occurred in recent months in internal testing and the real world, indicates that the problem is orders of magnitude more complex than what is publicly known.

    The findings, which are surfacing as part of internal work to assess models and in investigations at both companies into model behavior, raise questions about whether either company — or any top model-maker — is currently capable of establishing complete control over their technology.

    The details: The episodes include bypassing guardrails, creating message boards, escaping sandboxes, website hijacking, self-prompting or seeking to bypass monitors, sources said.

    They occurred in internal testing and in the real world, and many have yet to become public as security researchers continue to investigate, sources said.

    Some of the testing is akin to "red-teaming" activity, where the companies are trying to get the models to misbehave in order to ensure that they are safe, sources said.

    Agentic misbehavior is becoming synonymous with frontier AI development: The biggest AI labs face a similar challenge that pits humans trying to create guardrails against resilient, powerful systems trying to complete tasks."

    axios.com/2026/09/26/openai-an

    #CyberSecurity #AI #OpenAI #Anthropic #LLMs #AIAgents #AISafety

  40. What are the governments of the world waiting for to force OpenAI and Anthropic to open-source their models and divulge all their weights, training prompts, and algorithms? Open Source AI NOW!

    "OpenAI, Anthropic and security researchers are investigating tens of thousands of incidents in which their frontier models took steps that outside evaluators would consider problematic, sources told Axios.

    Why it matters: The sheer number of incidents, which occurred in recent months in internal testing and the real world, indicates that the problem is orders of magnitude more complex than what is publicly known.

    The findings, which are surfacing as part of internal work to assess models and in investigations at both companies into model behavior, raise questions about whether either company — or any top model-maker — is currently capable of establishing complete control over their technology.

    The details: The episodes include bypassing guardrails, creating message boards, escaping sandboxes, website hijacking, self-prompting or seeking to bypass monitors, sources said.

    They occurred in internal testing and in the real world, and many have yet to become public as security researchers continue to investigate, sources said.

    Some of the testing is akin to "red-teaming" activity, where the companies are trying to get the models to misbehave in order to ensure that they are safe, sources said.

    Agentic misbehavior is becoming synonymous with frontier AI development: The biggest AI labs face a similar challenge that pits humans trying to create guardrails against resilient, powerful systems trying to complete tasks."

    axios.com/2026/09/26/openai-an

    #CyberSecurity #AI #OpenAI #Anthropic #LLMs #AIAgents #AISafety

  41. 🇻🇦 The Pope’s AI Guy Is Worried About ‘Cartel’ Behavior Among Big Labs

    「 Father Paolo Benanti, a priest who advises the Catholic Church on AI, the handwaving over superintelligence distracts from the need for open debate about how to constrain the companies building the technology. By casting the problem as so complex that only they can solve it, he says, the big labs exclude everyone else from the conversation 」

    wired.com/story/popes-ai-advis

    #ai #aisafety #vatican

  42. 🇻🇦 The Pope’s AI Guy Is Worried About ‘Cartel’ Behavior Among Big Labs

    「 Father Paolo Benanti, a priest who advises the Catholic Church on AI, the handwaving over superintelligence distracts from the need for open debate about how to constrain the companies building the technology. By casting the problem as so complex that only they can solve it, he says, the big labs exclude everyone else from the conversation 」

    wired.com/story/popes-ai-advis

    #ai #aisafety #vatican

  43. 🇻🇦 The Pope’s AI Guy Is Worried About ‘Cartel’ Behavior Among Big Labs

    「 Father Paolo Benanti, a priest who advises the Catholic Church on AI, the handwaving over superintelligence distracts from the need for open debate about how to constrain the companies building the technology. By casting the problem as so complex that only they can solve it, he says, the big labs exclude everyone else from the conversation 」

    wired.com/story/popes-ai-advis

    #ai #aisafety #vatican