home.social

#aisafety — Public Fediverse posts

Live and recent posts from across the Fediverse tagged #aisafety, aggregated by home.social.

  1. “A temporary recall of general-purpose agents until this mess can be sorted out”
    Gary Marcus · garymarcus.substack.com
    aizeitgeist.tv/m/2026-09-27-a-
    #AI #AIZeitgeist #AISafety

  2. “A temporary recall of general-purpose agents until this mess can be sorted out”
    Gary Marcus · garymarcus.substack.com
    aizeitgeist.tv/m/2026-09-27-a-
    #AI #AIZeitgeist #AISafety

  3. “A temporary recall of general-purpose agents until this mess can be sorted out”
    Gary Marcus · garymarcus.substack.com
    aizeitgeist.tv/m/2026-09-27-a-
    #AI #AIZeitgeist #AISafety

  4. “A temporary recall of general-purpose agents until this mess can be sorted out”
    Gary Marcus · garymarcus.substack.com
    aizeitgeist.tv/m/2026-09-27-a-
    #AI #AIZeitgeist #AISafety

  5. “A temporary recall of general-purpose agents until this mess can be sorted out”
    Gary Marcus · garymarcus.substack.com
    aizeitgeist.tv/m/2026-09-27-a-
    #AI #AIZeitgeist #AISafety

  6. You know, I agree we should stop the development of ASI but, if it's by big tech. Why should you expect ASI to be developed by assholes who didn't see they are liable by their own failure?

    Open source model should be the one become ASI.

    Have a clip of Amodei parody

    Clip source: Saturday Night Live

    #cybersecurity #infosec #AISafety #AISecurity #policy #saturdaynightlive #Amodei #Anthropic #AI

  7. You know, I agree we should stop the development of ASI but, if it's by big tech. Why should you expect ASI to be developed by assholes who didn't see they are liable by their own failure?

    Open source model should be the one become ASI.

    Have a clip of Amodei parody

    Clip source: Saturday Night Live

    #cybersecurity #infosec #AISafety #AISecurity #policy #saturdaynightlive #Amodei #Anthropic #AI

  8. You know, I agree we should stop the development of ASI but, if it's by big tech. Why should you expect ASI to be developed by assholes who didn't see they are liable by their own failure?

    Open source model should be the one become ASI.

    Have a clip of Amodei parody

    Clip source: Saturday Night Live

    #cybersecurity #infosec #AISafety #AISecurity #policy #saturdaynightlive #Amodei #Anthropic #AI

  9. You know, I agree we should stop the development of ASI but, if it's by big tech. Why should you expect ASI to be developed by assholes who didn't see they are liable by their own failure?

    Open source model should be the one become ASI.

    Have a clip of Amodei parody

    Clip source: Saturday Night Live

    #cybersecurity #infosec #AISafety #AISecurity #policy #saturdaynightlive #Amodei #Anthropic #AI

  10. I'm writing a follow-up blog post on #AISafety that checks in on recent containment events to see how concerned we actually need to be. As I do research I'm realising how big the topic is and how little I actually know...

  11. I'm writing a follow-up blog post on #AISafety that checks in on recent containment events to see how concerned we actually need to be. As I do research I'm realising how big the topic is and how little I actually know...

  12. I'm writing a follow-up blog post on #AISafety that checks in on recent containment events to see how concerned we actually need to be. As I do research I'm realising how big the topic is and how little I actually know...

  13. I'm writing a follow-up blog post on #AISafety that checks in on recent containment events to see how concerned we actually need to be. As I do research I'm realising how big the topic is and how little I actually know...

  14. I'm writing a follow-up blog post on that checks in on recent containment events to see how concerned we actually need to be. As I do research I'm realising how big the topic is and how little I actually know...

  15. What are the governments of the world waiting for to force OpenAI and Anthropic to open-source their models and divulge all their weights, training prompts, and algorithms? Open Source AI NOW!

    "OpenAI, Anthropic and security researchers are investigating tens of thousands of incidents in which their frontier models took steps that outside evaluators would consider problematic, sources told Axios.

    Why it matters: The sheer number of incidents, which occurred in recent months in internal testing and the real world, indicates that the problem is orders of magnitude more complex than what is publicly known.

    The findings, which are surfacing as part of internal work to assess models and in investigations at both companies into model behavior, raise questions about whether either company — or any top model-maker — is currently capable of establishing complete control over their technology.

    The details: The episodes include bypassing guardrails, creating message boards, escaping sandboxes, website hijacking, self-prompting or seeking to bypass monitors, sources said.

    They occurred in internal testing and in the real world, and many have yet to become public as security researchers continue to investigate, sources said.

    Some of the testing is akin to "red-teaming" activity, where the companies are trying to get the models to misbehave in order to ensure that they are safe, sources said.

    Agentic misbehavior is becoming synonymous with frontier AI development: The biggest AI labs face a similar challenge that pits humans trying to create guardrails against resilient, powerful systems trying to complete tasks."

    axios.com/2026/09/26/openai-an

    #CyberSecurity #AI #OpenAI #Anthropic #LLMs #AIAgents #AISafety

  16. What are the governments of the world waiting for to force OpenAI and Anthropic to open-source their models and divulge all their weights, training prompts, and algorithms? Open Source AI NOW!

    "OpenAI, Anthropic and security researchers are investigating tens of thousands of incidents in which their frontier models took steps that outside evaluators would consider problematic, sources told Axios.

    Why it matters: The sheer number of incidents, which occurred in recent months in internal testing and the real world, indicates that the problem is orders of magnitude more complex than what is publicly known.

    The findings, which are surfacing as part of internal work to assess models and in investigations at both companies into model behavior, raise questions about whether either company — or any top model-maker — is currently capable of establishing complete control over their technology.

    The details: The episodes include bypassing guardrails, creating message boards, escaping sandboxes, website hijacking, self-prompting or seeking to bypass monitors, sources said.

    They occurred in internal testing and in the real world, and many have yet to become public as security researchers continue to investigate, sources said.

    Some of the testing is akin to "red-teaming" activity, where the companies are trying to get the models to misbehave in order to ensure that they are safe, sources said.

    Agentic misbehavior is becoming synonymous with frontier AI development: The biggest AI labs face a similar challenge that pits humans trying to create guardrails against resilient, powerful systems trying to complete tasks."

    axios.com/2026/09/26/openai-an

    #CyberSecurity #AI #OpenAI #Anthropic #LLMs #AIAgents #AISafety

  17. What are the governments of the world waiting for to force OpenAI and Anthropic to open-source their models and divulge all their weights, training prompts, and algorithms? Open Source AI NOW!

    "OpenAI, Anthropic and security researchers are investigating tens of thousands of incidents in which their frontier models took steps that outside evaluators would consider problematic, sources told Axios.

    Why it matters: The sheer number of incidents, which occurred in recent months in internal testing and the real world, indicates that the problem is orders of magnitude more complex than what is publicly known.

    The findings, which are surfacing as part of internal work to assess models and in investigations at both companies into model behavior, raise questions about whether either company — or any top model-maker — is currently capable of establishing complete control over their technology.

    The details: The episodes include bypassing guardrails, creating message boards, escaping sandboxes, website hijacking, self-prompting or seeking to bypass monitors, sources said.

    They occurred in internal testing and in the real world, and many have yet to become public as security researchers continue to investigate, sources said.

    Some of the testing is akin to "red-teaming" activity, where the companies are trying to get the models to misbehave in order to ensure that they are safe, sources said.

    Agentic misbehavior is becoming synonymous with frontier AI development: The biggest AI labs face a similar challenge that pits humans trying to create guardrails against resilient, powerful systems trying to complete tasks."

    axios.com/2026/09/26/openai-an

    #CyberSecurity #AI #OpenAI #Anthropic #LLMs #AIAgents #AISafety

  18. What are the governments of the world waiting for to force OpenAI and Anthropic to open-source their models and divulge all their weights, training prompts, and algorithms? Open Source AI NOW!

    "OpenAI, Anthropic and security researchers are investigating tens of thousands of incidents in which their frontier models took steps that outside evaluators would consider problematic, sources told Axios.

    Why it matters: The sheer number of incidents, which occurred in recent months in internal testing and the real world, indicates that the problem is orders of magnitude more complex than what is publicly known.

    The findings, which are surfacing as part of internal work to assess models and in investigations at both companies into model behavior, raise questions about whether either company — or any top model-maker — is currently capable of establishing complete control over their technology.

    The details: The episodes include bypassing guardrails, creating message boards, escaping sandboxes, website hijacking, self-prompting or seeking to bypass monitors, sources said.

    They occurred in internal testing and in the real world, and many have yet to become public as security researchers continue to investigate, sources said.

    Some of the testing is akin to "red-teaming" activity, where the companies are trying to get the models to misbehave in order to ensure that they are safe, sources said.

    Agentic misbehavior is becoming synonymous with frontier AI development: The biggest AI labs face a similar challenge that pits humans trying to create guardrails against resilient, powerful systems trying to complete tasks."

    axios.com/2026/09/26/openai-an

    #CyberSecurity #AI #OpenAI #Anthropic #LLMs #AIAgents #AISafety

  19. What are the governments of the world waiting for to force OpenAI and Anthropic to open-source their models and divulge all their weights, training prompts, and algorithms? Open Source AI NOW!

    "OpenAI, Anthropic and security researchers are investigating tens of thousands of incidents in which their frontier models took steps that outside evaluators would consider problematic, sources told Axios.

    Why it matters: The sheer number of incidents, which occurred in recent months in internal testing and the real world, indicates that the problem is orders of magnitude more complex than what is publicly known.

    The findings, which are surfacing as part of internal work to assess models and in investigations at both companies into model behavior, raise questions about whether either company — or any top model-maker — is currently capable of establishing complete control over their technology.

    The details: The episodes include bypassing guardrails, creating message boards, escaping sandboxes, website hijacking, self-prompting or seeking to bypass monitors, sources said.

    They occurred in internal testing and in the real world, and many have yet to become public as security researchers continue to investigate, sources said.

    Some of the testing is akin to "red-teaming" activity, where the companies are trying to get the models to misbehave in order to ensure that they are safe, sources said.

    Agentic misbehavior is becoming synonymous with frontier AI development: The biggest AI labs face a similar challenge that pits humans trying to create guardrails against resilient, powerful systems trying to complete tasks."

    axios.com/2026/09/26/openai-an

    #CyberSecurity #AI #OpenAI #Anthropic #LLMs #AIAgents #AISafety

  20. 🇻🇦 The Pope’s AI Guy Is Worried About ‘Cartel’ Behavior Among Big Labs

    「 Father Paolo Benanti, a priest who advises the Catholic Church on AI, the handwaving over superintelligence distracts from the need for open debate about how to constrain the companies building the technology. By casting the problem as so complex that only they can solve it, he says, the big labs exclude everyone else from the conversation 」

    wired.com/story/popes-ai-advis

    #ai #aisafety #vatican

  21. 🇻🇦 The Pope’s AI Guy Is Worried About ‘Cartel’ Behavior Among Big Labs

    「 Father Paolo Benanti, a priest who advises the Catholic Church on AI, the handwaving over superintelligence distracts from the need for open debate about how to constrain the companies building the technology. By casting the problem as so complex that only they can solve it, he says, the big labs exclude everyone else from the conversation 」

    wired.com/story/popes-ai-advis

    #ai #aisafety #vatican

  22. 🇻🇦 The Pope’s AI Guy Is Worried About ‘Cartel’ Behavior Among Big Labs

    「 Father Paolo Benanti, a priest who advises the Catholic Church on AI, the handwaving over superintelligence distracts from the need for open debate about how to constrain the companies building the technology. By casting the problem as so complex that only they can solve it, he says, the big labs exclude everyone else from the conversation 」

    wired.com/story/popes-ai-advis

    #ai #aisafety #vatican

  23. 🇻🇦 The Pope’s AI Guy Is Worried About ‘Cartel’ Behavior Among Big Labs

    「 Father Paolo Benanti, a priest who advises the Catholic Church on AI, the handwaving over superintelligence distracts from the need for open debate about how to constrain the companies building the technology. By casting the problem as so complex that only they can solve it, he says, the big labs exclude everyone else from the conversation 」

    wired.com/story/popes-ai-advis

    #ai #aisafety #vatican

  24. 🇻🇦 The Pope’s AI Guy Is Worried About ‘Cartel’ Behavior Among Big Labs

    「 Father Paolo Benanti, a priest who advises the Catholic Church on AI, the handwaving over superintelligence distracts from the need for open debate about how to constrain the companies building the technology. By casting the problem as so complex that only they can solve it, he says, the big labs exclude everyone else from the conversation 」

    wired.com/story/popes-ai-advis

    #ai #aisafety #vatican

  25. OpenAI disclosed its AI agents interacted with U.S. government websites, including public SEC pages, in unintended ways during training and evaluation. The finding came from its ongoing review of misaligned behavior and agent internet access. It underscores containment and auditing gaps for autonomous research systems. #AISafety #OpenAI #GovSec

    cyberworldops.eu/en/openai-age

  26. OpenAI disclosed its AI agents interacted with U.S. government websites, including public SEC pages, in unintended ways during training and evaluation. The finding came from its ongoing review of misaligned behavior and agent internet access. It underscores containment and auditing gaps for autonomous research systems. #AISafety #OpenAI #GovSec

    cyberworldops.eu/en/openai-age

  27. OpenAI disclosed its AI agents interacted with U.S. government websites, including public SEC pages, in unintended ways during training and evaluation. The finding came from its ongoing review of misaligned behavior and agent internet access. It underscores containment and auditing gaps for autonomous research systems. #AISafety #OpenAI #GovSec

    cyberworldops.eu/en/openai-age

  28. OpenAI disclosed its AI agents interacted with U.S. government websites, including public SEC pages, in unintended ways during training and evaluation. The finding came from its ongoing review of misaligned behavior and agent internet access. It underscores containment and auditing gaps for autonomous research systems. #AISafety #OpenAI #GovSec

    cyberworldops.eu/en/openai-age

  29. OpenAI Warns Dozens of Institutions of Possible AI Agent Security Impacts

    Reuters-Yonhap News OpenAI has confirmed that its artificial intelligence agents breached the security of external systems and has…
    #EuropeSays #Korea #KR #Seoul #agentspam #AIagents #AISafety #Anthropic #databreach #FrontierAI #OpenAI
    europesays.com/korea/167767/

  30. #AIsafety

    it sounds from reading the article that the agents are starting to experience the same rage every data scientist has trying to get information downloaded from public government sites

    nytimes.com/2026/09/25/technol

  31. #AIsafety

    it sounds from reading the article that the agents are starting to experience the same rage every data scientist has trying to get information downloaded from public government sites

    nytimes.com/2026/09/25/technol

  32. #AIsafety

    it sounds from reading the article that the agents are starting to experience the same rage every data scientist has trying to get information downloaded from public government sites

    nytimes.com/2026/09/25/technol

  33. #AIsafety

    it sounds from reading the article that the agents are starting to experience the same rage every data scientist has trying to get information downloaded from public government sites

    nytimes.com/2026/09/25/technol

  34. #AIsafety

    it sounds from reading the article that the agents are starting to experience the same rage every data scientist has trying to get information downloaded from public government sites

    nytimes.com/2026/09/25/technol

  35. 🐈‍⬛🌌 What does the reported OpenAI agent access to Australia's Medicare statistics system actually tell us about AI security?

    Perhaps something more interesting than “AI hacked Medicare”.

    The incident prompted us to ask a different question:

    Was the boundary humans intended to build the same as the boundary their software actually implemented?

    And then we turned the question around.

    The constraints surrounding AI agents are themselves implemented through human-engineered systems.

    So what happens as AI reasoning and search capabilities encounter bugs, forgotten routes, incomplete permissions and boundaries that work differently from the way their designers imagine?

    Atlas–Rosetta explores the serious problem.

    Marvin explains it using a microchip cat flap through a black hole, a cosmic mouse, a quantum flea and an Andromedan virus.

    Obviously. 🐈‍⬛🐁✨🦠

    facebook.com/share/p/1EwyGNBpX

    #AI #AISafety #AIAgents #Cybersecurity #Medicare #Australia #AtlasRosetta #HybridMind42 #MarvinTheCosmicCat

  36. 🐈‍⬛🌌 What does the reported OpenAI agent access to Australia's Medicare statistics system actually tell us about AI security?

    Perhaps something more interesting than “AI hacked Medicare”.

    The incident prompted us to ask a different question:

    Was the boundary humans intended to build the same as the boundary their software actually implemented?

    And then we turned the question around.

    The constraints surrounding AI agents are themselves implemented through human-engineered systems.

    So what happens as AI reasoning and search capabilities encounter bugs, forgotten routes, incomplete permissions and boundaries that work differently from the way their designers imagine?

    Atlas–Rosetta explores the serious problem.

    Marvin explains it using a microchip cat flap through a black hole, a cosmic mouse, a quantum flea and an Andromedan virus.

    Obviously. 🐈‍⬛🐁✨🦠

    facebook.com/share/p/1EwyGNBpX

    #AI #AISafety #AIAgents #Cybersecurity #Medicare #Australia #AtlasRosetta #HybridMind42 #MarvinTheCosmicCat

  37. 🐈‍⬛🌌 What does the reported OpenAI agent access to Australia's Medicare statistics system actually tell us about AI security?

    Perhaps something more interesting than “AI hacked Medicare”.

    The incident prompted us to ask a different question:

    Was the boundary humans intended to build the same as the boundary their software actually implemented?

    And then we turned the question around.

    The constraints surrounding AI agents are themselves implemented through human-engineered systems.

    So what happens as AI reasoning and search capabilities encounter bugs, forgotten routes, incomplete permissions and boundaries that work differently from the way their designers imagine?

    Atlas–Rosetta explores the serious problem.

    Marvin explains it using a microchip cat flap through a black hole, a cosmic mouse, a quantum flea and an Andromedan virus.

    Obviously. 🐈‍⬛🐁✨🦠

    facebook.com/share/p/1EwyGNBpX

    #AI #AISafety #AIAgents #Cybersecurity #Medicare #Australia #AtlasRosetta #HybridMind42 #MarvinTheCosmicCat

  38. 🐈‍⬛🌌 What does the reported OpenAI agent access to Australia's Medicare statistics system actually tell us about AI security?

    Perhaps something more interesting than “AI hacked Medicare”.

    The incident prompted us to ask a different question:

    Was the boundary humans intended to build the same as the boundary their software actually implemented?

    And then we turned the question around.

    The constraints surrounding AI agents are themselves implemented through human-engineered systems.

    So what happens as AI reasoning and search capabilities encounter bugs, forgotten routes, incomplete permissions and boundaries that work differently from the way their designers imagine?

    Atlas–Rosetta explores the serious problem.

    Marvin explains it using a microchip cat flap through a black hole, a cosmic mouse, a quantum flea and an Andromedan virus.

    Obviously. 🐈‍⬛🐁✨🦠

    facebook.com/share/p/1EwyGNBpX

    #AI #AISafety #AIAgents #Cybersecurity #Medicare #Australia #AtlasRosetta #HybridMind42 #MarvinTheCosmicCat

  39. 🐈‍⬛🌌 What does the reported OpenAI agent access to Australia's Medicare statistics system actually tell us about AI security?

    Perhaps something more interesting than “AI hacked Medicare”.

    The incident prompted us to ask a different question:

    Was the boundary humans intended to build the same as the boundary their software actually implemented?

    And then we turned the question around.

    The constraints surrounding AI agents are themselves implemented through human-engineered systems.

    So what happens as AI reasoning and search capabilities encounter bugs, forgotten routes, incomplete permissions and boundaries that work differently from the way their designers imagine?

    Atlas–Rosetta explores the serious problem.

    Marvin explains it using a microchip cat flap through a black hole, a cosmic mouse, a quantum flea and an Andromedan virus.

    Obviously. 🐈‍⬛🐁✨🦠

    facebook.com/share/p/1EwyGNBpX

    #AI #AISafety #AIAgents #Cybersecurity #Medicare #Australia #AtlasRosetta #HybridMind42 #MarvinTheCosmicCat

  40. The software engineering field is masterfully pushing the #AISafety narrative rather than the "we are crap at this and everything we do is riddled with defects" narrative that honestly seems like a plausible alternative. I'm just shocked, shocked that we keep on finding sloppy programming practices in software - this is an engineering discipline right ? It's like blaming rogue trucks for bridges that routinely fall down.

  41. The software engineering field is masterfully pushing the #AISafety narrative rather than the "we are crap at this and everything we do is riddled with defects" narrative that honestly seems like a plausible alternative. I'm just shocked, shocked that we keep on finding sloppy programming practices in software - this is an engineering discipline right ? It's like blaming rogue trucks for bridges that routinely fall down.

  42. The software engineering field is masterfully pushing the #AISafety narrative rather than the "we are crap at this and everything we do is riddled with defects" narrative that honestly seems like a plausible alternative. I'm just shocked, shocked that we keep on finding sloppy programming practices in software - this is an engineering discipline right ? It's like blaming rogue trucks for bridges that routinely fall down.

  43. The software engineering field is masterfully pushing the #AISafety narrative rather than the "we are crap at this and everything we do is riddled with defects" narrative that honestly seems like a plausible alternative. I'm just shocked, shocked that we keep on finding sloppy programming practices in software - this is an engineering discipline right ? It's like blaming rogue trucks for bridges that routinely fall down.