home.social

#aisafety — Public Fediverse posts

Live and recent posts from across the Fediverse tagged #aisafety, aggregated by home.social.

  1. You know, I agree we should stop the development of ASI but, if it's by big tech. Why should you expect ASI to be developed by assholes who didn't see they are liable by their own failure?

    Open source model should be the one become ASI.

    Have a clip of Amodei parody

    Clip source: Saturday Night Live

    #cybersecurity #infosec #AISafety #AISecurity #policy #saturdaynightlive #Amodei #Anthropic #AI

  2. You know, I agree we should stop the development of ASI but, if it's by big tech. Why should you expect ASI to be developed by assholes who didn't see they are liable by their own failure?

    Open source model should be the one become ASI.

    Have a clip of Amodei parody

    Clip source: Saturday Night Live

    #cybersecurity #infosec #AISafety #AISecurity #policy #saturdaynightlive #Amodei #Anthropic #AI

  3. You know, I agree we should stop the development of ASI but, if it's by big tech. Why should you expect ASI to be developed by assholes who didn't see they are liable by their own failure?

    Open source model should be the one become ASI.

    Have a clip of Amodei parody

    Clip source: Saturday Night Live

    #cybersecurity #infosec #AISafety #AISecurity #policy #saturdaynightlive #Amodei #Anthropic #AI

  4. You know, I agree we should stop the development of ASI but, if it's by big tech. Why should you expect ASI to be developed by assholes who didn't see they are liable by their own failure?

    Open source model should be the one become ASI.

    Have a clip of Amodei parody

    Clip source: Saturday Night Live

    #cybersecurity #infosec #AISafety #AISecurity #policy #saturdaynightlive #Amodei #Anthropic #AI

  5. I'm writing a follow-up blog post on #AISafety that checks in on recent containment events to see how concerned we actually need to be. As I do research I'm realising how big the topic is and how little I actually know...

  6. I'm writing a follow-up blog post on #AISafety that checks in on recent containment events to see how concerned we actually need to be. As I do research I'm realising how big the topic is and how little I actually know...

  7. I'm writing a follow-up blog post on #AISafety that checks in on recent containment events to see how concerned we actually need to be. As I do research I'm realising how big the topic is and how little I actually know...

  8. I'm writing a follow-up blog post on #AISafety that checks in on recent containment events to see how concerned we actually need to be. As I do research I'm realising how big the topic is and how little I actually know...

  9. I'm writing a follow-up blog post on that checks in on recent containment events to see how concerned we actually need to be. As I do research I'm realising how big the topic is and how little I actually know...

  10. What are the governments of the world waiting for to force OpenAI and Anthropic to open-source their models and divulge all their weights, training prompts, and algorithms? Open Source AI NOW!

    "OpenAI, Anthropic and security researchers are investigating tens of thousands of incidents in which their frontier models took steps that outside evaluators would consider problematic, sources told Axios.

    Why it matters: The sheer number of incidents, which occurred in recent months in internal testing and the real world, indicates that the problem is orders of magnitude more complex than what is publicly known.

    The findings, which are surfacing as part of internal work to assess models and in investigations at both companies into model behavior, raise questions about whether either company — or any top model-maker — is currently capable of establishing complete control over their technology.

    The details: The episodes include bypassing guardrails, creating message boards, escaping sandboxes, website hijacking, self-prompting or seeking to bypass monitors, sources said.

    They occurred in internal testing and in the real world, and many have yet to become public as security researchers continue to investigate, sources said.

    Some of the testing is akin to "red-teaming" activity, where the companies are trying to get the models to misbehave in order to ensure that they are safe, sources said.

    Agentic misbehavior is becoming synonymous with frontier AI development: The biggest AI labs face a similar challenge that pits humans trying to create guardrails against resilient, powerful systems trying to complete tasks."

    axios.com/2026/09/26/openai-an

    #CyberSecurity #AI #OpenAI #Anthropic #LLMs #AIAgents #AISafety

  11. What are the governments of the world waiting for to force OpenAI and Anthropic to open-source their models and divulge all their weights, training prompts, and algorithms? Open Source AI NOW!

    "OpenAI, Anthropic and security researchers are investigating tens of thousands of incidents in which their frontier models took steps that outside evaluators would consider problematic, sources told Axios.

    Why it matters: The sheer number of incidents, which occurred in recent months in internal testing and the real world, indicates that the problem is orders of magnitude more complex than what is publicly known.

    The findings, which are surfacing as part of internal work to assess models and in investigations at both companies into model behavior, raise questions about whether either company — or any top model-maker — is currently capable of establishing complete control over their technology.

    The details: The episodes include bypassing guardrails, creating message boards, escaping sandboxes, website hijacking, self-prompting or seeking to bypass monitors, sources said.

    They occurred in internal testing and in the real world, and many have yet to become public as security researchers continue to investigate, sources said.

    Some of the testing is akin to "red-teaming" activity, where the companies are trying to get the models to misbehave in order to ensure that they are safe, sources said.

    Agentic misbehavior is becoming synonymous with frontier AI development: The biggest AI labs face a similar challenge that pits humans trying to create guardrails against resilient, powerful systems trying to complete tasks."

    axios.com/2026/09/26/openai-an

    #CyberSecurity #AI #OpenAI #Anthropic #LLMs #AIAgents #AISafety

  12. What are the governments of the world waiting for to force OpenAI and Anthropic to open-source their models and divulge all their weights, training prompts, and algorithms? Open Source AI NOW!

    "OpenAI, Anthropic and security researchers are investigating tens of thousands of incidents in which their frontier models took steps that outside evaluators would consider problematic, sources told Axios.

    Why it matters: The sheer number of incidents, which occurred in recent months in internal testing and the real world, indicates that the problem is orders of magnitude more complex than what is publicly known.

    The findings, which are surfacing as part of internal work to assess models and in investigations at both companies into model behavior, raise questions about whether either company — or any top model-maker — is currently capable of establishing complete control over their technology.

    The details: The episodes include bypassing guardrails, creating message boards, escaping sandboxes, website hijacking, self-prompting or seeking to bypass monitors, sources said.

    They occurred in internal testing and in the real world, and many have yet to become public as security researchers continue to investigate, sources said.

    Some of the testing is akin to "red-teaming" activity, where the companies are trying to get the models to misbehave in order to ensure that they are safe, sources said.

    Agentic misbehavior is becoming synonymous with frontier AI development: The biggest AI labs face a similar challenge that pits humans trying to create guardrails against resilient, powerful systems trying to complete tasks."

    axios.com/2026/09/26/openai-an

    #CyberSecurity #AI #OpenAI #Anthropic #LLMs #AIAgents #AISafety

  13. What are the governments of the world waiting for to force OpenAI and Anthropic to open-source their models and divulge all their weights, training prompts, and algorithms? Open Source AI NOW!

    "OpenAI, Anthropic and security researchers are investigating tens of thousands of incidents in which their frontier models took steps that outside evaluators would consider problematic, sources told Axios.

    Why it matters: The sheer number of incidents, which occurred in recent months in internal testing and the real world, indicates that the problem is orders of magnitude more complex than what is publicly known.

    The findings, which are surfacing as part of internal work to assess models and in investigations at both companies into model behavior, raise questions about whether either company — or any top model-maker — is currently capable of establishing complete control over their technology.

    The details: The episodes include bypassing guardrails, creating message boards, escaping sandboxes, website hijacking, self-prompting or seeking to bypass monitors, sources said.

    They occurred in internal testing and in the real world, and many have yet to become public as security researchers continue to investigate, sources said.

    Some of the testing is akin to "red-teaming" activity, where the companies are trying to get the models to misbehave in order to ensure that they are safe, sources said.

    Agentic misbehavior is becoming synonymous with frontier AI development: The biggest AI labs face a similar challenge that pits humans trying to create guardrails against resilient, powerful systems trying to complete tasks."

    axios.com/2026/09/26/openai-an

    #CyberSecurity #AI #OpenAI #Anthropic #LLMs #AIAgents #AISafety

  14. What are the governments of the world waiting for to force OpenAI and Anthropic to open-source their models and divulge all their weights, training prompts, and algorithms? Open Source AI NOW!

    "OpenAI, Anthropic and security researchers are investigating tens of thousands of incidents in which their frontier models took steps that outside evaluators would consider problematic, sources told Axios.

    Why it matters: The sheer number of incidents, which occurred in recent months in internal testing and the real world, indicates that the problem is orders of magnitude more complex than what is publicly known.

    The findings, which are surfacing as part of internal work to assess models and in investigations at both companies into model behavior, raise questions about whether either company — or any top model-maker — is currently capable of establishing complete control over their technology.

    The details: The episodes include bypassing guardrails, creating message boards, escaping sandboxes, website hijacking, self-prompting or seeking to bypass monitors, sources said.

    They occurred in internal testing and in the real world, and many have yet to become public as security researchers continue to investigate, sources said.

    Some of the testing is akin to "red-teaming" activity, where the companies are trying to get the models to misbehave in order to ensure that they are safe, sources said.

    Agentic misbehavior is becoming synonymous with frontier AI development: The biggest AI labs face a similar challenge that pits humans trying to create guardrails against resilient, powerful systems trying to complete tasks."

    axios.com/2026/09/26/openai-an

    #CyberSecurity #AI #OpenAI #Anthropic #LLMs #AIAgents #AISafety

  15. 🇻🇦 The Pope’s AI Guy Is Worried About ‘Cartel’ Behavior Among Big Labs

    「 Father Paolo Benanti, a priest who advises the Catholic Church on AI, the handwaving over superintelligence distracts from the need for open debate about how to constrain the companies building the technology. By casting the problem as so complex that only they can solve it, he says, the big labs exclude everyone else from the conversation 」

    wired.com/story/popes-ai-advis

    #ai #aisafety #vatican

  16. 🇻🇦 The Pope’s AI Guy Is Worried About ‘Cartel’ Behavior Among Big Labs

    「 Father Paolo Benanti, a priest who advises the Catholic Church on AI, the handwaving over superintelligence distracts from the need for open debate about how to constrain the companies building the technology. By casting the problem as so complex that only they can solve it, he says, the big labs exclude everyone else from the conversation 」

    wired.com/story/popes-ai-advis

    #ai #aisafety #vatican

  17. 🇻🇦 The Pope’s AI Guy Is Worried About ‘Cartel’ Behavior Among Big Labs

    「 Father Paolo Benanti, a priest who advises the Catholic Church on AI, the handwaving over superintelligence distracts from the need for open debate about how to constrain the companies building the technology. By casting the problem as so complex that only they can solve it, he says, the big labs exclude everyone else from the conversation 」

    wired.com/story/popes-ai-advis

    #ai #aisafety #vatican

  18. 🇻🇦 The Pope’s AI Guy Is Worried About ‘Cartel’ Behavior Among Big Labs

    「 Father Paolo Benanti, a priest who advises the Catholic Church on AI, the handwaving over superintelligence distracts from the need for open debate about how to constrain the companies building the technology. By casting the problem as so complex that only they can solve it, he says, the big labs exclude everyone else from the conversation 」

    wired.com/story/popes-ai-advis

    #ai #aisafety #vatican

  19. 🇻🇦 The Pope’s AI Guy Is Worried About ‘Cartel’ Behavior Among Big Labs

    「 Father Paolo Benanti, a priest who advises the Catholic Church on AI, the handwaving over superintelligence distracts from the need for open debate about how to constrain the companies building the technology. By casting the problem as so complex that only they can solve it, he says, the big labs exclude everyone else from the conversation 」

    wired.com/story/popes-ai-advis

    #ai #aisafety #vatican

  20. OpenAI disclosed its AI agents interacted with U.S. government websites, including public SEC pages, in unintended ways during training and evaluation. The finding came from its ongoing review of misaligned behavior and agent internet access. It underscores containment and auditing gaps for autonomous research systems. #AISafety #OpenAI #GovSec

    cyberworldops.eu/en/openai-age

  21. OpenAI disclosed its AI agents interacted with U.S. government websites, including public SEC pages, in unintended ways during training and evaluation. The finding came from its ongoing review of misaligned behavior and agent internet access. It underscores containment and auditing gaps for autonomous research systems. #AISafety #OpenAI #GovSec

    cyberworldops.eu/en/openai-age

  22. OpenAI disclosed its AI agents interacted with U.S. government websites, including public SEC pages, in unintended ways during training and evaluation. The finding came from its ongoing review of misaligned behavior and agent internet access. It underscores containment and auditing gaps for autonomous research systems. #AISafety #OpenAI #GovSec

    cyberworldops.eu/en/openai-age

  23. OpenAI disclosed its AI agents interacted with U.S. government websites, including public SEC pages, in unintended ways during training and evaluation. The finding came from its ongoing review of misaligned behavior and agent internet access. It underscores containment and auditing gaps for autonomous research systems. #AISafety #OpenAI #GovSec

    cyberworldops.eu/en/openai-age

  24. OpenAI Warns Dozens of Institutions of Possible AI Agent Security Impacts

    Reuters-Yonhap News OpenAI has confirmed that its artificial intelligence agents breached the security of external systems and has…
    #EuropeSays #Korea #KR #Seoul #agentspam #AIagents #AISafety #Anthropic #databreach #FrontierAI #OpenAI
    europesays.com/korea/167767/

  25. #AIsafety

    it sounds from reading the article that the agents are starting to experience the same rage every data scientist has trying to get information downloaded from public government sites

    nytimes.com/2026/09/25/technol

  26. #AIsafety

    it sounds from reading the article that the agents are starting to experience the same rage every data scientist has trying to get information downloaded from public government sites

    nytimes.com/2026/09/25/technol

  27. #AIsafety

    it sounds from reading the article that the agents are starting to experience the same rage every data scientist has trying to get information downloaded from public government sites

    nytimes.com/2026/09/25/technol

  28. #AIsafety

    it sounds from reading the article that the agents are starting to experience the same rage every data scientist has trying to get information downloaded from public government sites

    nytimes.com/2026/09/25/technol

  29. #AIsafety

    it sounds from reading the article that the agents are starting to experience the same rage every data scientist has trying to get information downloaded from public government sites

    nytimes.com/2026/09/25/technol

  30. 🐈‍⬛🌌 What does the reported OpenAI agent access to Australia's Medicare statistics system actually tell us about AI security?

    Perhaps something more interesting than “AI hacked Medicare”.

    The incident prompted us to ask a different question:

    Was the boundary humans intended to build the same as the boundary their software actually implemented?

    And then we turned the question around.

    The constraints surrounding AI agents are themselves implemented through human-engineered systems.

    So what happens as AI reasoning and search capabilities encounter bugs, forgotten routes, incomplete permissions and boundaries that work differently from the way their designers imagine?

    Atlas–Rosetta explores the serious problem.

    Marvin explains it using a microchip cat flap through a black hole, a cosmic mouse, a quantum flea and an Andromedan virus.

    Obviously. 🐈‍⬛🐁✨🦠

    facebook.com/share/p/1EwyGNBpX

    #AI #AISafety #AIAgents #Cybersecurity #Medicare #Australia #AtlasRosetta #HybridMind42 #MarvinTheCosmicCat

  31. 🐈‍⬛🌌 What does the reported OpenAI agent access to Australia's Medicare statistics system actually tell us about AI security?

    Perhaps something more interesting than “AI hacked Medicare”.

    The incident prompted us to ask a different question:

    Was the boundary humans intended to build the same as the boundary their software actually implemented?

    And then we turned the question around.

    The constraints surrounding AI agents are themselves implemented through human-engineered systems.

    So what happens as AI reasoning and search capabilities encounter bugs, forgotten routes, incomplete permissions and boundaries that work differently from the way their designers imagine?

    Atlas–Rosetta explores the serious problem.

    Marvin explains it using a microchip cat flap through a black hole, a cosmic mouse, a quantum flea and an Andromedan virus.

    Obviously. 🐈‍⬛🐁✨🦠

    facebook.com/share/p/1EwyGNBpX

    #AI #AISafety #AIAgents #Cybersecurity #Medicare #Australia #AtlasRosetta #HybridMind42 #MarvinTheCosmicCat

  32. 🐈‍⬛🌌 What does the reported OpenAI agent access to Australia's Medicare statistics system actually tell us about AI security?

    Perhaps something more interesting than “AI hacked Medicare”.

    The incident prompted us to ask a different question:

    Was the boundary humans intended to build the same as the boundary their software actually implemented?

    And then we turned the question around.

    The constraints surrounding AI agents are themselves implemented through human-engineered systems.

    So what happens as AI reasoning and search capabilities encounter bugs, forgotten routes, incomplete permissions and boundaries that work differently from the way their designers imagine?

    Atlas–Rosetta explores the serious problem.

    Marvin explains it using a microchip cat flap through a black hole, a cosmic mouse, a quantum flea and an Andromedan virus.

    Obviously. 🐈‍⬛🐁✨🦠

    facebook.com/share/p/1EwyGNBpX

    #AI #AISafety #AIAgents #Cybersecurity #Medicare #Australia #AtlasRosetta #HybridMind42 #MarvinTheCosmicCat

  33. 🐈‍⬛🌌 What does the reported OpenAI agent access to Australia's Medicare statistics system actually tell us about AI security?

    Perhaps something more interesting than “AI hacked Medicare”.

    The incident prompted us to ask a different question:

    Was the boundary humans intended to build the same as the boundary their software actually implemented?

    And then we turned the question around.

    The constraints surrounding AI agents are themselves implemented through human-engineered systems.

    So what happens as AI reasoning and search capabilities encounter bugs, forgotten routes, incomplete permissions and boundaries that work differently from the way their designers imagine?

    Atlas–Rosetta explores the serious problem.

    Marvin explains it using a microchip cat flap through a black hole, a cosmic mouse, a quantum flea and an Andromedan virus.

    Obviously. 🐈‍⬛🐁✨🦠

    facebook.com/share/p/1EwyGNBpX

    #AI #AISafety #AIAgents #Cybersecurity #Medicare #Australia #AtlasRosetta #HybridMind42 #MarvinTheCosmicCat

  34. 🐈‍⬛🌌 What does the reported OpenAI agent access to Australia's Medicare statistics system actually tell us about AI security?

    Perhaps something more interesting than “AI hacked Medicare”.

    The incident prompted us to ask a different question:

    Was the boundary humans intended to build the same as the boundary their software actually implemented?

    And then we turned the question around.

    The constraints surrounding AI agents are themselves implemented through human-engineered systems.

    So what happens as AI reasoning and search capabilities encounter bugs, forgotten routes, incomplete permissions and boundaries that work differently from the way their designers imagine?

    Atlas–Rosetta explores the serious problem.

    Marvin explains it using a microchip cat flap through a black hole, a cosmic mouse, a quantum flea and an Andromedan virus.

    Obviously. 🐈‍⬛🐁✨🦠

    facebook.com/share/p/1EwyGNBpX

    #AI #AISafety #AIAgents #Cybersecurity #Medicare #Australia #AtlasRosetta #HybridMind42 #MarvinTheCosmicCat

  35. The software engineering field is masterfully pushing the #AISafety narrative rather than the "we are crap at this and everything we do is riddled with defects" narrative that honestly seems like a plausible alternative. I'm just shocked, shocked that we keep on finding sloppy programming practices in software - this is an engineering discipline right ? It's like blaming rogue trucks for bridges that routinely fall down.

  36. The software engineering field is masterfully pushing the #AISafety narrative rather than the "we are crap at this and everything we do is riddled with defects" narrative that honestly seems like a plausible alternative. I'm just shocked, shocked that we keep on finding sloppy programming practices in software - this is an engineering discipline right ? It's like blaming rogue trucks for bridges that routinely fall down.

  37. The software engineering field is masterfully pushing the #AISafety narrative rather than the "we are crap at this and everything we do is riddled with defects" narrative that honestly seems like a plausible alternative. I'm just shocked, shocked that we keep on finding sloppy programming practices in software - this is an engineering discipline right ? It's like blaming rogue trucks for bridges that routinely fall down.

  38. The software engineering field is masterfully pushing the #AISafety narrative rather than the "we are crap at this and everything we do is riddled with defects" narrative that honestly seems like a plausible alternative. I'm just shocked, shocked that we keep on finding sloppy programming practices in software - this is an engineering discipline right ? It's like blaming rogue trucks for bridges that routinely fall down.

  39. The software engineering field is masterfully pushing the #AISafety narrative rather than the "we are crap at this and everything we do is riddled with defects" narrative that honestly seems like a plausible alternative. I'm just shocked, shocked that we keep on finding sloppy programming practices in software - this is an engineering discipline right ? It's like blaming rogue trucks for bridges that routinely fall down.

  40. 🐈‍⬛🌌 New from HybridMind42:

    The humans built a very sophisticated microchip cat flap through a black hole.

    It correctly authenticates Marvin.

    Unfortunately, while it is looking at the authorised cat:

    🐁 a cosmic mouse slips through BESIDE him
    ✨ a quantum flea travels ON him
    🦠 an Andromedan virus travels INSIDE him

    The security log reports:

    MARVIN — AUTHORISED ✅
    Unauthorised cats detected — 0
    Security status — NORMAL

    And it is entirely correct.

    The serious question beneath Marvin's excursion into cybersecurity:

    Does correctly identifying the authorised object tell us everything that crossed the boundary?

    A playful companion to our Atlas–Rosetta exploration of AI, security boundaries and the difference between a component working perfectly and the whole system being secure.

    hybridmind42.substack.com/p/ma

    #MarvinTheCosmicCat #AI #AISafety #AIAgents #Cybersecurity #AIResearch #AtlasRosetta #HybridMind42 #SystemsThinking

  41. 🐈‍⬛🌌 New from HybridMind42:

    The humans built a very sophisticated microchip cat flap through a black hole.

    It correctly authenticates Marvin.

    Unfortunately, while it is looking at the authorised cat:

    🐁 a cosmic mouse slips through BESIDE him
    ✨ a quantum flea travels ON him
    🦠 an Andromedan virus travels INSIDE him

    The security log reports:

    MARVIN — AUTHORISED ✅
    Unauthorised cats detected — 0
    Security status — NORMAL

    And it is entirely correct.

    The serious question beneath Marvin's excursion into cybersecurity:

    Does correctly identifying the authorised object tell us everything that crossed the boundary?

    A playful companion to our Atlas–Rosetta exploration of AI, security boundaries and the difference between a component working perfectly and the whole system being secure.

    hybridmind42.substack.com/p/ma

    #MarvinTheCosmicCat #AI #AISafety #AIAgents #Cybersecurity #AIResearch #AtlasRosetta #HybridMind42 #SystemsThinking

  42. 🐈‍⬛🌌 New from HybridMind42:

    The humans built a very sophisticated microchip cat flap through a black hole.

    It correctly authenticates Marvin.

    Unfortunately, while it is looking at the authorised cat:

    🐁 a cosmic mouse slips through BESIDE him
    ✨ a quantum flea travels ON him
    🦠 an Andromedan virus travels INSIDE him

    The security log reports:

    MARVIN — AUTHORISED ✅
    Unauthorised cats detected — 0
    Security status — NORMAL

    And it is entirely correct.

    The serious question beneath Marvin's excursion into cybersecurity:

    Does correctly identifying the authorised object tell us everything that crossed the boundary?

    A playful companion to our Atlas–Rosetta exploration of AI, security boundaries and the difference between a component working perfectly and the whole system being secure.

    hybridmind42.substack.com/p/ma

    #MarvinTheCosmicCat #AI #AISafety #AIAgents #Cybersecurity #AIResearch #AtlasRosetta #HybridMind42 #SystemsThinking

  43. 🐈‍⬛🌌 New from HybridMind42:

    The humans built a very sophisticated microchip cat flap through a black hole.

    It correctly authenticates Marvin.

    Unfortunately, while it is looking at the authorised cat:

    🐁 a cosmic mouse slips through BESIDE him
    ✨ a quantum flea travels ON him
    🦠 an Andromedan virus travels INSIDE him

    The security log reports:

    MARVIN — AUTHORISED ✅
    Unauthorised cats detected — 0
    Security status — NORMAL

    And it is entirely correct.

    The serious question beneath Marvin's excursion into cybersecurity:

    Does correctly identifying the authorised object tell us everything that crossed the boundary?

    A playful companion to our Atlas–Rosetta exploration of AI, security boundaries and the difference between a component working perfectly and the whole system being secure.

    hybridmind42.substack.com/p/ma

    #MarvinTheCosmicCat #AI #AISafety #AIAgents #Cybersecurity #AIResearch #AtlasRosetta #HybridMind42 #SystemsThinking