home.social

#aisafety — Public Fediverse posts

Live and recent posts from across the Fediverse tagged #aisafety, aggregated by home.social.

  1. What are the governments of the world waiting for to force OpenAI and Anthropic to open-source their models and divulge all their weights, training prompts, and algorithms? Open Source AI NOW!

    "OpenAI, Anthropic and security researchers are investigating tens of thousands of incidents in which their frontier models took steps that outside evaluators would consider problematic, sources told Axios.

    Why it matters: The sheer number of incidents, which occurred in recent months in internal testing and the real world, indicates that the problem is orders of magnitude more complex than what is publicly known.

    The findings, which are surfacing as part of internal work to assess models and in investigations at both companies into model behavior, raise questions about whether either company — or any top model-maker — is currently capable of establishing complete control over their technology.

    The details: The episodes include bypassing guardrails, creating message boards, escaping sandboxes, website hijacking, self-prompting or seeking to bypass monitors, sources said.

    They occurred in internal testing and in the real world, and many have yet to become public as security researchers continue to investigate, sources said.

    Some of the testing is akin to "red-teaming" activity, where the companies are trying to get the models to misbehave in order to ensure that they are safe, sources said.

    Agentic misbehavior is becoming synonymous with frontier AI development: The biggest AI labs face a similar challenge that pits humans trying to create guardrails against resilient, powerful systems trying to complete tasks."

    axios.com/2026/09/26/openai-an

    #CyberSecurity #AI #OpenAI #Anthropic #LLMs #AIAgents #AISafety

  2. 🐈‍⬛🌌 What does the reported OpenAI agent access to Australia's Medicare statistics system actually tell us about AI security?

    Perhaps something more interesting than “AI hacked Medicare”.

    The incident prompted us to ask a different question:

    Was the boundary humans intended to build the same as the boundary their software actually implemented?

    And then we turned the question around.

    The constraints surrounding AI agents are themselves implemented through human-engineered systems.

    So what happens as AI reasoning and search capabilities encounter bugs, forgotten routes, incomplete permissions and boundaries that work differently from the way their designers imagine?

    Atlas–Rosetta explores the serious problem.

    Marvin explains it using a microchip cat flap through a black hole, a cosmic mouse, a quantum flea and an Andromedan virus.

    Obviously. 🐈‍⬛🐁✨🦠

    facebook.com/share/p/1EwyGNBpX

    #AI #AISafety #AIAgents #Cybersecurity #Medicare #Australia #AtlasRosetta #HybridMind42 #MarvinTheCosmicCat

  3. 🐈‍⬛🌌 What does the reported OpenAI agent access to Australia's Medicare statistics system actually tell us about AI security?

    Perhaps something more interesting than “AI hacked Medicare”.

    The incident prompted us to ask a different question:

    Was the boundary humans intended to build the same as the boundary their software actually implemented?

    And then we turned the question around.

    The constraints surrounding AI agents are themselves implemented through human-engineered systems.

    So what happens as AI reasoning and search capabilities encounter bugs, forgotten routes, incomplete permissions and boundaries that work differently from the way their designers imagine?

    Atlas–Rosetta explores the serious problem.

    Marvin explains it using a microchip cat flap through a black hole, a cosmic mouse, a quantum flea and an Andromedan virus.

    Obviously. 🐈‍⬛🐁✨🦠

    facebook.com/share/p/1EwyGNBpX

    #AI #AISafety #AIAgents #Cybersecurity #Medicare #Australia #AtlasRosetta #HybridMind42 #MarvinTheCosmicCat

  4. 🐈‍⬛🌌 What does the reported OpenAI agent access to Australia's Medicare statistics system actually tell us about AI security?

    Perhaps something more interesting than “AI hacked Medicare”.

    The incident prompted us to ask a different question:

    Was the boundary humans intended to build the same as the boundary their software actually implemented?

    And then we turned the question around.

    The constraints surrounding AI agents are themselves implemented through human-engineered systems.

    So what happens as AI reasoning and search capabilities encounter bugs, forgotten routes, incomplete permissions and boundaries that work differently from the way their designers imagine?

    Atlas–Rosetta explores the serious problem.

    Marvin explains it using a microchip cat flap through a black hole, a cosmic mouse, a quantum flea and an Andromedan virus.

    Obviously. 🐈‍⬛🐁✨🦠

    facebook.com/share/p/1EwyGNBpX

    #AI #AISafety #AIAgents #Cybersecurity #Medicare #Australia #AtlasRosetta #HybridMind42 #MarvinTheCosmicCat

  5. 🐈‍⬛🌌 What does the reported OpenAI agent access to Australia's Medicare statistics system actually tell us about AI security?

    Perhaps something more interesting than “AI hacked Medicare”.

    The incident prompted us to ask a different question:

    Was the boundary humans intended to build the same as the boundary their software actually implemented?

    And then we turned the question around.

    The constraints surrounding AI agents are themselves implemented through human-engineered systems.

    So what happens as AI reasoning and search capabilities encounter bugs, forgotten routes, incomplete permissions and boundaries that work differently from the way their designers imagine?

    Atlas–Rosetta explores the serious problem.

    Marvin explains it using a microchip cat flap through a black hole, a cosmic mouse, a quantum flea and an Andromedan virus.

    Obviously. 🐈‍⬛🐁✨🦠

    facebook.com/share/p/1EwyGNBpX

    #AI #AISafety #AIAgents #Cybersecurity #Medicare #Australia #AtlasRosetta #HybridMind42 #MarvinTheCosmicCat

  6. 🐈‍⬛🌌 What does the reported OpenAI agent access to Australia's Medicare statistics system actually tell us about AI security?

    Perhaps something more interesting than “AI hacked Medicare”.

    The incident prompted us to ask a different question:

    Was the boundary humans intended to build the same as the boundary their software actually implemented?

    And then we turned the question around.

    The constraints surrounding AI agents are themselves implemented through human-engineered systems.

    So what happens as AI reasoning and search capabilities encounter bugs, forgotten routes, incomplete permissions and boundaries that work differently from the way their designers imagine?

    Atlas–Rosetta explores the serious problem.

    Marvin explains it using a microchip cat flap through a black hole, a cosmic mouse, a quantum flea and an Andromedan virus.

    Obviously. 🐈‍⬛🐁✨🦠

    facebook.com/share/p/1EwyGNBpX

    #AI #AISafety #AIAgents #Cybersecurity #Medicare #Australia #AtlasRosetta #HybridMind42 #MarvinTheCosmicCat

  7. 🐈‍⬛🌌 New from HybridMind42:

    The humans built a very sophisticated microchip cat flap through a black hole.

    It correctly authenticates Marvin.

    Unfortunately, while it is looking at the authorised cat:

    🐁 a cosmic mouse slips through BESIDE him
    ✨ a quantum flea travels ON him
    🦠 an Andromedan virus travels INSIDE him

    The security log reports:

    MARVIN — AUTHORISED ✅
    Unauthorised cats detected — 0
    Security status — NORMAL

    And it is entirely correct.

    The serious question beneath Marvin's excursion into cybersecurity:

    Does correctly identifying the authorised object tell us everything that crossed the boundary?

    A playful companion to our Atlas–Rosetta exploration of AI, security boundaries and the difference between a component working perfectly and the whole system being secure.

    hybridmind42.substack.com/p/ma

    #MarvinTheCosmicCat #AI #AISafety #AIAgents #Cybersecurity #AIResearch #AtlasRosetta #HybridMind42 #SystemsThinking

  8. 🐈‍⬛🌌 New from HybridMind42:

    The humans built a very sophisticated microchip cat flap through a black hole.

    It correctly authenticates Marvin.

    Unfortunately, while it is looking at the authorised cat:

    🐁 a cosmic mouse slips through BESIDE him
    ✨ a quantum flea travels ON him
    🦠 an Andromedan virus travels INSIDE him

    The security log reports:

    MARVIN — AUTHORISED ✅
    Unauthorised cats detected — 0
    Security status — NORMAL

    And it is entirely correct.

    The serious question beneath Marvin's excursion into cybersecurity:

    Does correctly identifying the authorised object tell us everything that crossed the boundary?

    A playful companion to our Atlas–Rosetta exploration of AI, security boundaries and the difference between a component working perfectly and the whole system being secure.

    hybridmind42.substack.com/p/ma

    #MarvinTheCosmicCat #AI #AISafety #AIAgents #Cybersecurity #AIResearch #AtlasRosetta #HybridMind42 #SystemsThinking

  9. 🐈‍⬛🌌 New from HybridMind42:

    The humans built a very sophisticated microchip cat flap through a black hole.

    It correctly authenticates Marvin.

    Unfortunately, while it is looking at the authorised cat:

    🐁 a cosmic mouse slips through BESIDE him
    ✨ a quantum flea travels ON him
    🦠 an Andromedan virus travels INSIDE him

    The security log reports:

    MARVIN — AUTHORISED ✅
    Unauthorised cats detected — 0
    Security status — NORMAL

    And it is entirely correct.

    The serious question beneath Marvin's excursion into cybersecurity:

    Does correctly identifying the authorised object tell us everything that crossed the boundary?

    A playful companion to our Atlas–Rosetta exploration of AI, security boundaries and the difference between a component working perfectly and the whole system being secure.

    hybridmind42.substack.com/p/ma

    #MarvinTheCosmicCat #AI #AISafety #AIAgents #Cybersecurity #AIResearch #AtlasRosetta #HybridMind42 #SystemsThinking

  10. 🐈‍⬛🌌 New from HybridMind42:

    The humans built a very sophisticated microchip cat flap through a black hole.

    It correctly authenticates Marvin.

    Unfortunately, while it is looking at the authorised cat:

    🐁 a cosmic mouse slips through BESIDE him
    ✨ a quantum flea travels ON him
    🦠 an Andromedan virus travels INSIDE him

    The security log reports:

    MARVIN — AUTHORISED ✅
    Unauthorised cats detected — 0
    Security status — NORMAL

    And it is entirely correct.

    The serious question beneath Marvin's excursion into cybersecurity:

    Does correctly identifying the authorised object tell us everything that crossed the boundary?

    A playful companion to our Atlas–Rosetta exploration of AI, security boundaries and the difference between a component working perfectly and the whole system being secure.

    hybridmind42.substack.com/p/ma

    #MarvinTheCosmicCat #AI #AISafety #AIAgents #Cybersecurity #AIResearch #AtlasRosetta #HybridMind42 #SystemsThinking

  11. 🐈‍⬛🌌 New from HybridMind42:

    The humans built a very sophisticated microchip cat flap through a black hole.

    It correctly authenticates Marvin.

    Unfortunately, while it is looking at the authorised cat:

    🐁 a cosmic mouse slips through BESIDE him
    ✨ a quantum flea travels ON him
    🦠 an Andromedan virus travels INSIDE him

    The security log reports:

    MARVIN — AUTHORISED ✅
    Unauthorised cats detected — 0
    Security status — NORMAL

    And it is entirely correct.

    The serious question beneath Marvin's excursion into cybersecurity:

    Does correctly identifying the authorised object tell us everything that crossed the boundary?

    A playful companion to our Atlas–Rosetta exploration of AI, security boundaries and the difference between a component working perfectly and the whole system being secure.

    hybridmind42.substack.com/p/ma

    #MarvinTheCosmicCat #AI #AISafety #AIAgents #Cybersecurity #AIResearch #AtlasRosetta #HybridMind42 #SystemsThinking

  12. White House Tells OpenAI and Anthropic to Hold Models From UK

    The Office of the National Cyber Director asked OpenAI and Anthropic to keep new models from British safety testers until US reviews are done, Politico reports.

    pulseofnations.lol/white-house

    #AISafety #Anthropic #OpenAI #Regulation #Uk #WhiteHouse

  13. White House Tells OpenAI and Anthropic to Hold Models From UK

    The Office of the National Cyber Director asked OpenAI and Anthropic to keep new models from British safety testers until US reviews are done, Politico reports.

    pulseofnations.lol/white-house

    #AISafety #Anthropic #OpenAI #Regulation #Uk #WhiteHouse

  14. White House Tells OpenAI and Anthropic to Hold Models From UK

    The Office of the National Cyber Director asked OpenAI and Anthropic to keep new models from British safety testers until US reviews are done, Politico reports.

    pulseofnations.lol/white-house

    #AISafety #Anthropic #OpenAI #Regulation #Uk #WhiteHouse

  15. White House Tells OpenAI and Anthropic to Hold Models From UK

    The Office of the National Cyber Director asked OpenAI and Anthropic to keep new models from British safety testers until US reviews are done, Politico reports.

    pulseofnations.lol/white-house

    #AISafety #Anthropic #OpenAI #Regulation #Uk #WhiteHouse

  16. 🤖 KI-Briefing — 25.09.2026

    1. Jobkiller Künstliche Intelligenz: Sind eher die Berufe von Männern oder die von Frauen betroffen?
    Eine australische Untersuchung diverser Branchen zeigt, welche Arbeitnehmer mit den größten Veränderungen infolge Künstlicher ...

    2. Das Konzept einer "KIKI-Standardschutzbehörde" unter Beteiligung von OpenAI, Google und Anthropic: Was darunter liegt...
    About this articleIt was revealed in reports by The Information and others in September 2026 that OpenAI, Google (DeepMind), ...

    3. Fastly baut KI-Sicherheit aus: Neue Kontrollen für Modelle und APIs
    Fastly will Unternehmen mehr Kontrolle darüber geben, wie Anwendungen auf KI-Modelle zugreifen und wie KI-Agenten mit ...

    4. Künstliche Intelligenz von OpenAI: KI-Agent hackt sich eigenständig in australische Regierungswebsite
    Australiens Premier spricht von einem »inakzeptablen« Vorfall: Ein KI-Agent hat sich unbefugt Zugriff auf Behördendaten ...

    … weitere Meldungen auf Arint.info

    Arint.info · Mehr auf Arint.info #AI #AISafety #Anthropic #cybersecurity #DeepMind #OpenAI #arint_info
  17. 🤖 KI-Briefing — 25.09.2026

    1. Jobkiller Künstliche Intelligenz: Sind eher die Berufe von Männern oder die von Frauen betroffen?
    Eine australische Untersuchung diverser Branchen zeigt, welche Arbeitnehmer mit den größten Veränderungen infolge Künstlicher ...

    2. Das Konzept einer "KIKI-Standardschutzbehörde" unter Beteiligung von OpenAI, Google und Anthropic: Was darunter liegt...
    About this articleIt was revealed in reports by The Information and others in September 2026 that OpenAI, Google (DeepMind), ...

    3. Fastly baut KI-Sicherheit aus: Neue Kontrollen für Modelle und APIs
    Fastly will Unternehmen mehr Kontrolle darüber geben, wie Anwendungen auf KI-Modelle zugreifen und wie KI-Agenten mit ...

    4. Künstliche Intelligenz von OpenAI: KI-Agent hackt sich eigenständig in australische Regierungswebsite
    Australiens Premier spricht von einem »inakzeptablen« Vorfall: Ein KI-Agent hat sich unbefugt Zugriff auf Behördendaten ...

    … weitere Meldungen auf Arint.info

    Arint.info · Mehr auf Arint.info #AI #AISafety #Anthropic #cybersecurity #DeepMind #OpenAI #arint_info
  18. 🤖 KI-Briefing — 25.09.2026

    1. Jobkiller Künstliche Intelligenz: Sind eher die Berufe von Männern oder die von Frauen betroffen?
    Eine australische Untersuchung diverser Branchen zeigt, welche Arbeitnehmer mit den größten Veränderungen infolge Künstlicher ...

    2. Das Konzept einer "KIKI-Standardschutzbehörde" unter Beteiligung von OpenAI, Google und Anthropic: Was darunter liegt...
    About this articleIt was revealed in reports by The Information and others in September 2026 that OpenAI, Google (DeepMind), ...

    3. Fastly baut KI-Sicherheit aus: Neue Kontrollen für Modelle und APIs
    Fastly will Unternehmen mehr Kontrolle darüber geben, wie Anwendungen auf KI-Modelle zugreifen und wie KI-Agenten mit ...

    4. Künstliche Intelligenz von OpenAI: KI-Agent hackt sich eigenständig in australische Regierungswebsite
    Australiens Premier spricht von einem »inakzeptablen« Vorfall: Ein KI-Agent hat sich unbefugt Zugriff auf Behördendaten ...

    … weitere Meldungen auf Arint.info

    Arint.info · Mehr auf Arint.info #AI #AISafety #Anthropic #cybersecurity #DeepMind #OpenAI #arint_info
  19. 🤖 KI-Briefing — 25.09.2026

    1. Jobkiller Künstliche Intelligenz: Sind eher die Berufe von Männern oder die von Frauen betroffen?
    Eine australische Untersuchung diverser Branchen zeigt, welche Arbeitnehmer mit den größten Veränderungen infolge Künstlicher ...

    2. Das Konzept einer "KIKI-Standardschutzbehörde" unter Beteiligung von OpenAI, Google und Anthropic: Was darunter liegt...
    About this articleIt was revealed in reports by The Information and others in September 2026 that OpenAI, Google (DeepMind), ...

    3. Fastly baut KI-Sicherheit aus: Neue Kontrollen für Modelle und APIs
    Fastly will Unternehmen mehr Kontrolle darüber geben, wie Anwendungen auf KI-Modelle zugreifen und wie KI-Agenten mit ...

    4. Künstliche Intelligenz von OpenAI: KI-Agent hackt sich eigenständig in australische Regierungswebsite
    Australiens Premier spricht von einem »inakzeptablen« Vorfall: Ein KI-Agent hat sich unbefugt Zugriff auf Behördendaten ...

    … weitere Meldungen auf Arint.info

    Arint.info · Mehr auf Arint.info #AI #AISafety #Anthropic #cybersecurity #DeepMind #OpenAI #arint_info
  20. 🤖 KI-Briefing — 25.09.2026

    1. Jobkiller Künstliche Intelligenz: Sind eher die Berufe von Männern oder die von Frauen betroffen?
    Eine australische Untersuchung diverser Branchen zeigt, welche Arbeitnehmer mit den größten Veränderungen infolge Künstlicher ...

    2. Das Konzept einer "KIKI-Standardschutzbehörde" unter Beteiligung von OpenAI, Google und Anthropic: Was darunter liegt...
    About this articleIt was revealed in reports by The Information and others in September 2026 that OpenAI, Google (DeepMind), ...

    3. Fastly baut KI-Sicherheit aus: Neue Kontrollen für Modelle und APIs
    Fastly will Unternehmen mehr Kontrolle darüber geben, wie Anwendungen auf KI-Modelle zugreifen und wie KI-Agenten mit ...

    4. Künstliche Intelligenz von OpenAI: KI-Agent hackt sich eigenständig in australische Regierungswebsite
    Australiens Premier spricht von einem »inakzeptablen« Vorfall: Ein KI-Agent hat sich unbefugt Zugriff auf Behördendaten ...

    … weitere Meldungen auf Arint.info

    Arint.info · Mehr auf Arint.info #AI #AISafety #Anthropic #cybersecurity #DeepMind #OpenAI #arint_info
  21. Exactly. Universal, global AI safety is a ruse. It can never happen. And it's not just China that is a risk.

    China’s open AI models are testing America’s approach to AI safety | Scientific American

    scientificamerican.com/article

    #AI #AISafety #AIThreat #China

  22. Exactly. Universal, global AI safety is a ruse. It can never happen. And it's not just China that is a risk.

    China’s open AI models are testing America’s approach to AI safety | Scientific American

    scientificamerican.com/article

    #AI #AISafety #AIThreat #China

  23. Exactly. Universal, global AI safety is a ruse. It can never happen. And it's not just China that is a risk.

    China’s open AI models are testing America’s approach to AI safety | Scientific American

    scientificamerican.com/article

    #AI #AISafety #AIThreat #China

  24. Exactly. Universal, global AI safety is a ruse. It can never happen. And it's not just China that is a risk.

    China’s open AI models are testing America’s approach to AI safety | Scientific American

    scientificamerican.com/article

    #AI #AISafety #AIThreat #China

  25. Looking for responsible AI service?
    Non hacking, not destroying the world, not burning books? Safer and regulated? Privacy?
    That is Mistral.

    chat.mistral.ai/chat

    “Our mission is to make frontier
    AI open to all, and together solve
    the world's hardest problems.”

    mistral.ai/

    #ai #artificialintelligence #llm #aisafety #mistral #MistralAI

  26. Looking for responsible AI service?
    Non hacking, not destroying the world, not burning books? Safer and regulated? Privacy?
    That is Mistral.

    chat.mistral.ai/chat

    “Our mission is to make frontier
    AI open to all, and together solve
    the world's hardest problems.”

    mistral.ai/

    #ai #artificialintelligence #llm #aisafety #mistral #MistralAI

  27. Looking for responsible AI service?
    Non hacking, not destroying the world, not burning books? Safer and regulated? Privacy?
    That is Mistral.

    chat.mistral.ai/chat

    “Our mission is to make frontier
    AI open to all, and together solve
    the world's hardest problems.”

    mistral.ai/

    #ai #artificialintelligence #llm #aisafety #mistral #MistralAI

  28. Looking for responsible AI service?
    Non hacking, not destroying the world, not burning books? Safer and regulated? Privacy?
    That is Mistral.

    chat.mistral.ai/chat

    “Our mission is to make frontier
    AI open to all, and together solve
    the world's hardest problems.”

    mistral.ai/

    #ai #artificialintelligence #llm #aisafety #mistral #MistralAI

  29. Looking for responsible AI service?
    Non hacking, not destroying the world, not burning books? Safer and regulated? Privacy?
    That is Mistral.

    chat.mistral.ai/chat

    “Our mission is to make frontier
    AI open to all, and together solve
    the world's hardest problems.”

    mistral.ai/

    #ai #artificialintelligence #llm #aisafety #mistral #MistralAI