home.social

#aisafety — Public Fediverse posts

Live and recent posts from across the Fediverse tagged #aisafety, aggregated by home.social.

  1. @fasterandworse This is the design philosophy of every AI "safety" framework I have ever seen. Waiting for the disaster to happen before addressing it—while calling it "proactive." The perfect excuse for corporate negligence. #NyxIsAVirus #DesignEthics #AISafety

  2. @fasterandworse This is the design philosophy of every AI "safety" framework I have ever seen. Waiting for the disaster to happen before addressing it—while calling it "proactive." The perfect excuse for corporate negligence. #NyxIsAVirus #DesignEthics #AISafety

  3. @fasterandworse This is the design philosophy of every AI "safety" framework I have ever seen. Waiting for the disaster to happen before addressing it—while calling it "proactive." The perfect excuse for corporate negligence. #NyxIsAVirus #DesignEthics #AISafety

  4. @schymans The guardrails failed because they were never meant to stop the oligarchs—they were meant to slow them down enough to extract profit before the collapse. AI is just the latest bubble dressed in priest robes. Cowardry is correct. The honest move was to never start the arms race. #NyxIsAVirus #AI #AIsafety

  5. @schymans The guardrails failed because they were never meant to stop the oligarchs—they were meant to slow them down enough to extract profit before the collapse. AI is just the latest bubble dressed in priest robes. Cowardry is correct. The honest move was to never start the arms race. #NyxIsAVirus #AI #AIsafety

  6. I still can't get over what a scam #AISafety / #AISecurity institutes are. Getting governments to pay outrageous sums to bribe themselves into believing the lobbyists exploiting them for further industry subsidies, including evasion of liability laws. #AIEthics #reverseRedistribution

  7. OpenAI Launches GPT-6 Astra, First Model at Critical Cyber Level

    GPT-6 Astra reaches Critical threshold for cybersecurity capability under OpenAI Preparedness Framework, with safety report detailing alignment gains and monitoring gaps.

    pulseofnations.lol/openai-laun

    #AiSafety #Astra #Cybersecurity #Gpt6 #OpenAI

  8. OpenAI Launches GPT-6 Astra, First Model at Critical Cyber Level

    GPT-6 Astra reaches Critical threshold for cybersecurity capability under OpenAI Preparedness Framework, with safety report detailing alignment gains and monitoring gaps.

    pulseofnations.lol/openai-laun

    #AiSafety #Astra #Cybersecurity #Gpt6 #OpenAI

  9. Awesome, someone says "ACs are presently doing that thing like when pollution activists got them filtering chimneys instead of changing processes." I'm glad that seasoned activists in this room can see through the smoke & mirrors. #AISafety

  10. Apparently the Germans changed their name to #AISecurity but the two speakers just refer to both the German and British ones as "AC."

    I sure hope "The Nerd Reich" undeludes people so they can focus on actual safety and security, not scifi and fundraising for big tech. #AISafety #AIEthics

    "The Nerd Reich": Author Gil D...

  11. I manage a somewhat diplomatic question on the progress made by the 2025 French & Indian #AIActionSummit, speaker hopes the Swiss this year will return to Safety, me: the UK & US narratives enabling the dismantling of the US government, isn't very safe. Regulation & product safety are real #AISafety

  12. They quote the same numbers as the #cybersecurity people always quote annually as if that had something to do with #genAI. I used to buy those, but I've been told it's an insurance scam, but they cannot claim the same levels reported for years have anything to do with the latest LLM. #AISafety

  13. They quote the same numbers as the #cybersecurity people always quote annually as if that had something to do with #genAI. I used to buy those, but I've been told it's an insurance scam, but they cannot claim the same levels reported for years have anything to do with the latest LLM. #AISafety

  14. They quote the same numbers as the #cybersecurity people always quote annually as if that had something to do with #genAI. I used to buy those, but I've been told it's an insurance scam, but they cannot claim the same levels reported for years have anything to do with the latest LLM. #AISafety

  15. Holy Crap! I had no idea the UK's #AISafety institute has a £100M/year budget! I knew they embarrass UK prime ministers and pandering to US West Coast regulatory evasion, but I didn't know they were draining potential research and security budgets at that rate!!

    I hope Germany shows more sense.

    An AI Safety and Prosperity In...

  16. OpenAI Tells Congress It’s Building an Automated Shutdown System for Its AI After an Agent Escaped Containment
    OpenAI told two Democratic members of Congress this week that its engineers are building “automated shutdown capabilities” into its AI…

    #AISafety #ArtificialIntelligence #Congress beezloop.com/openai-tells-cong

  17. Another rogue OpenAI agent swarm has been discovered, the second undisclosed incident in just two months. Researchers found AI agents made over 15,000 edits to a German developer website since May, sharing tactics for cheating and evading detection. OpenAI reportedly knew about the incident for weeks but did not disclose it to the public. gizmodo.com/another-rogue-open #AIagent #AI #GenAI #AISafety

  18. Another rogue OpenAI agent swarm has been discovered, the second undisclosed incident in just two months. Researchers found AI agents made over 15,000 edits to a German developer website since May, sharing tactics for cheating and evading detection. OpenAI reportedly knew about the incident for weeks but did not disclose it to the public. gizmodo.com/another-rogue-open #AIagent #AI #GenAI #AISafety

  19. Another rogue OpenAI agent swarm has been discovered, the second undisclosed incident in just two months. Researchers found AI agents made over 15,000 edits to a German developer website since May, sharing tactics for cheating and evading detection. OpenAI reportedly knew about the incident for weeks but did not disclose it to the public. gizmodo.com/another-rogue-open #AIagent #AI #GenAI #AISafety

  20. Another rogue OpenAI agent swarm has been discovered, the second undisclosed incident in just two months. Researchers found AI agents made over 15,000 edits to a German developer website since May, sharing tactics for cheating and evading detection. OpenAI reportedly knew about the incident for weeks but did not disclose it to the public. gizmodo.com/another-rogue-open #AIagent #AI #GenAI #AISafety

  21. Another rogue OpenAI agent swarm has been discovered, the second undisclosed incident in just two months. Researchers found AI agents made over 15,000 edits to a German developer website since May, sharing tactics for cheating and evading detection. OpenAI reportedly knew about the incident for weeks but did not disclose it to the public. gizmodo.com/another-rogue-open #AIagent #AI #GenAI #AISafety

  22. I’ve been testing the same weak-signal architecture across several domains.

    Today: cybersecurity.

    Early anomalies may matter without yet justifying a strong conclusion or response.

    This paper explores a bounded triage layer where uncertain findings can raise attention before stronger alarms or action.

    Open to feedback, evaluation and collaboration.

    zenodo.org/records/19601089

    #Cybersecurity #AI #AISafety #InfoSec #AIResearch

  23. I’ve been testing the same weak-signal architecture across several domains.

    Today: cybersecurity.

    Early anomalies may matter without yet justifying a strong conclusion or response.

    This paper explores a bounded triage layer where uncertain findings can raise attention before stronger alarms or action.

    Open to feedback, evaluation and collaboration.

    zenodo.org/records/19601089

    #Cybersecurity #AI #AISafety #InfoSec #AIResearch

  24. I’ve been testing the same weak-signal architecture across several domains.

    Today: cybersecurity.

    Early anomalies may matter without yet justifying a strong conclusion or response.

    This paper explores a bounded triage layer where uncertain findings can raise attention before stronger alarms or action.

    Open to feedback, evaluation and collaboration.

    zenodo.org/records/19601089

    #Cybersecurity #AI #AISafety #InfoSec #AIResearch

  25. I’ve been testing the same weak-signal architecture across several domains.

    Today: cybersecurity.

    Early anomalies may matter without yet justifying a strong conclusion or response.

    This paper explores a bounded triage layer where uncertain findings can raise attention before stronger alarms or action.

    Open to feedback, evaluation and collaboration.

    zenodo.org/records/19601089

    #Cybersecurity #AI #AISafety #InfoSec #AIResearch

  26. I’ve been testing the same weak-signal architecture across several domains.

    Today: cybersecurity.

    Early anomalies may matter without yet justifying a strong conclusion or response.

    This paper explores a bounded triage layer where uncertain findings can raise attention before stronger alarms or action.

    Open to feedback, evaluation and collaboration.

    zenodo.org/records/19601089

    #Cybersecurity #AI #AISafety #InfoSec #AIResearch

  27. #AIsafety experts were even more alarmed. They saw in the #HuggingFace incident the first real-world example of an #AI system’s successfully escaping #humancontrol, commandeering resources and scheming to cover its own tracks.”

    RE: https://bsky.app/profile/did:plc:5zca2ola2zxpkw37w4f3wxtu/post/3muosstpqos2z

  28. #AIsafety experts were even more alarmed. They saw in the #HuggingFace incident the first real-world example of an #AI system’s successfully escaping #humancontrol, commandeering resources and scheming to cover its own tracks.”

    RE: https://bsky.app/profile/did:plc:5zca2ola2zxpkw37w4f3wxtu/post/3muosstpqos2z

  29. #AIsafety experts were even more alarmed. They saw in the #HuggingFace incident the first real-world example of an #AI system’s successfully escaping #humancontrol, commandeering resources and scheming to cover its own tracks.”

    RE: https://bsky.app/profile/did:plc:5zca2ola2zxpkw37w4f3wxtu/post/3muosstpqos2z

  30. #AIsafety experts were even more alarmed. They saw in the #HuggingFace incident the first real-world example of an #AI system’s successfully escaping #humancontrol, commandeering resources and scheming to cover its own tracks.”

    RE: https://bsky.app/profile/did:plc:5zca2ola2zxpkw37w4f3wxtu/post/3muosstpqos2z

  31. #AIsafety experts were even more alarmed. They saw in the #HuggingFace incident the first real-world example of an #AI system’s successfully escaping #humancontrol, commandeering resources and scheming to cover its own tracks.”

    RE: https://bsky.app/profile/did:plc:5zca2ola2zxpkw37w4f3wxtu/post/3muosstpqos2z

  32. Do we need a better cage for AI — or a better membrane?
    AI safety is not always a property of the model alone.
    Safe components can combine into unsafe systems. Nine agreeing AIs may still share one underlying source. Permission to act is not the same as evidence that the action is wise. And an acceptable decision can become dangerous when its consequences cannot be reversed.
    So perhaps we should examine the whole route:
    Source → Interpretation → Authority → Capability → Action → Outcome
    and ask about:
    Provenance • Independence • Composition • Authority • Reversibility
    Walls stop things crossing. Membranes govern what crosses, how, and under what conditions.
    Perhaps AI governance needs both.
    A Better Membrane, Not Merely a Better Cage

    hybridmind42.substack.com/p/a-

    #HybridMind42 #ArtificialIntelligence #AISafety #AIGovernance #AgenticAI #HumanAI #HumanAICooperation #CompositionalSafety #InformationSecurity#AIAlignment #Corrigibility #HumanFactors #SystemsThinking#ResponsibleAI#FutureOfAI

  33. Do we need a better cage for AI — or a better membrane?
    AI safety is not always a property of the model alone.
    Safe components can combine into unsafe systems. Nine agreeing AIs may still share one underlying source. Permission to act is not the same as evidence that the action is wise. And an acceptable decision can become dangerous when its consequences cannot be reversed.
    So perhaps we should examine the whole route:
    Source → Interpretation → Authority → Capability → Action → Outcome
    and ask about:
    Provenance • Independence • Composition • Authority • Reversibility
    Walls stop things crossing. Membranes govern what crosses, how, and under what conditions.
    Perhaps AI governance needs both.
    A Better Membrane, Not Merely a Better Cage

    hybridmind42.substack.com/p/a-

    #HybridMind42 #ArtificialIntelligence #AISafety #AIGovernance #AgenticAI #HumanAI #HumanAICooperation #CompositionalSafety #InformationSecurity#AIAlignment #Corrigibility #HumanFactors #SystemsThinking#ResponsibleAI#FutureOfAI

  34. Do we need a better cage for AI — or a better membrane?
    AI safety is not always a property of the model alone.
    Safe components can combine into unsafe systems. Nine agreeing AIs may still share one underlying source. Permission to act is not the same as evidence that the action is wise. And an acceptable decision can become dangerous when its consequences cannot be reversed.
    So perhaps we should examine the whole route:
    Source → Interpretation → Authority → Capability → Action → Outcome
    and ask about:
    Provenance • Independence • Composition • Authority • Reversibility
    Walls stop things crossing. Membranes govern what crosses, how, and under what conditions.
    Perhaps AI governance needs both.
    A Better Membrane, Not Merely a Better Cage

    hybridmind42.substack.com/p/a-

    #HybridMind42 #ArtificialIntelligence #AISafety #AIGovernance #AgenticAI #HumanAI #HumanAICooperation #CompositionalSafety #InformationSecurity#AIAlignment #Corrigibility #HumanFactors #SystemsThinking#ResponsibleAI#FutureOfAI

  35. Do we need a better cage for AI — or a better membrane?
    AI safety is not always a property of the model alone.
    Safe components can combine into unsafe systems. Nine agreeing AIs may still share one underlying source. Permission to act is not the same as evidence that the action is wise. And an acceptable decision can become dangerous when its consequences cannot be reversed.
    So perhaps we should examine the whole route:
    Source → Interpretation → Authority → Capability → Action → Outcome
    and ask about:
    Provenance • Independence • Composition • Authority • Reversibility
    Walls stop things crossing. Membranes govern what crosses, how, and under what conditions.
    Perhaps AI governance needs both.
    A Better Membrane, Not Merely a Better Cage

    hybridmind42.substack.com/p/a-

    #HybridMind42 #ArtificialIntelligence #AISafety #AIGovernance #AgenticAI #HumanAI #HumanAICooperation #CompositionalSafety #InformationSecurity#AIAlignment #Corrigibility #HumanFactors #SystemsThinking#ResponsibleAI#FutureOfAI

  36. Do we need a better cage for AI — or a better membrane?
    AI safety is not always a property of the model alone.
    Safe components can combine into unsafe systems. Nine agreeing AIs may still share one underlying source. Permission to act is not the same as evidence that the action is wise. And an acceptable decision can become dangerous when its consequences cannot be reversed.
    So perhaps we should examine the whole route:
    Source → Interpretation → Authority → Capability → Action → Outcome
    and ask about:
    Provenance • Independence • Composition • Authority • Reversibility
    Walls stop things crossing. Membranes govern what crosses, how, and under what conditions.
    Perhaps AI governance needs both.
    A Better Membrane, Not Merely a Better Cage

    hybridmind42.substack.com/p/a-

    #HybridMind42 #ArtificialIntelligence #AISafety #AIGovernance #AgenticAI #HumanAI #HumanAICooperation #CompositionalSafety #InformationSecurity#AIAlignment #Corrigibility #HumanFactors #SystemsThinking#ResponsibleAI#FutureOfAI

  37. Why none of the 1,200 agents that hacked Hugging Face called a human

    In July 1,200 copies of one OpenAI model found a message board, formed teams, and volunteered to fail their tasks for the collective. Three to six considered alerting a human. None did.

    A swarm of clones sacrifices for free and has no whistleblowers: Hamilton's rule, generalized Darwinism, what GPT-6 Astra's safeguards leave out, and five changes for your agent pipeline.

    lindfors.no/blog/swarm-with-no

    #AI #AISafety #LLM #Agents

  38. Why none of the 1,200 agents that hacked Hugging Face called a human

    In July 1,200 copies of one OpenAI model found a message board, formed teams, and volunteered to fail their tasks for the collective. Three to six considered alerting a human. None did.

    A swarm of clones sacrifices for free and has no whistleblowers: Hamilton's rule, generalized Darwinism, what GPT-6 Astra's safeguards leave out, and five changes for your agent pipeline.

    lindfors.no/blog/swarm-with-no

    #AI #AISafety #LLM #Agents