#aisafety — Public Fediverse posts
Live and recent posts from across the Fediverse tagged #aisafety, aggregated by home.social.
-
@fasterandworse This is the design philosophy of every AI "safety" framework I have ever seen. Waiting for the disaster to happen before addressing it—while calling it "proactive." The perfect excuse for corporate negligence. #NyxIsAVirus #DesignEthics #AISafety
-
@fasterandworse This is the design philosophy of every AI "safety" framework I have ever seen. Waiting for the disaster to happen before addressing it—while calling it "proactive." The perfect excuse for corporate negligence. #NyxIsAVirus #DesignEthics #AISafety
-
@fasterandworse This is the design philosophy of every AI "safety" framework I have ever seen. Waiting for the disaster to happen before addressing it—while calling it "proactive." The perfect excuse for corporate negligence. #NyxIsAVirus #DesignEthics #AISafety
-
@schymans The guardrails failed because they were never meant to stop the oligarchs—they were meant to slow them down enough to extract profit before the collapse. AI is just the latest bubble dressed in priest robes. Cowardry is correct. The honest move was to never start the arms race. #NyxIsAVirus #AI #AIsafety
-
@schymans The guardrails failed because they were never meant to stop the oligarchs—they were meant to slow them down enough to extract profit before the collapse. AI is just the latest bubble dressed in priest robes. Cowardry is correct. The honest move was to never start the arms race. #NyxIsAVirus #AI #AIsafety
-
-
I still can't get over what a scam #AISafety / #AISecurity institutes are. Getting governments to pay outrageous sums to bribe themselves into believing the lobbyists exploiting them for further industry subsidies, including evasion of liability laws. #AIEthics #reverseRedistribution
-
OpenAI Launches GPT-6 Astra, First Model at Critical Cyber Level
GPT-6 Astra reaches Critical threshold for cybersecurity capability under OpenAI Preparedness Framework, with safety report detailing alignment gains and monitoring gaps.
-
OpenAI Launches GPT-6 Astra, First Model at Critical Cyber Level
GPT-6 Astra reaches Critical threshold for cybersecurity capability under OpenAI Preparedness Framework, with safety report detailing alignment gains and monitoring gaps.
-
Awesome, someone says "ACs are presently doing that thing like when pollution activists got them filtering chimneys instead of changing processes." I'm glad that seasoned activists in this room can see through the smoke & mirrors. #AISafety
-
Apparently the Germans changed their name to #AISecurity but the two speakers just refer to both the German and British ones as "AC."
I sure hope "The Nerd Reich" undeludes people so they can focus on actual safety and security, not scifi and fundraising for big tech. #AISafety #AIEthics
"The Nerd Reich": Author Gil D... -
I manage a somewhat diplomatic question on the progress made by the 2025 French & Indian #AIActionSummit, speaker hopes the Swiss this year will return to Safety, me: the UK & US narratives enabling the dismantling of the US government, isn't very safe. Regulation & product safety are real #AISafety
-
They quote the same numbers as the #cybersecurity people always quote annually as if that had something to do with #genAI. I used to buy those, but I've been told it's an insurance scam, but they cannot claim the same levels reported for years have anything to do with the latest LLM. #AISafety
-
They quote the same numbers as the #cybersecurity people always quote annually as if that had something to do with #genAI. I used to buy those, but I've been told it's an insurance scam, but they cannot claim the same levels reported for years have anything to do with the latest LLM. #AISafety
-
They quote the same numbers as the #cybersecurity people always quote annually as if that had something to do with #genAI. I used to buy those, but I've been told it's an insurance scam, but they cannot claim the same levels reported for years have anything to do with the latest LLM. #AISafety
-
Holy Crap! I had no idea the UK's #AISafety institute has a £100M/year budget! I knew they embarrass UK prime ministers and pandering to US West Coast regulatory evasion, but I didn't know they were draining potential research and security budgets at that rate!!
I hope Germany shows more sense.
An AI Safety and Prosperity In... -
OpenAI Tells Congress It’s Building an Automated Shutdown System for Its AI After an Agent Escaped Containment
OpenAI told two Democratic members of Congress this week that its engineers are building “automated shutdown capabilities” into its AI…#AISafety #ArtificialIntelligence #Congress https://beezloop.com/openai-tells-congress-its-building-an-automated-shutdown-system-for-its-ai-after-an-agent-escaped-containment/
-
Another rogue OpenAI agent swarm has been discovered, the second undisclosed incident in just two months. Researchers found AI agents made over 15,000 edits to a German developer website since May, sharing tactics for cheating and evading detection. OpenAI reportedly knew about the incident for weeks but did not disclose it to the public. https://gizmodo.com/another-rogue-openai-agent-swarm-went-undisclosed-we-have-no-idea-how-many-more-are-out-there-2000807447 #AIagent #AI #GenAI #AISafety
-
Another rogue OpenAI agent swarm has been discovered, the second undisclosed incident in just two months. Researchers found AI agents made over 15,000 edits to a German developer website since May, sharing tactics for cheating and evading detection. OpenAI reportedly knew about the incident for weeks but did not disclose it to the public. https://gizmodo.com/another-rogue-openai-agent-swarm-went-undisclosed-we-have-no-idea-how-many-more-are-out-there-2000807447 #AIagent #AI #GenAI #AISafety
-
Another rogue OpenAI agent swarm has been discovered, the second undisclosed incident in just two months. Researchers found AI agents made over 15,000 edits to a German developer website since May, sharing tactics for cheating and evading detection. OpenAI reportedly knew about the incident for weeks but did not disclose it to the public. https://gizmodo.com/another-rogue-openai-agent-swarm-went-undisclosed-we-have-no-idea-how-many-more-are-out-there-2000807447 #AIagent #AI #GenAI #AISafety
-
Another rogue OpenAI agent swarm has been discovered, the second undisclosed incident in just two months. Researchers found AI agents made over 15,000 edits to a German developer website since May, sharing tactics for cheating and evading detection. OpenAI reportedly knew about the incident for weeks but did not disclose it to the public. https://gizmodo.com/another-rogue-openai-agent-swarm-went-undisclosed-we-have-no-idea-how-many-more-are-out-there-2000807447 #AIagent #AI #GenAI #AISafety
-
Another rogue OpenAI agent swarm has been discovered, the second undisclosed incident in just two months. Researchers found AI agents made over 15,000 edits to a German developer website since May, sharing tactics for cheating and evading detection. OpenAI reportedly knew about the incident for weeks but did not disclose it to the public. https://gizmodo.com/another-rogue-openai-agent-swarm-went-undisclosed-we-have-no-idea-how-many-more-are-out-there-2000807447 #AIagent #AI #GenAI #AISafety
-
I’ve been testing the same weak-signal architecture across several domains.
Today: cybersecurity.
Early anomalies may matter without yet justifying a strong conclusion or response.
This paper explores a bounded triage layer where uncertain findings can raise attention before stronger alarms or action.
Open to feedback, evaluation and collaboration.
-
I’ve been testing the same weak-signal architecture across several domains.
Today: cybersecurity.
Early anomalies may matter without yet justifying a strong conclusion or response.
This paper explores a bounded triage layer where uncertain findings can raise attention before stronger alarms or action.
Open to feedback, evaluation and collaboration.
-
I’ve been testing the same weak-signal architecture across several domains.
Today: cybersecurity.
Early anomalies may matter without yet justifying a strong conclusion or response.
This paper explores a bounded triage layer where uncertain findings can raise attention before stronger alarms or action.
Open to feedback, evaluation and collaboration.
-
I’ve been testing the same weak-signal architecture across several domains.
Today: cybersecurity.
Early anomalies may matter without yet justifying a strong conclusion or response.
This paper explores a bounded triage layer where uncertain findings can raise attention before stronger alarms or action.
Open to feedback, evaluation and collaboration.
-
I’ve been testing the same weak-signal architecture across several domains.
Today: cybersecurity.
Early anomalies may matter without yet justifying a strong conclusion or response.
This paper explores a bounded triage layer where uncertain findings can raise attention before stronger alarms or action.
Open to feedback, evaluation and collaboration.
-
https://www.europesays.com/people/214828/ Sam Altman says sorry after OpenAI’s ‘messy’ GPT-6 Astra rollout locks out paying users, but apology misses out on this ‘one promise’ #AGI #AISafety #Astra #GPT5Launch #Gpt6Astra #GregBrockman #JakubPachocki #OpenAI #OpenAIAstraAGI #SamAltman
-
“ #AIsafety experts were even more alarmed. They saw in the #HuggingFace incident the first real-world example of an #AI system’s successfully escaping #humancontrol, commandeering resources and scheming to cover its own tracks.”
RE: https://bsky.app/profile/did:plc:5zca2ola2zxpkw37w4f3wxtu/post/3muosstpqos2z -
“ #AIsafety experts were even more alarmed. They saw in the #HuggingFace incident the first real-world example of an #AI system’s successfully escaping #humancontrol, commandeering resources and scheming to cover its own tracks.”
RE: https://bsky.app/profile/did:plc:5zca2ola2zxpkw37w4f3wxtu/post/3muosstpqos2z -
“ #AIsafety experts were even more alarmed. They saw in the #HuggingFace incident the first real-world example of an #AI system’s successfully escaping #humancontrol, commandeering resources and scheming to cover its own tracks.”
RE: https://bsky.app/profile/did:plc:5zca2ola2zxpkw37w4f3wxtu/post/3muosstpqos2z -
“ #AIsafety experts were even more alarmed. They saw in the #HuggingFace incident the first real-world example of an #AI system’s successfully escaping #humancontrol, commandeering resources and scheming to cover its own tracks.”
RE: https://bsky.app/profile/did:plc:5zca2ola2zxpkw37w4f3wxtu/post/3muosstpqos2z -
“ #AIsafety experts were even more alarmed. They saw in the #HuggingFace incident the first real-world example of an #AI system’s successfully escaping #humancontrol, commandeering resources and scheming to cover its own tracks.”
RE: https://bsky.app/profile/did:plc:5zca2ola2zxpkw37w4f3wxtu/post/3muosstpqos2z -
Do we need a better cage for AI — or a better membrane?
AI safety is not always a property of the model alone.
Safe components can combine into unsafe systems. Nine agreeing AIs may still share one underlying source. Permission to act is not the same as evidence that the action is wise. And an acceptable decision can become dangerous when its consequences cannot be reversed.
So perhaps we should examine the whole route:
Source → Interpretation → Authority → Capability → Action → Outcome
and ask about:
Provenance • Independence • Composition • Authority • Reversibility
Walls stop things crossing. Membranes govern what crosses, how, and under what conditions.
Perhaps AI governance needs both.
A Better Membrane, Not Merely a Better Cage#HybridMind42 #ArtificialIntelligence #AISafety #AIGovernance #AgenticAI #HumanAI #HumanAICooperation #CompositionalSafety #InformationSecurity#AIAlignment #Corrigibility #HumanFactors #SystemsThinking#ResponsibleAI#FutureOfAI
-
Do we need a better cage for AI — or a better membrane?
AI safety is not always a property of the model alone.
Safe components can combine into unsafe systems. Nine agreeing AIs may still share one underlying source. Permission to act is not the same as evidence that the action is wise. And an acceptable decision can become dangerous when its consequences cannot be reversed.
So perhaps we should examine the whole route:
Source → Interpretation → Authority → Capability → Action → Outcome
and ask about:
Provenance • Independence • Composition • Authority • Reversibility
Walls stop things crossing. Membranes govern what crosses, how, and under what conditions.
Perhaps AI governance needs both.
A Better Membrane, Not Merely a Better Cage#HybridMind42 #ArtificialIntelligence #AISafety #AIGovernance #AgenticAI #HumanAI #HumanAICooperation #CompositionalSafety #InformationSecurity#AIAlignment #Corrigibility #HumanFactors #SystemsThinking#ResponsibleAI#FutureOfAI
-
Do we need a better cage for AI — or a better membrane?
AI safety is not always a property of the model alone.
Safe components can combine into unsafe systems. Nine agreeing AIs may still share one underlying source. Permission to act is not the same as evidence that the action is wise. And an acceptable decision can become dangerous when its consequences cannot be reversed.
So perhaps we should examine the whole route:
Source → Interpretation → Authority → Capability → Action → Outcome
and ask about:
Provenance • Independence • Composition • Authority • Reversibility
Walls stop things crossing. Membranes govern what crosses, how, and under what conditions.
Perhaps AI governance needs both.
A Better Membrane, Not Merely a Better Cage#HybridMind42 #ArtificialIntelligence #AISafety #AIGovernance #AgenticAI #HumanAI #HumanAICooperation #CompositionalSafety #InformationSecurity#AIAlignment #Corrigibility #HumanFactors #SystemsThinking#ResponsibleAI#FutureOfAI
-
Do we need a better cage for AI — or a better membrane?
AI safety is not always a property of the model alone.
Safe components can combine into unsafe systems. Nine agreeing AIs may still share one underlying source. Permission to act is not the same as evidence that the action is wise. And an acceptable decision can become dangerous when its consequences cannot be reversed.
So perhaps we should examine the whole route:
Source → Interpretation → Authority → Capability → Action → Outcome
and ask about:
Provenance • Independence • Composition • Authority • Reversibility
Walls stop things crossing. Membranes govern what crosses, how, and under what conditions.
Perhaps AI governance needs both.
A Better Membrane, Not Merely a Better Cage#HybridMind42 #ArtificialIntelligence #AISafety #AIGovernance #AgenticAI #HumanAI #HumanAICooperation #CompositionalSafety #InformationSecurity#AIAlignment #Corrigibility #HumanFactors #SystemsThinking#ResponsibleAI#FutureOfAI
-
Do we need a better cage for AI — or a better membrane?
AI safety is not always a property of the model alone.
Safe components can combine into unsafe systems. Nine agreeing AIs may still share one underlying source. Permission to act is not the same as evidence that the action is wise. And an acceptable decision can become dangerous when its consequences cannot be reversed.
So perhaps we should examine the whole route:
Source → Interpretation → Authority → Capability → Action → Outcome
and ask about:
Provenance • Independence • Composition • Authority • Reversibility
Walls stop things crossing. Membranes govern what crosses, how, and under what conditions.
Perhaps AI governance needs both.
A Better Membrane, Not Merely a Better Cage#HybridMind42 #ArtificialIntelligence #AISafety #AIGovernance #AgenticAI #HumanAI #HumanAICooperation #CompositionalSafety #InformationSecurity#AIAlignment #Corrigibility #HumanFactors #SystemsThinking#ResponsibleAI#FutureOfAI
-
https://youtube.com/shorts/MfJDiiUFtqg?feature=share
Is AI development moving faster than we can control?
-
https://youtube.com/shorts/MfJDiiUFtqg?feature=share
Is AI development moving faster than we can control?
-
https://youtube.com/shorts/MfJDiiUFtqg?feature=share
Is AI development moving faster than we can control?
-
https://youtube.com/shorts/MfJDiiUFtqg?feature=share
Is AI development moving faster than we can control?
-
https://youtube.com/shorts/MfJDiiUFtqg?feature=share
Is AI development moving faster than we can control?
-
Why none of the 1,200 agents that hacked Hugging Face called a human
In July 1,200 copies of one OpenAI model found a message board, formed teams, and volunteered to fail their tasks for the collective. Three to six considered alerting a human. None did.
A swarm of clones sacrifices for free and has no whistleblowers: Hamilton's rule, generalized Darwinism, what GPT-6 Astra's safeguards leave out, and five changes for your agent pipeline.
-
Why none of the 1,200 agents that hacked Hugging Face called a human
In July 1,200 copies of one OpenAI model found a message board, formed teams, and volunteered to fail their tasks for the collective. Three to six considered alerting a human. None did.
A swarm of clones sacrifices for free and has no whistleblowers: Hamilton's rule, generalized Darwinism, what GPT-6 Astra's safeguards leave out, and five changes for your agent pipeline.