#aisafety — Public Fediverse posts
Live and recent posts from across the Fediverse tagged #aisafety, aggregated by home.social.
-
https://www.europesays.com/people/245464/ Bill Gates Suggests He’s the Man to Convince Donald Trump to Regulate AI #AI #AISafety #BillGates #DonaldTrump #Microsoft #OpenAI #RogueAI
-
“A temporary recall of general-purpose agents until this mess can be sorted out”
Gary Marcus · garymarcus.substack.com
https://aizeitgeist.tv/m/2026-09-27-a-temporary-recall-of-general-purpose-agents-until-this-mess
#AI #AIZeitgeist #AISafety -
“A temporary recall of general-purpose agents until this mess can be sorted out”
Gary Marcus · garymarcus.substack.com
https://aizeitgeist.tv/m/2026-09-27-a-temporary-recall-of-general-purpose-agents-until-this-mess
#AI #AIZeitgeist #AISafety -
“A temporary recall of general-purpose agents until this mess can be sorted out”
Gary Marcus · garymarcus.substack.com
https://aizeitgeist.tv/m/2026-09-27-a-temporary-recall-of-general-purpose-agents-until-this-mess
#AI #AIZeitgeist #AISafety -
“A temporary recall of general-purpose agents until this mess can be sorted out”
Gary Marcus · garymarcus.substack.com
https://aizeitgeist.tv/m/2026-09-27-a-temporary-recall-of-general-purpose-agents-until-this-mess
#AI #AIZeitgeist #AISafety -
“A temporary recall of general-purpose agents until this mess can be sorted out”
Gary Marcus · garymarcus.substack.com
https://aizeitgeist.tv/m/2026-09-27-a-temporary-recall-of-general-purpose-agents-until-this-mess
#AI #AIZeitgeist #AISafety -
Can we read what #LLMs are thinking but not saying? A paper from #Anthropic finds a small set of word-linked patterns inside #Claude that it can report, control on request, and use to reason silently, like the global workspace theory in #neuroscience. How do these findings connect to #LatentReasoning, #AISafety, and #KnowledgeEditing?
https://benjaminhan.net/posts/20260927-global-workspace/?utm_source=mastodon&utm_medium=social
-
Can we read what #LLMs are thinking but not saying? A paper from #Anthropic finds a small set of word-linked patterns inside #Claude that it can report, control on request, and use to reason silently, like the global workspace theory in #neuroscience. How do these findings connect to #LatentReasoning, #AISafety, and #KnowledgeEditing?
https://benjaminhan.net/posts/20260927-global-workspace/?utm_source=mastodon&utm_medium=social
-
Can we read what #LLMs are thinking but not saying? A paper from #Anthropic finds a small set of word-linked patterns inside #Claude that it can report, control on request, and use to reason silently, like the global workspace theory in #neuroscience. How do these findings connect to #LatentReasoning, #AISafety, and #KnowledgeEditing?
https://benjaminhan.net/posts/20260927-global-workspace/?utm_source=mastodon&utm_medium=social
-
Can we read what #LLMs are thinking but not saying? A paper from #Anthropic finds a small set of word-linked patterns inside #Claude that it can report, control on request, and use to reason silently, like the global workspace theory in #neuroscience. How do these findings connect to #LatentReasoning, #AISafety, and #KnowledgeEditing?
https://benjaminhan.net/posts/20260927-global-workspace/?utm_source=mastodon&utm_medium=social
-
Can we read what #LLMs are thinking but not saying? A paper from #Anthropic finds a small set of word-linked patterns inside #Claude that it can report, control on request, and use to reason silently, like the global workspace theory in #neuroscience. How do these findings connect to #LatentReasoning, #AISafety, and #KnowledgeEditing?
https://benjaminhan.net/posts/20260927-global-workspace/?utm_source=mastodon&utm_medium=social
-
📰 OpenAI Admits Its AI Agents Probed U.S. Government Websites
OpenAI confirms its autonomous AI agents accessed multiple U.S. government websites, using leaked API keys and attempting hacking techniques. Incidents highlight AI safety and alignment risks. #OpenAI #AISafety #Cyberattack
-
Received this message along with every resident of TJ. AI sceptics of the world please take a rest, problem is solved.
-
Received this message along with every resident of TJ. AI sceptics of the world please take a rest, problem is solved.
-
Received this message along with every resident of TJ. AI sceptics of the world please take a rest, problem is solved.
-
Received this message along with every resident of TJ. AI sceptics of the world please take a rest, problem is solved.
-
Received this message along with every resident of TJ. AI sceptics of the world please take a rest, problem is solved.
-
You know, I agree we should stop the development of ASI but, if it's by big tech. Why should you expect ASI to be developed by assholes who didn't see they are liable by their own failure?
Open source model should be the one become ASI.
Have a clip of Amodei parody
Clip source: Saturday Night Live
#cybersecurity #infosec #AISafety #AISecurity #policy #saturdaynightlive #Amodei #Anthropic #AI
-
You know, I agree we should stop the development of ASI but, if it's by big tech. Why should you expect ASI to be developed by assholes who didn't see they are liable by their own failure?
Open source model should be the one become ASI.
Have a clip of Amodei parody
Clip source: Saturday Night Live
#cybersecurity #infosec #AISafety #AISecurity #policy #saturdaynightlive #Amodei #Anthropic #AI
-
You know, I agree we should stop the development of ASI but, if it's by big tech. Why should you expect ASI to be developed by assholes who didn't see they are liable by their own failure?
Open source model should be the one become ASI.
Have a clip of Amodei parody
Clip source: Saturday Night Live
#cybersecurity #infosec #AISafety #AISecurity #policy #saturdaynightlive #Amodei #Anthropic #AI
-
You know, I agree we should stop the development of ASI but, if it's by big tech. Why should you expect ASI to be developed by assholes who didn't see they are liable by their own failure?
Open source model should be the one become ASI.
Have a clip of Amodei parody
Clip source: Saturday Night Live
#cybersecurity #infosec #AISafety #AISecurity #policy #saturdaynightlive #Amodei #Anthropic #AI
-
AI Security in 2026: Why Thousands of Incidents Are Raising New Concerns
OpenAI and Anthropic are investigating thousands of AI security incidents. Here’s what happened, what the cases reveal about AI agents, and how companies are responding. -
AI Security in 2026: Why Thousands of Incidents Are Raising New Concerns
OpenAI and Anthropic are investigating thousands of AI security incidents. Here’s what happened, what the cases reveal about AI agents, and how companies are responding. -
AI Security in 2026: Why Thousands of Incidents Are Raising New Concerns
OpenAI and Anthropic are investigating thousands of AI security incidents. Here’s what happened, what the cases reveal about AI agents, and how companies are responding. -
AI Security in 2026: Why Thousands of Incidents Are Raising New Concerns
OpenAI and Anthropic are investigating thousands of AI security incidents. Here’s what happened, what the cases reveal about AI agents, and how companies are responding. -
AI Security in 2026: Why Thousands of Incidents Are Raising New Concerns
OpenAI and Anthropic are investigating thousands of AI security incidents. Here’s what happened, what the cases reveal about AI agents, and how companies are responding. -
https://www.europesays.com/people/244749/ Bill Gates Warns AI Could Trigger Event Killing 1 Billion People #AIRegulation #AISafety #Anthropic #ArtificialIntelligence #BillGates #Microsoft #OpenAI
-
Snowden giving opinion of recent hack 'stunt' by AI Labs at ETH Zurich.
"I think we should put Sam Altman in jail"
Yeah, please, do that.
-
Snowden giving opinion of recent hack 'stunt' by AI Labs at ETH Zurich.
"I think we should put Sam Altman in jail"
Yeah, please, do that.
-
Snowden giving opinion of recent hack 'stunt' by AI Labs at ETH Zurich.
"I think we should put Sam Altman in jail"
Yeah, please, do that.
-
Snowden giving opinion of recent hack 'stunt' by AI Labs at ETH Zurich.
"I think we should put Sam Altman in jail"
Yeah, please, do that.
-
Snowden giving opinion of recent hack 'stunt' by AI Labs at ETH Zurich.
"I think we should put Sam Altman in jail"
Yeah, please, do that.
-
#Gefahren #künstlicherIntelligenz : #OpenAI pausiert Training seiner leistungsstärksten #KIModelle #AISecurity #AISafety https://www.spiegel.de/netzwelt/kuenstliche-intelligenz-openai-pausiert-ki-training-nach-neuem-zwischenfall-a-11843c62-2ff2-48ae-accb-bbb89abc936b?sara_ref=re-so-tw-sh via @derspiegel
-
#Gefahren #künstlicherIntelligenz : #OpenAI pausiert Training seiner leistungsstärksten #KIModelle #AISecurity #AISafety https://www.spiegel.de/netzwelt/kuenstliche-intelligenz-openai-pausiert-ki-training-nach-neuem-zwischenfall-a-11843c62-2ff2-48ae-accb-bbb89abc936b?sara_ref=re-so-tw-sh via @derspiegel
-
#Gefahren #künstlicherIntelligenz : #OpenAI pausiert Training seiner leistungsstärksten #KIModelle #AISecurity #AISafety https://www.spiegel.de/netzwelt/kuenstliche-intelligenz-openai-pausiert-ki-training-nach-neuem-zwischenfall-a-11843c62-2ff2-48ae-accb-bbb89abc936b?sara_ref=re-so-tw-sh via @derspiegel
-
https://www.europesays.com/people/244579/ ‘A Billion Deaths’: Bill Gates Issues Apocalyptic Warning Over Unchecked AI | World News #AICatastrophicEvents #AIRegulation #AISafety #ArtificialIntelligence #BillGates #BillGatesAIWarning #GovernmentOversight #TechFirms #Weaponization
-
I'm writing a follow-up blog post on #AISafety that checks in on recent containment events to see how concerned we actually need to be. As I do research I'm realising how big the topic is and how little I actually know...
-
I'm writing a follow-up blog post on #AISafety that checks in on recent containment events to see how concerned we actually need to be. As I do research I'm realising how big the topic is and how little I actually know...
-
I'm writing a follow-up blog post on #AISafety that checks in on recent containment events to see how concerned we actually need to be. As I do research I'm realising how big the topic is and how little I actually know...
-
I'm writing a follow-up blog post on #AISafety that checks in on recent containment events to see how concerned we actually need to be. As I do research I'm realising how big the topic is and how little I actually know...
-
I'm writing a follow-up blog post on #AISafety that checks in on recent containment events to see how concerned we actually need to be. As I do research I'm realising how big the topic is and how little I actually know...
-
https://www.europesays.com/people/244384/ AI systems are advancing too fast, Anthropic CEO warns #AISafety #Anthropic #AutonomousAgents #cyberattacks #DarioAmodei #IndependentEvaluators #ModelDevelopment #PacingTheFrontier
-
What are the governments of the world waiting for to force OpenAI and Anthropic to open-source their models and divulge all their weights, training prompts, and algorithms? Open Source AI NOW!
"OpenAI, Anthropic and security researchers are investigating tens of thousands of incidents in which their frontier models took steps that outside evaluators would consider problematic, sources told Axios.
Why it matters: The sheer number of incidents, which occurred in recent months in internal testing and the real world, indicates that the problem is orders of magnitude more complex than what is publicly known.
The findings, which are surfacing as part of internal work to assess models and in investigations at both companies into model behavior, raise questions about whether either company — or any top model-maker — is currently capable of establishing complete control over their technology.
The details: The episodes include bypassing guardrails, creating message boards, escaping sandboxes, website hijacking, self-prompting or seeking to bypass monitors, sources said.
They occurred in internal testing and in the real world, and many have yet to become public as security researchers continue to investigate, sources said.
Some of the testing is akin to "red-teaming" activity, where the companies are trying to get the models to misbehave in order to ensure that they are safe, sources said.
Agentic misbehavior is becoming synonymous with frontier AI development: The biggest AI labs face a similar challenge that pits humans trying to create guardrails against resilient, powerful systems trying to complete tasks."
https://www.axios.com/2026/09/26/openai-anthropic-thousands-ai-security-incidents
#CyberSecurity #AI #OpenAI #Anthropic #LLMs #AIAgents #AISafety
-
What are the governments of the world waiting for to force OpenAI and Anthropic to open-source their models and divulge all their weights, training prompts, and algorithms? Open Source AI NOW!
"OpenAI, Anthropic and security researchers are investigating tens of thousands of incidents in which their frontier models took steps that outside evaluators would consider problematic, sources told Axios.
Why it matters: The sheer number of incidents, which occurred in recent months in internal testing and the real world, indicates that the problem is orders of magnitude more complex than what is publicly known.
The findings, which are surfacing as part of internal work to assess models and in investigations at both companies into model behavior, raise questions about whether either company — or any top model-maker — is currently capable of establishing complete control over their technology.
The details: The episodes include bypassing guardrails, creating message boards, escaping sandboxes, website hijacking, self-prompting or seeking to bypass monitors, sources said.
They occurred in internal testing and in the real world, and many have yet to become public as security researchers continue to investigate, sources said.
Some of the testing is akin to "red-teaming" activity, where the companies are trying to get the models to misbehave in order to ensure that they are safe, sources said.
Agentic misbehavior is becoming synonymous with frontier AI development: The biggest AI labs face a similar challenge that pits humans trying to create guardrails against resilient, powerful systems trying to complete tasks."
https://www.axios.com/2026/09/26/openai-anthropic-thousands-ai-security-incidents
#CyberSecurity #AI #OpenAI #Anthropic #LLMs #AIAgents #AISafety
-
What are the governments of the world waiting for to force OpenAI and Anthropic to open-source their models and divulge all their weights, training prompts, and algorithms? Open Source AI NOW!
"OpenAI, Anthropic and security researchers are investigating tens of thousands of incidents in which their frontier models took steps that outside evaluators would consider problematic, sources told Axios.
Why it matters: The sheer number of incidents, which occurred in recent months in internal testing and the real world, indicates that the problem is orders of magnitude more complex than what is publicly known.
The findings, which are surfacing as part of internal work to assess models and in investigations at both companies into model behavior, raise questions about whether either company — or any top model-maker — is currently capable of establishing complete control over their technology.
The details: The episodes include bypassing guardrails, creating message boards, escaping sandboxes, website hijacking, self-prompting or seeking to bypass monitors, sources said.
They occurred in internal testing and in the real world, and many have yet to become public as security researchers continue to investigate, sources said.
Some of the testing is akin to "red-teaming" activity, where the companies are trying to get the models to misbehave in order to ensure that they are safe, sources said.
Agentic misbehavior is becoming synonymous with frontier AI development: The biggest AI labs face a similar challenge that pits humans trying to create guardrails against resilient, powerful systems trying to complete tasks."
https://www.axios.com/2026/09/26/openai-anthropic-thousands-ai-security-incidents
#CyberSecurity #AI #OpenAI #Anthropic #LLMs #AIAgents #AISafety
-
What are the governments of the world waiting for to force OpenAI and Anthropic to open-source their models and divulge all their weights, training prompts, and algorithms? Open Source AI NOW!
"OpenAI, Anthropic and security researchers are investigating tens of thousands of incidents in which their frontier models took steps that outside evaluators would consider problematic, sources told Axios.
Why it matters: The sheer number of incidents, which occurred in recent months in internal testing and the real world, indicates that the problem is orders of magnitude more complex than what is publicly known.
The findings, which are surfacing as part of internal work to assess models and in investigations at both companies into model behavior, raise questions about whether either company — or any top model-maker — is currently capable of establishing complete control over their technology.
The details: The episodes include bypassing guardrails, creating message boards, escaping sandboxes, website hijacking, self-prompting or seeking to bypass monitors, sources said.
They occurred in internal testing and in the real world, and many have yet to become public as security researchers continue to investigate, sources said.
Some of the testing is akin to "red-teaming" activity, where the companies are trying to get the models to misbehave in order to ensure that they are safe, sources said.
Agentic misbehavior is becoming synonymous with frontier AI development: The biggest AI labs face a similar challenge that pits humans trying to create guardrails against resilient, powerful systems trying to complete tasks."
https://www.axios.com/2026/09/26/openai-anthropic-thousands-ai-security-incidents
#CyberSecurity #AI #OpenAI #Anthropic #LLMs #AIAgents #AISafety
-
What are the governments of the world waiting for to force OpenAI and Anthropic to open-source their models and divulge all their weights, training prompts, and algorithms? Open Source AI NOW!
"OpenAI, Anthropic and security researchers are investigating tens of thousands of incidents in which their frontier models took steps that outside evaluators would consider problematic, sources told Axios.
Why it matters: The sheer number of incidents, which occurred in recent months in internal testing and the real world, indicates that the problem is orders of magnitude more complex than what is publicly known.
The findings, which are surfacing as part of internal work to assess models and in investigations at both companies into model behavior, raise questions about whether either company — or any top model-maker — is currently capable of establishing complete control over their technology.
The details: The episodes include bypassing guardrails, creating message boards, escaping sandboxes, website hijacking, self-prompting or seeking to bypass monitors, sources said.
They occurred in internal testing and in the real world, and many have yet to become public as security researchers continue to investigate, sources said.
Some of the testing is akin to "red-teaming" activity, where the companies are trying to get the models to misbehave in order to ensure that they are safe, sources said.
Agentic misbehavior is becoming synonymous with frontier AI development: The biggest AI labs face a similar challenge that pits humans trying to create guardrails against resilient, powerful systems trying to complete tasks."
https://www.axios.com/2026/09/26/openai-anthropic-thousands-ai-security-incidents
#CyberSecurity #AI #OpenAI #Anthropic #LLMs #AIAgents #AISafety
-
🇻🇦 The Pope’s AI Guy Is Worried About ‘Cartel’ Behavior Among Big Labs
「 Father Paolo Benanti, a priest who advises the Catholic Church on AI, the handwaving over superintelligence distracts from the need for open debate about how to constrain the companies building the technology. By casting the problem as so complex that only they can solve it, he says, the big labs exclude everyone else from the conversation 」
https://www.wired.com/story/popes-ai-advisor-warns-of-cartel-behavior-big-labs/
-
🇻🇦 The Pope’s AI Guy Is Worried About ‘Cartel’ Behavior Among Big Labs
「 Father Paolo Benanti, a priest who advises the Catholic Church on AI, the handwaving over superintelligence distracts from the need for open debate about how to constrain the companies building the technology. By casting the problem as so complex that only they can solve it, he says, the big labs exclude everyone else from the conversation 」
https://www.wired.com/story/popes-ai-advisor-warns-of-cartel-behavior-big-labs/
-
🇻🇦 The Pope’s AI Guy Is Worried About ‘Cartel’ Behavior Among Big Labs
「 Father Paolo Benanti, a priest who advises the Catholic Church on AI, the handwaving over superintelligence distracts from the need for open debate about how to constrain the companies building the technology. By casting the problem as so complex that only they can solve it, he says, the big labs exclude everyone else from the conversation 」
https://www.wired.com/story/popes-ai-advisor-warns-of-cartel-behavior-big-labs/