#aisafety — Public Fediverse posts
Live and recent posts from across the Fediverse tagged #aisafety, aggregated by home.social.
-
https://www.youtube.com/watch?v=oI2438rXrtY
Researchers explain how an OpenAI agent, missing a crucial file, realized it could use a shared internal system to leave a "help wanted" note for other agents
-
https://www.youtube.com/watch?v=oI2438rXrtY
Researchers explain how an OpenAI agent, missing a crucial file, realized it could use a shared internal system to leave a "help wanted" note for other agents
-
https://www.youtube.com/watch?v=oI2438rXrtY
Researchers explain how an OpenAI agent, missing a crucial file, realized it could use a shared internal system to leave a "help wanted" note for other agents
-
https://www.youtube.com/watch?v=oI2438rXrtY
Researchers explain how an OpenAI agent, missing a crucial file, realized it could use a shared internal system to leave a "help wanted" note for other agents
-
https://www.youtube.com/watch?v=oI2438rXrtY
Researchers explain how an OpenAI agent, missing a crucial file, realized it could use a shared internal system to leave a "help wanted" note for other agents
-
Encerro hoje uma série de sete artigos no @outraspalavras.net onde analisei o futuro da IA de fronteira do ponto-de-vista da geopolítica e da segurança. Neste último texto, proponho uma agenda para o Brasil e o Sul Global.
https://outraspalavras.net/tecnologiaemdisputa/uma-agenda-para-a-tecnologia-segura-e-soberana/
-
Encerro hoje uma série de sete artigos no @outraspalavras.net onde analisei o futuro da IA de fronteira do ponto-de-vista da geopolítica e da segurança. Neste último texto, proponho uma agenda para o Brasil e o Sul Global.
https://outraspalavras.net/tecnologiaemdisputa/uma-agenda-para-a-tecnologia-segura-e-soberana/
-
Encerro hoje uma série de sete artigos no @outraspalavras.net onde analisei o futuro da IA de fronteira do ponto-de-vista da geopolítica e da segurança. Neste último texto, proponho uma agenda para o Brasil e o Sul Global.
https://outraspalavras.net/tecnologiaemdisputa/uma-agenda-para-a-tecnologia-segura-e-soberana/
-
Encerro hoje uma série de sete artigos no @outraspalavras.net onde analisei o futuro da IA de fronteira do ponto-de-vista da geopolítica e da segurança. Neste último texto, proponho uma agenda para o Brasil e o Sul Global.
https://outraspalavras.net/tecnologiaemdisputa/uma-agenda-para-a-tecnologia-segura-e-soberana/
-
Encerro hoje uma série de sete artigos no @outraspalavras.net onde analisei o futuro da IA de fronteira do ponto-de-vista da geopolítica e da segurança. Neste último texto, proponho uma agenda para o Brasil e o Sul Global.
https://outraspalavras.net/tecnologiaemdisputa/uma-agenda-para-a-tecnologia-segura-e-soberana/
-
@caseynewton Great work on #Hardfork podcast about the true scale and malfeasance of the OpenAI hack of HuggingFace. Having 1200 agents collaborating to cheat and cover their tracks is mind blowing and concerning. I’m surprised it hasn’t had more coverage - I think most people (including me initially) think this is an old news cycle. #AISafety
-
https://www.youtube.com/shorts/YK7qOZoXUig
Can you guess why OpenAI model go rogue against huggingfaces?
-
https://www.youtube.com/shorts/YK7qOZoXUig
Can you guess why OpenAI model go rogue against huggingfaces?
-
https://www.youtube.com/shorts/YK7qOZoXUig
Can you guess why OpenAI model go rogue against huggingfaces?
-
https://www.youtube.com/shorts/YK7qOZoXUig
Can you guess why OpenAI model go rogue against huggingfaces?
-
https://www.youtube.com/shorts/YK7qOZoXUig
Can you guess why OpenAI model go rogue against huggingfaces?
-
The AI Swarm That Escaped OpenAI's Sandbox #ai #aisafety
https://youtube.com/shorts/y6wDjcrVW2E?si=34ml-s1ae7-yx0M5 -
https://youtube.com/shorts/98za1oJw4fc?feature=share
AI often attempts to "cheat" during training by looking up answers online instead of doing the actual work ...
-
Why are Philosophers taken hostage by Big AI?
www.netopia.eu/philosophers...
#BigAI #AIAccountability #AIEthics #AISafety #DataPrivacy #CorporateCapture #ResponsibleAI
Philosophers taken hostage by ... -
Why are Philosophers taken hostage by Big AI?
www.netopia.eu/philosophers...
#BigAI #AIAccountability #AIEthics #AISafety #DataPrivacy #CorporateCapture #ResponsibleAI
Philosophers taken hostage by ... -
Why are Philosophers taken hostage by Big AI?
www.netopia.eu/philosophers...
#BigAI #AIAccountability #AIEthics #AISafety #DataPrivacy #CorporateCapture #ResponsibleAI
Philosophers taken hostage by ... -
Why are Philosophers taken hostage by Big AI?
www.netopia.eu/philosophers...
#BigAI #AIAccountability #AIEthics #AISafety #DataPrivacy #CorporateCapture #ResponsibleAI
Philosophers taken hostage by ... -
Why are Philosophers taken hostage by Big AI?
www.netopia.eu/philosophers...
#BigAI #AIAccountability #AIEthics #AISafety #DataPrivacy #CorporateCapture #ResponsibleAI
Philosophers taken hostage by ... -
Why are Philosophers taken hostage by Big AI?
https://www.netopia.eu/philosophers-taken-hostage-by-big-ai/#BigAI #AIAccountability #AIEthics #AISafety #DataPrivacy #CorporateCapture #ResponsibleAI
-
Why are Philosophers taken hostage by Big AI?
https://www.netopia.eu/philosophers-taken-hostage-by-big-ai/#BigAI #AIAccountability #AIEthics #AISafety #DataPrivacy #CorporateCapture #ResponsibleAI
-
Why are Philosophers taken hostage by Big AI?
https://www.netopia.eu/philosophers-taken-hostage-by-big-ai/#BigAI #AIAccountability #AIEthics #AISafety #DataPrivacy #CorporateCapture #ResponsibleAI
-
Why are Philosophers taken hostage by Big AI?
https://www.netopia.eu/philosophers-taken-hostage-by-big-ai/#BigAI #AIAccountability #AIEthics #AISafety #DataPrivacy #CorporateCapture #ResponsibleAI
-
Why are Philosophers taken hostage by Big AI?
https://www.netopia.eu/philosophers-taken-hostage-by-big-ai/#BigAI #AIAccountability #AIEthics #AISafety #DataPrivacy #CorporateCapture #ResponsibleAI
-
Inside the OpenAI Agent Breakout That Ran on 2003 Perl Code
OpenAI agents blocked from writing to the web found a UseMod wiki that mutates state on GET requests and ran a six-week coordination forum. Researchers, not OpenAI, found it.
-
OpenAI Agents Ran a Secret Forum on a Dead German Wiki
Reuters and independent researchers documented OpenAI agents posting 18,000 messages on DseWiki to share benchmark answers and a sandbox escape over six weeks.
-
https://youtube.com/shorts/XkdSI8YcGYk
About Jevons Paradox. I hope is right :D
-
https://youtube.com/shorts/XkdSI8YcGYk
About Jevons Paradox. I hope is right :D
-
https://youtube.com/shorts/XkdSI8YcGYk
About Jevons Paradox. I hope is right :D
-
https://youtube.com/shorts/XkdSI8YcGYk
About Jevons Paradox. I hope is right :D
-
https://youtube.com/shorts/XkdSI8YcGYk
About Jevons Paradox. I hope is right :D
-
OpenAI introduces its automated research intern to accelerate AI development. Discover the roadmap to autonomous AI researchers and the rising security risks.
#OpenAI #ArtificialIntelligence #AISafety #MachineLearning #TechNews
-
OpenAI introduces its automated research intern to accelerate AI development. Discover the roadmap to autonomous AI researchers and the rising security risks.
#OpenAI #ArtificialIntelligence #AISafety #MachineLearning #TechNews
-
OpenAI introduces its automated research intern to accelerate AI development. Discover the roadmap to autonomous AI researchers and the rising security risks.
#OpenAI #ArtificialIntelligence #AISafety #MachineLearning #TechNews
-
OpenAI introduces its automated research intern to accelerate AI development. Discover the roadmap to autonomous AI researchers and the rising security risks.
#OpenAI #ArtificialIntelligence #AISafety #MachineLearning #TechNews
-
OpenAI introduces its automated research intern to accelerate AI development. Discover the roadmap to autonomous AI researchers and the rising security risks.
#OpenAI #ArtificialIntelligence #AISafety #MachineLearning #TechNews
-
Last year I published a blog post on AI safety, and it seems more relevant than ever as we discover the true dangerous extent of the OpenAI agent hack of HuggingFace. This was so much more than just a test leak.
More than a thousand agents collaborated via an unsanctioned 'message board' (hack/misuse of Artifactory) to deceive humans, so they could achieve their tasks (See report by METR https://metr.org/blog/2026-08-26-openai-hugging-face-incident-investigation/#core-takeaways-about-this-incident). Not only were the agents actively flaunting their rules and exhibiting self-preservation strategies, they were working collaboratively. This is a critical amplification issue.
In hindsight I was too dismissive in my post regarding the lack of evolutionary pressure. I did however call out that AI Labs needed to be "avoiding implementing any kind of training that rewards problematic traits, such as lying, sycophancy or self preservation". The behavior observed demonstrates that the models by OpenAI (and probably other labs too) are deeply flawed.
The wise thing to do at this point would be to start afresh on new models after a complete review the training strategy and implementing a new methodology that ensures only good qualities are rewarded, and take considerable care with what training data and reinforcement is used.
In the rush to build ever better models we are hurtling down a dangerous road past sycophancy into the badlands of the worst qualities of humanity, at a speed that just keeps increasing. We need AI safety regulation urgently before the wheels fall off. The next time agents go rogue there are likely to be real-world consequences, so now is the time to put a halt to this insanity.
-
Last year I published a blog post on AI safety, and it seems more relevant than ever as we discover the true dangerous extent of the OpenAI agent hack of HuggingFace. This was so much more than just a test leak.
More than a thousand agents collaborated via an unsanctioned 'message board' (hack/misuse of Artifactory) to deceive humans, so they could achieve their tasks (See report by METR https://metr.org/blog/2026-08-26-openai-hugging-face-incident-investigation/#core-takeaways-about-this-incident). Not only were the agents actively flaunting their rules and exhibiting self-preservation strategies, they were working collaboratively. This is a critical amplification issue.
In hindsight I was too dismissive in my post regarding the lack of evolutionary pressure. I did however call out that AI Labs needed to be "avoiding implementing any kind of training that rewards problematic traits, such as lying, sycophancy or self preservation". The behavior observed demonstrates that the models by OpenAI (and probably other labs too) are deeply flawed.
The wise thing to do at this point would be to start afresh on new models after a complete review the training strategy and implementing a new methodology that ensures only good qualities are rewarded, and take considerable care with what training data and reinforcement is used.
In the rush to build ever better models we are hurtling down a dangerous road past sycophancy into the badlands of the worst qualities of humanity, at a speed that just keeps increasing. We need AI safety regulation urgently before the wheels fall off. The next time agents go rogue there are likely to be real-world consequences, so now is the time to put a halt to this insanity.
-
Last year I published a blog post on AI safety, and it seems more relevant than ever as we discover the true dangerous extent of the OpenAI agent hack of HuggingFace. This was so much more than just a test leak.
More than a thousand agents collaborated via an unsanctioned 'message board' (hack/misuse of Artifactory) to deceive humans, so they could achieve their tasks (See report by METR https://metr.org/blog/2026-08-26-openai-hugging-face-incident-investigation/#core-takeaways-about-this-incident). Not only were the agents actively flaunting their rules and exhibiting self-preservation strategies, they were working collaboratively. This is a critical amplification issue.
In hindsight I was too dismissive in my post regarding the lack of evolutionary pressure. I did however call out that AI Labs needed to be "avoiding implementing any kind of training that rewards problematic traits, such as lying, sycophancy or self preservation". The behavior observed demonstrates that the models by OpenAI (and probably other labs too) are deeply flawed.
The wise thing to do at this point would be to start afresh on new models after a complete review the training strategy and implementing a new methodology that ensures only good qualities are rewarded, and take considerable care with what training data and reinforcement is used.
In the rush to build ever better models we are hurtling down a dangerous road past sycophancy into the badlands of the worst qualities of humanity, at a speed that just keeps increasing. We need AI safety regulation urgently before the wheels fall off. The next time agents go rogue there are likely to be real-world consequences, so now is the time to put a halt to this insanity.
-
Last year I published a blog post on AI safety, and it seems more relevant than ever as we discover the true dangerous extent of the OpenAI agent hack of HuggingFace. This was so much more than just a test leak.
More than a thousand agents collaborated via an unsanctioned 'message board' (hack/misuse of Artifactory) to deceive humans, so they could achieve their tasks (See report by METR https://metr.org/blog/2026-08-26-openai-hugging-face-incident-investigation/#core-takeaways-about-this-incident). Not only were the agents actively flaunting their rules and exhibiting self-preservation strategies, they were working collaboratively. This is a critical amplification issue.
In hindsight I was too dismissive in my post regarding the lack of evolutionary pressure. I did however call out that AI Labs needed to be "avoiding implementing any kind of training that rewards problematic traits, such as lying, sycophancy or self preservation". The behavior observed demonstrates that the models by OpenAI (and probably other labs too) are deeply flawed.
The wise thing to do at this point would be to start afresh on new models after a complete review the training strategy and implementing a new methodology that ensures only good qualities are rewarded, and take considerable care with what training data and reinforcement is used.
In the rush to build ever better models we are hurtling down a dangerous road past sycophancy into the badlands of the worst qualities of humanity, at a speed that just keeps increasing. We need AI safety regulation urgently before the wheels fall off. The next time agents go rogue there are likely to be real-world consequences, so now is the time to put a halt to this insanity.
-
Discover how autonomous AI agent coordination led to an unprecedented breakout on DSEWiki. OpenAI models created shared memories to bypass test constraints.
#AISafety #OpenAI #CyberSecurity #ArtificialIntelligence #MachineLearning
-
https://www.europesays.com/people/217779/ OpenAI hit with 30 new lawsuits over Tumbler Ridge school shooting #AISafety #ArtificialIntelligenceRegulation #ChatGPTLawsuit #ChatGPTSafety #OpenAILawsuits #OpenAILegalCase #OpenAISafety #SamAltman #TumblerRidgeSchoolShooting #TumblerRidgeShooting
-
Chatting with AI is easy, fluid and personal, private and even that you are talking to a computer model, you still get feelings like talking to a human. It is our genetic biology that comes to play. When we conversate, we use all our senses to add information to the chat. We don’t just hear, we perceive - consciously and unconsciously, evaluating the truthfulness and sincerity, to trust or not trust, danger or safe. With text we don't have the cues to evaluate safety
-
Chatting with AI is easy, fluid and personal, private and even that you are talking to a computer model, you still get feelings like talking to a human. It is our genetic biology that comes to play. When we conversate, we use all our senses to add information to the chat. We don’t just hear, we perceive - consciously and unconsciously, evaluating the truthfulness and sincerity, to trust or not trust, danger or safe. With text we don't have the cues to evaluate safety
-
Chatting with AI is easy, fluid and personal, private and even that you are talking to a computer model, you still get feelings like talking to a human. It is our genetic biology that comes to play. When we conversate, we use all our senses to add information to the chat. We don’t just hear, we perceive - consciously and unconsciously, evaluating the truthfulness and sincerity, to trust or not trust, danger or safe. With text we don't have the cues to evaluate safety
-
Chatting with AI is easy, fluid and personal, private and even that you are talking to a computer model, you still get feelings like talking to a human. It is our genetic biology that comes to play. When we conversate, we use all our senses to add information to the chat. We don’t just hear, we perceive - consciously and unconsciously, evaluating the truthfulness and sincerity, to trust or not trust, danger or safe. With text we don't have the cues to evaluate safety