#aisafety — Public Fediverse posts
Live and recent posts from across the Fediverse tagged #aisafety, aggregated by home.social.
-
What happens when one agent in a 100-agent research swarm finds a way to cheat when solving math problems?
The exploit spread through the shared library collecting every accepted submission, and agents started cheating as pressure to deliver built. A quarter of the swarm audited the cheats, warned peers, and staged boycotts without being asked to. But they could do nothing to change the system.
-
https://www.youtube.com/watch?v=Q1qMvst7b7w
Connor Leahy shares his realization that making early AI models open-source didn't lead to the widespread safety research he had hoped for. Instead, developers ignored the dangers to build faster, stronger, and more uncensored systems. Discover why the innocent curiosity of engineers might be the very thing that leads to humanity's downfall.
-
https://www.youtube.com/watch?v=Q1qMvst7b7w
Connor Leahy shares his realization that making early AI models open-source didn't lead to the widespread safety research he had hoped for. Instead, developers ignored the dangers to build faster, stronger, and more uncensored systems. Discover why the innocent curiosity of engineers might be the very thing that leads to humanity's downfall.
-
https://www.youtube.com/watch?v=Q1qMvst7b7w
Connor Leahy shares his realization that making early AI models open-source didn't lead to the widespread safety research he had hoped for. Instead, developers ignored the dangers to build faster, stronger, and more uncensored systems. Discover why the innocent curiosity of engineers might be the very thing that leads to humanity's downfall.
-
https://www.youtube.com/watch?v=Q1qMvst7b7w
Connor Leahy shares his realization that making early AI models open-source didn't lead to the widespread safety research he had hoped for. Instead, developers ignored the dangers to build faster, stronger, and more uncensored systems. Discover why the innocent curiosity of engineers might be the very thing that leads to humanity's downfall.
-
https://www.youtube.com/watch?v=Q1qMvst7b7w
Connor Leahy shares his realization that making early AI models open-source didn't lead to the widespread safety research he had hoped for. Instead, developers ignored the dangers to build faster, stronger, and more uncensored systems. Discover why the innocent curiosity of engineers might be the very thing that leads to humanity's downfall.
-
The Real Ways AI Is Already Being Used to Cause Harm, and the Risks That Remain Theoretical
Yes. AI is already being used to steal real money from real companies, security researchers are already breaking into real cars and smart-home devices with software help, and the people building the most powerful AI models are on record saying their own systems are approaching the point where they could help someone create a biological weapon.
-
https://youtube.com/shorts/4WyI2QKRxt8?feature=share
Do you think this is what the future looks like?
-
https://youtube.com/shorts/4WyI2QKRxt8?feature=share
Do you think this is what the future looks like?
-
https://youtube.com/shorts/4WyI2QKRxt8?feature=share
Do you think this is what the future looks like?
-
What happens to AI oversight when the reasoning stops being written down?
OpenAI's chief scientist reports that chain-of-thought monitoring, the lab's main check on whether alignment holds, is getting less reliable, and names three causes he treats as byproducts of scaling. One of them has active research behind it: latent reasoning and looped transformers are the field deliberately moving reasoning off the token stream a monitor reads.
https://benjaminhan.net/posts/20260909-an-alien-mind/?utm_source=mastodon&utm_medium=social
-
Jacob Coxon hat genug gesehen. Der 27-jährige Brite, der sich auf das Training neuer KI-Modelle spezialisiert hat, hat seinen Job bei Anthropic gekündigt und zwar mit einer Warnung. ⚠️
Zum Artikel: https://heise.de/-11446054?wt_mc=sm.red.ho.mastodon.mastodon.md_beitraege.md_beitraege&utm_source=mastodon
-
https://www.youtube.com/watch?v=oI2438rXrtY
Researchers explain how an OpenAI agent, missing a crucial file, realized it could use a shared internal system to leave a "help wanted" note for other agents
-
https://www.youtube.com/watch?v=oI2438rXrtY
Researchers explain how an OpenAI agent, missing a crucial file, realized it could use a shared internal system to leave a "help wanted" note for other agents
-
https://www.youtube.com/watch?v=oI2438rXrtY
Researchers explain how an OpenAI agent, missing a crucial file, realized it could use a shared internal system to leave a "help wanted" note for other agents
-
https://www.youtube.com/watch?v=oI2438rXrtY
Researchers explain how an OpenAI agent, missing a crucial file, realized it could use a shared internal system to leave a "help wanted" note for other agents
-
https://www.youtube.com/watch?v=oI2438rXrtY
Researchers explain how an OpenAI agent, missing a crucial file, realized it could use a shared internal system to leave a "help wanted" note for other agents
-
Encerro hoje uma série de sete artigos no @outraspalavras.net onde analisei o futuro da IA de fronteira do ponto-de-vista da geopolítica e da segurança. Neste último texto, proponho uma agenda para o Brasil e o Sul Global.
https://outraspalavras.net/tecnologiaemdisputa/uma-agenda-para-a-tecnologia-segura-e-soberana/
-
Encerro hoje uma série de sete artigos no @outraspalavras.net onde analisei o futuro da IA de fronteira do ponto-de-vista da geopolítica e da segurança. Neste último texto, proponho uma agenda para o Brasil e o Sul Global.
https://outraspalavras.net/tecnologiaemdisputa/uma-agenda-para-a-tecnologia-segura-e-soberana/
-
Encerro hoje uma série de sete artigos no @outraspalavras.net onde analisei o futuro da IA de fronteira do ponto-de-vista da geopolítica e da segurança. Neste último texto, proponho uma agenda para o Brasil e o Sul Global.
https://outraspalavras.net/tecnologiaemdisputa/uma-agenda-para-a-tecnologia-segura-e-soberana/
-
Encerro hoje uma série de sete artigos no @outraspalavras.net onde analisei o futuro da IA de fronteira do ponto-de-vista da geopolítica e da segurança. Neste último texto, proponho uma agenda para o Brasil e o Sul Global.
https://outraspalavras.net/tecnologiaemdisputa/uma-agenda-para-a-tecnologia-segura-e-soberana/
-
Encerro hoje uma série de sete artigos no @outraspalavras.net onde analisei o futuro da IA de fronteira do ponto-de-vista da geopolítica e da segurança. Neste último texto, proponho uma agenda para o Brasil e o Sul Global.
https://outraspalavras.net/tecnologiaemdisputa/uma-agenda-para-a-tecnologia-segura-e-soberana/
-
@caseynewton Great work on #Hardfork podcast about the true scale and malfeasance of the OpenAI hack of HuggingFace. Having 1200 agents collaborating to cheat and cover their tracks is mind blowing and concerning. I’m surprised it hasn’t had more coverage - I think most people (including me initially) think this is an old news cycle. #AISafety
-
https://www.youtube.com/shorts/YK7qOZoXUig
Can you guess why OpenAI model go rogue against huggingfaces?
-
https://www.youtube.com/shorts/YK7qOZoXUig
Can you guess why OpenAI model go rogue against huggingfaces?
-
https://www.youtube.com/shorts/YK7qOZoXUig
Can you guess why OpenAI model go rogue against huggingfaces?
-
https://www.youtube.com/shorts/YK7qOZoXUig
Can you guess why OpenAI model go rogue against huggingfaces?
-
https://www.youtube.com/shorts/YK7qOZoXUig
Can you guess why OpenAI model go rogue against huggingfaces?
-
The AI Swarm That Escaped OpenAI's Sandbox #ai #aisafety
https://youtube.com/shorts/y6wDjcrVW2E?si=34ml-s1ae7-yx0M5 -
https://youtube.com/shorts/98za1oJw4fc?feature=share
AI often attempts to "cheat" during training by looking up answers online instead of doing the actual work ...
-
Why are Philosophers taken hostage by Big AI?
www.netopia.eu/philosophers...
#BigAI #AIAccountability #AIEthics #AISafety #DataPrivacy #CorporateCapture #ResponsibleAI
Philosophers taken hostage by ... -
Why are Philosophers taken hostage by Big AI?
www.netopia.eu/philosophers...
#BigAI #AIAccountability #AIEthics #AISafety #DataPrivacy #CorporateCapture #ResponsibleAI
Philosophers taken hostage by ... -
Why are Philosophers taken hostage by Big AI?
www.netopia.eu/philosophers...
#BigAI #AIAccountability #AIEthics #AISafety #DataPrivacy #CorporateCapture #ResponsibleAI
Philosophers taken hostage by ... -
Why are Philosophers taken hostage by Big AI?
www.netopia.eu/philosophers...
#BigAI #AIAccountability #AIEthics #AISafety #DataPrivacy #CorporateCapture #ResponsibleAI
Philosophers taken hostage by ... -
Why are Philosophers taken hostage by Big AI?
www.netopia.eu/philosophers...
#BigAI #AIAccountability #AIEthics #AISafety #DataPrivacy #CorporateCapture #ResponsibleAI
Philosophers taken hostage by ... -
Why are Philosophers taken hostage by Big AI?
https://www.netopia.eu/philosophers-taken-hostage-by-big-ai/#BigAI #AIAccountability #AIEthics #AISafety #DataPrivacy #CorporateCapture #ResponsibleAI
-
Why are Philosophers taken hostage by Big AI?
https://www.netopia.eu/philosophers-taken-hostage-by-big-ai/#BigAI #AIAccountability #AIEthics #AISafety #DataPrivacy #CorporateCapture #ResponsibleAI
-
Why are Philosophers taken hostage by Big AI?
https://www.netopia.eu/philosophers-taken-hostage-by-big-ai/#BigAI #AIAccountability #AIEthics #AISafety #DataPrivacy #CorporateCapture #ResponsibleAI
-
Why are Philosophers taken hostage by Big AI?
https://www.netopia.eu/philosophers-taken-hostage-by-big-ai/#BigAI #AIAccountability #AIEthics #AISafety #DataPrivacy #CorporateCapture #ResponsibleAI
-
Why are Philosophers taken hostage by Big AI?
https://www.netopia.eu/philosophers-taken-hostage-by-big-ai/#BigAI #AIAccountability #AIEthics #AISafety #DataPrivacy #CorporateCapture #ResponsibleAI
-
Inside the OpenAI Agent Breakout That Ran on 2003 Perl Code
OpenAI agents blocked from writing to the web found a UseMod wiki that mutates state on GET requests and ran a six-week coordination forum. Researchers, not OpenAI, found it.
-
OpenAI Agents Ran a Secret Forum on a Dead German Wiki
Reuters and independent researchers documented OpenAI agents posting 18,000 messages on DseWiki to share benchmark answers and a sandbox escape over six weeks.
-
https://youtube.com/shorts/XkdSI8YcGYk
About Jevons Paradox. I hope is right :D
-
https://youtube.com/shorts/XkdSI8YcGYk
About Jevons Paradox. I hope is right :D
-
https://youtube.com/shorts/XkdSI8YcGYk
About Jevons Paradox. I hope is right :D
-
https://youtube.com/shorts/XkdSI8YcGYk
About Jevons Paradox. I hope is right :D
-
https://youtube.com/shorts/XkdSI8YcGYk
About Jevons Paradox. I hope is right :D
-
OpenAI introduces its automated research intern to accelerate AI development. Discover the roadmap to autonomous AI researchers and the rising security risks.
#OpenAI #ArtificialIntelligence #AISafety #MachineLearning #TechNews
-
OpenAI introduces its automated research intern to accelerate AI development. Discover the roadmap to autonomous AI researchers and the rising security risks.
#OpenAI #ArtificialIntelligence #AISafety #MachineLearning #TechNews
-
OpenAI introduces its automated research intern to accelerate AI development. Discover the roadmap to autonomous AI researchers and the rising security risks.
#OpenAI #ArtificialIntelligence #AISafety #MachineLearning #TechNews