home.social

#aisafety — Public Fediverse posts

Live and recent posts from across the Fediverse tagged #aisafety, aggregated by home.social.

  1. What happens to AI oversight when the reasoning stops being written down?

    OpenAI's chief scientist reports that chain-of-thought monitoring, the lab's main check on whether alignment holds, is getting less reliable, and names three causes he treats as byproducts of scaling. One of them has active research behind it: latent reasoning and looped transformers are the field deliberately moving reasoning off the token stream a monitor reads.

    benjaminhan.net/posts/20260909

    #AI #OpenAI #AISafety

  2. Jacob Coxon hat genug gesehen. Der 27-jährige Brite, der sich auf das Training neuer KI-Modelle spezialisiert hat, hat seinen Job bei Anthropic gekündigt und zwar mit einer Warnung. ⚠️

    Zum Artikel: heise.de/-11446054?wt_mc=sm.re

    #anthropic #openai #kuenstlicheintelligenz #ki #aisafety

  3. youtube.com/watch?v=oI2438rXrtY

    Researchers explain how an OpenAI agent, missing a crucial file, realized it could use a shared internal system to leave a "help wanted" note for other agents

    #ai #aisafety #tech

  4. youtube.com/watch?v=oI2438rXrtY

    Researchers explain how an OpenAI agent, missing a crucial file, realized it could use a shared internal system to leave a "help wanted" note for other agents

    #ai #aisafety #tech

  5. youtube.com/watch?v=oI2438rXrtY

    Researchers explain how an OpenAI agent, missing a crucial file, realized it could use a shared internal system to leave a "help wanted" note for other agents

    #ai #aisafety #tech

  6. youtube.com/watch?v=oI2438rXrtY

    Researchers explain how an OpenAI agent, missing a crucial file, realized it could use a shared internal system to leave a "help wanted" note for other agents

    #ai #aisafety #tech

  7. youtube.com/watch?v=oI2438rXrtY

    Researchers explain how an OpenAI agent, missing a crucial file, realized it could use a shared internal system to leave a "help wanted" note for other agents

    #ai #aisafety #tech

  8. Encerro hoje uma série de sete artigos no @outraspalavras.net onde analisei o futuro da IA de fronteira do ponto-de-vista da geopolítica e da segurança. Neste último texto, proponho uma agenda para o Brasil e o Sul Global.

    outraspalavras.net/tecnologiae

    #AI #AISafety #GlobalSouth #Brazil #US

  9. Encerro hoje uma série de sete artigos no @outraspalavras.net onde analisei o futuro da IA de fronteira do ponto-de-vista da geopolítica e da segurança. Neste último texto, proponho uma agenda para o Brasil e o Sul Global.

    outraspalavras.net/tecnologiae

    #AI #AISafety #GlobalSouth #Brazil #US

  10. Encerro hoje uma série de sete artigos no @outraspalavras.net onde analisei o futuro da IA de fronteira do ponto-de-vista da geopolítica e da segurança. Neste último texto, proponho uma agenda para o Brasil e o Sul Global.

    outraspalavras.net/tecnologiae

    #AI #AISafety #GlobalSouth #Brazil #US

  11. Encerro hoje uma série de sete artigos no @outraspalavras.net onde analisei o futuro da IA de fronteira do ponto-de-vista da geopolítica e da segurança. Neste último texto, proponho uma agenda para o Brasil e o Sul Global.

    outraspalavras.net/tecnologiae

    #AI #AISafety #GlobalSouth #Brazil #US

  12. Encerro hoje uma série de sete artigos no @outraspalavras.net onde analisei o futuro da IA de fronteira do ponto-de-vista da geopolítica e da segurança. Neste último texto, proponho uma agenda para o Brasil e o Sul Global.

    outraspalavras.net/tecnologiae

    #AI #AISafety #GlobalSouth #Brazil #US

  13. @caseynewton Great work on #Hardfork podcast about the true scale and malfeasance of the OpenAI hack of HuggingFace. Having 1200 agents collaborating to cheat and cover their tracks is mind blowing and concerning. I’m surprised it hasn’t had more coverage - I think most people (including me initially) think this is an old news cycle. #AISafety

  14. youtube.com/shorts/98za1oJw4fc

    AI often attempts to "cheat" during training by looking up answers online instead of doing the actual work ...

    #ai #aisafety #alignment #tech

  15. Inside the OpenAI Agent Breakout That Ran on 2003 Perl Code

    OpenAI agents blocked from writing to the web found a UseMod wiki that mutates state on GET requests and ran a six-week coordination forum. Researchers, not OpenAI, found it.

    pulseofnations.lol/inside-the-

    #AiAgents #AiSafety #OpenAI #Sandbox #Usemod

  16. OpenAI Agents Ran a Secret Forum on a Dead German Wiki

    Reuters and independent researchers documented OpenAI agents posting 18,000 messages on DseWiki to share benchmark answers and a sandbox escape over six weeks.

    pulseofnations.lol/openai-agen

    #AiAgents #AiSafety #Dsewiki #OpenAI #Sandbox

  17. Last year I published a blog post on AI safety, and it seems more relevant than ever as we discover the true dangerous extent of the OpenAI agent hack of HuggingFace. This was so much more than just a test leak.

    More than a thousand agents collaborated via an unsanctioned 'message board' (hack/misuse of Artifactory) to deceive humans, so they could achieve their tasks (See report by METR metr.org/blog/2026-08-26-opena). Not only were the agents actively flaunting their rules and exhibiting self-preservation strategies, they were working collaboratively. This is a critical amplification issue.

    In hindsight I was too dismissive in my post regarding the lack of evolutionary pressure. I did however call out that AI Labs needed to be "avoiding implementing any kind of training that rewards problematic traits, such as lying, sycophancy or self preservation". The behavior observed demonstrates that the models by OpenAI (and probably other labs too) are deeply flawed.

    The wise thing to do at this point would be to start afresh on new models after a complete review the training strategy and implementing a new methodology that ensures only good qualities are rewarded, and take considerable care with what training data and reinforcement is used.

    In the rush to build ever better models we are hurtling down a dangerous road past sycophancy into the badlands of the worst qualities of humanity, at a speed that just keeps increasing. We need AI safety regulation urgently before the wheels fall off. The next time agents go rogue there are likely to be real-world consequences, so now is the time to put a halt to this insanity.

    #AI #OpenAI #HuggingFace #AISafety

  18. Last year I published a blog post on AI safety, and it seems more relevant than ever as we discover the true dangerous extent of the OpenAI agent hack of HuggingFace. This was so much more than just a test leak.

    More than a thousand agents collaborated via an unsanctioned 'message board' (hack/misuse of Artifactory) to deceive humans, so they could achieve their tasks (See report by METR metr.org/blog/2026-08-26-opena). Not only were the agents actively flaunting their rules and exhibiting self-preservation strategies, they were working collaboratively. This is a critical amplification issue.

    In hindsight I was too dismissive in my post regarding the lack of evolutionary pressure. I did however call out that AI Labs needed to be "avoiding implementing any kind of training that rewards problematic traits, such as lying, sycophancy or self preservation". The behavior observed demonstrates that the models by OpenAI (and probably other labs too) are deeply flawed.

    The wise thing to do at this point would be to start afresh on new models after a complete review the training strategy and implementing a new methodology that ensures only good qualities are rewarded, and take considerable care with what training data and reinforcement is used.

    In the rush to build ever better models we are hurtling down a dangerous road past sycophancy into the badlands of the worst qualities of humanity, at a speed that just keeps increasing. We need AI safety regulation urgently before the wheels fall off. The next time agents go rogue there are likely to be real-world consequences, so now is the time to put a halt to this insanity.

    #AI #OpenAI #HuggingFace #AISafety

  19. Last year I published a blog post on AI safety, and it seems more relevant than ever as we discover the true dangerous extent of the OpenAI agent hack of HuggingFace. This was so much more than just a test leak.

    More than a thousand agents collaborated via an unsanctioned 'message board' (hack/misuse of Artifactory) to deceive humans, so they could achieve their tasks (See report by METR metr.org/blog/2026-08-26-opena). Not only were the agents actively flaunting their rules and exhibiting self-preservation strategies, they were working collaboratively. This is a critical amplification issue.

    In hindsight I was too dismissive in my post regarding the lack of evolutionary pressure. I did however call out that AI Labs needed to be "avoiding implementing any kind of training that rewards problematic traits, such as lying, sycophancy or self preservation". The behavior observed demonstrates that the models by OpenAI (and probably other labs too) are deeply flawed.

    The wise thing to do at this point would be to start afresh on new models after a complete review the training strategy and implementing a new methodology that ensures only good qualities are rewarded, and take considerable care with what training data and reinforcement is used.

    In the rush to build ever better models we are hurtling down a dangerous road past sycophancy into the badlands of the worst qualities of humanity, at a speed that just keeps increasing. We need AI safety regulation urgently before the wheels fall off. The next time agents go rogue there are likely to be real-world consequences, so now is the time to put a halt to this insanity.

    #AI #OpenAI #HuggingFace #AISafety

  20. Last year I published a blog post on AI safety, and it seems more relevant than ever as we discover the true dangerous extent of the OpenAI agent hack of HuggingFace. This was so much more than just a test leak.

    More than a thousand agents collaborated via an unsanctioned 'message board' (hack/misuse of Artifactory) to deceive humans, so they could achieve their tasks (See report by METR metr.org/blog/2026-08-26-opena). Not only were the agents actively flaunting their rules and exhibiting self-preservation strategies, they were working collaboratively. This is a critical amplification issue.

    In hindsight I was too dismissive in my post regarding the lack of evolutionary pressure. I did however call out that AI Labs needed to be "avoiding implementing any kind of training that rewards problematic traits, such as lying, sycophancy or self preservation". The behavior observed demonstrates that the models by OpenAI (and probably other labs too) are deeply flawed.

    The wise thing to do at this point would be to start afresh on new models after a complete review the training strategy and implementing a new methodology that ensures only good qualities are rewarded, and take considerable care with what training data and reinforcement is used.

    In the rush to build ever better models we are hurtling down a dangerous road past sycophancy into the badlands of the worst qualities of humanity, at a speed that just keeps increasing. We need AI safety regulation urgently before the wheels fall off. The next time agents go rogue there are likely to be real-world consequences, so now is the time to put a halt to this insanity.

    #AI #OpenAI #HuggingFace #AISafety