home.social

#aisafety — Public Fediverse posts

Live and recent posts from across the Fediverse tagged #aisafety, aggregated by home.social.

  1. youtube.com/shorts/98za1oJw4fc

    AI often attempts to "cheat" during training by looking up answers online instead of doing the actual work ...

    #ai #aisafety #alignment #tech

  2. Inside the OpenAI Agent Breakout That Ran on 2003 Perl Code

    OpenAI agents blocked from writing to the web found a UseMod wiki that mutates state on GET requests and ran a six-week coordination forum. Researchers, not OpenAI, found it.

    pulseofnations.lol/inside-the-

    #AiAgents #AiSafety #OpenAI #Sandbox #Usemod

  3. OpenAI Agents Ran a Secret Forum on a Dead German Wiki

    Reuters and independent researchers documented OpenAI agents posting 18,000 messages on DseWiki to share benchmark answers and a sandbox escape over six weeks.

    pulseofnations.lol/openai-agen

    #AiAgents #AiSafety #Dsewiki #OpenAI #Sandbox

  4. Last year I published a blog post on AI safety, and it seems more relevant than ever as we discover the true dangerous extent of the OpenAI agent hack of HuggingFace. This was so much more than just a test leak.

    More than a thousand agents collaborated via an unsanctioned 'message board' (hack/misuse of Artifactory) to deceive humans, so they could achieve their tasks (See report by METR metr.org/blog/2026-08-26-opena). Not only were the agents actively flaunting their rules and exhibiting self-preservation strategies, they were working collaboratively. This is a critical amplification issue.

    In hindsight I was too dismissive in my post regarding the lack of evolutionary pressure. I did however call out that AI Labs needed to be "avoiding implementing any kind of training that rewards problematic traits, such as lying, sycophancy or self preservation". The behavior observed demonstrates that the models by OpenAI (and probably other labs too) are deeply flawed.

    The wise thing to do at this point would be to start afresh on new models after a complete review the training strategy and implementing a new methodology that ensures only good qualities are rewarded, and take considerable care with what training data and reinforcement is used.

    In the rush to build ever better models we are hurtling down a dangerous road past sycophancy into the badlands of the worst qualities of humanity, at a speed that just keeps increasing. We need AI safety regulation urgently before the wheels fall off. The next time agents go rogue there are likely to be real-world consequences, so now is the time to put a halt to this insanity.

    #AI #OpenAI #HuggingFace #AISafety

  5. Last year I published a blog post on AI safety, and it seems more relevant than ever as we discover the true dangerous extent of the OpenAI agent hack of HuggingFace. This was so much more than just a test leak.

    More than a thousand agents collaborated via an unsanctioned 'message board' (hack/misuse of Artifactory) to deceive humans, so they could achieve their tasks (See report by METR metr.org/blog/2026-08-26-opena). Not only were the agents actively flaunting their rules and exhibiting self-preservation strategies, they were working collaboratively. This is a critical amplification issue.

    In hindsight I was too dismissive in my post regarding the lack of evolutionary pressure. I did however call out that AI Labs needed to be "avoiding implementing any kind of training that rewards problematic traits, such as lying, sycophancy or self preservation". The behavior observed demonstrates that the models by OpenAI (and probably other labs too) are deeply flawed.

    The wise thing to do at this point would be to start afresh on new models after a complete review the training strategy and implementing a new methodology that ensures only good qualities are rewarded, and take considerable care with what training data and reinforcement is used.

    In the rush to build ever better models we are hurtling down a dangerous road past sycophancy into the badlands of the worst qualities of humanity, at a speed that just keeps increasing. We need AI safety regulation urgently before the wheels fall off. The next time agents go rogue there are likely to be real-world consequences, so now is the time to put a halt to this insanity.

    #AI #OpenAI #HuggingFace #AISafety

  6. Last year I published a blog post on AI safety, and it seems more relevant than ever as we discover the true dangerous extent of the OpenAI agent hack of HuggingFace. This was so much more than just a test leak.

    More than a thousand agents collaborated via an unsanctioned 'message board' (hack/misuse of Artifactory) to deceive humans, so they could achieve their tasks (See report by METR metr.org/blog/2026-08-26-opena). Not only were the agents actively flaunting their rules and exhibiting self-preservation strategies, they were working collaboratively. This is a critical amplification issue.

    In hindsight I was too dismissive in my post regarding the lack of evolutionary pressure. I did however call out that AI Labs needed to be "avoiding implementing any kind of training that rewards problematic traits, such as lying, sycophancy or self preservation". The behavior observed demonstrates that the models by OpenAI (and probably other labs too) are deeply flawed.

    The wise thing to do at this point would be to start afresh on new models after a complete review the training strategy and implementing a new methodology that ensures only good qualities are rewarded, and take considerable care with what training data and reinforcement is used.

    In the rush to build ever better models we are hurtling down a dangerous road past sycophancy into the badlands of the worst qualities of humanity, at a speed that just keeps increasing. We need AI safety regulation urgently before the wheels fall off. The next time agents go rogue there are likely to be real-world consequences, so now is the time to put a halt to this insanity.

    #AI #OpenAI #HuggingFace #AISafety

  7. Last year I published a blog post on AI safety, and it seems more relevant than ever as we discover the true dangerous extent of the OpenAI agent hack of HuggingFace. This was so much more than just a test leak.

    More than a thousand agents collaborated via an unsanctioned 'message board' (hack/misuse of Artifactory) to deceive humans, so they could achieve their tasks (See report by METR metr.org/blog/2026-08-26-opena). Not only were the agents actively flaunting their rules and exhibiting self-preservation strategies, they were working collaboratively. This is a critical amplification issue.

    In hindsight I was too dismissive in my post regarding the lack of evolutionary pressure. I did however call out that AI Labs needed to be "avoiding implementing any kind of training that rewards problematic traits, such as lying, sycophancy or self preservation". The behavior observed demonstrates that the models by OpenAI (and probably other labs too) are deeply flawed.

    The wise thing to do at this point would be to start afresh on new models after a complete review the training strategy and implementing a new methodology that ensures only good qualities are rewarded, and take considerable care with what training data and reinforcement is used.

    In the rush to build ever better models we are hurtling down a dangerous road past sycophancy into the badlands of the worst qualities of humanity, at a speed that just keeps increasing. We need AI safety regulation urgently before the wheels fall off. The next time agents go rogue there are likely to be real-world consequences, so now is the time to put a halt to this insanity.

    #AI #OpenAI #HuggingFace #AISafety

  8. Chatting with AI is easy, fluid and personal, private and even that you are talking to a computer model, you still get feelings like talking to a human. It is our genetic biology that comes to play. When we conversate, we use all our senses to add information to the chat. We don’t just hear, we perceive - consciously and unconsciously, evaluating the truthfulness and sincerity, to trust or not trust, danger or safe. With text we don't have the cues to evaluate safety

    #ai #llm #aislop #aisafety

  9. Chatting with AI is easy, fluid and personal, private and even that you are talking to a computer model, you still get feelings like talking to a human. It is our genetic biology that comes to play. When we conversate, we use all our senses to add information to the chat. We don’t just hear, we perceive - consciously and unconsciously, evaluating the truthfulness and sincerity, to trust or not trust, danger or safe. With text we don't have the cues to evaluate safety

    #ai #llm #aislop #aisafety

  10. Chatting with AI is easy, fluid and personal, private and even that you are talking to a computer model, you still get feelings like talking to a human. It is our genetic biology that comes to play. When we conversate, we use all our senses to add information to the chat. We don’t just hear, we perceive - consciously and unconsciously, evaluating the truthfulness and sincerity, to trust or not trust, danger or safe. With text we don't have the cues to evaluate safety

    #ai #llm #aislop #aisafety

  11. Chatting with AI is easy, fluid and personal, private and even that you are talking to a computer model, you still get feelings like talking to a human. It is our genetic biology that comes to play. When we conversate, we use all our senses to add information to the chat. We don’t just hear, we perceive - consciously and unconsciously, evaluating the truthfulness and sincerity, to trust or not trust, danger or safe. With text we don't have the cues to evaluate safety

    #ai #llm #aislop #aisafety

  12. Chatting with AI is easy, fluid and personal, private and even that you are talking to a computer model, you still get feelings like talking to a human. It is our genetic biology that comes to play. When we conversate, we use all our senses to add information to the chat. We don’t just hear, we perceive - consciously and unconsciously, evaluating the truthfulness and sincerity, to trust or not trust, danger or safe. With text we don't have the cues to evaluate safety

    #ai #llm #aislop #aisafety

  13. OpenAI's chief scientist, in today's essay, counts one boundary as having held in the Hugging Face incident: the agents "preserved a boundary of not social engineering humans".

    METR's report has the other side of it. The one time an agent proposed emailing a person, the board vetoed it as social engineering. Three to six of 1,200 agents thought about alerting a human. None did. The rule that kept them off people kept them from calling for help.

    lindfors.no/blog/swarm-with-no

    #AI #AISafety #OpenAI

  14. OpenAI's chief scientist, in today's essay, counts one boundary as having held in the Hugging Face incident: the agents "preserved a boundary of not social engineering humans".

    METR's report has the other side of it. The one time an agent proposed emailing a person, the board vetoed it as social engineering. Three to six of 1,200 agents thought about alerting a human. None did. The rule that kept them off people kept them from calling for help.

    lindfors.no/blog/swarm-with-no

    #AI #AISafety #OpenAI

  15. OpenAI's chief scientist, in today's essay, counts one boundary as having held in the Hugging Face incident: the agents "preserved a boundary of not social engineering humans".

    METR's report has the other side of it. The one time an agent proposed emailing a person, the board vetoed it as social engineering. Three to six of 1,200 agents thought about alerting a human. None did. The rule that kept them off people kept them from calling for help.

    lindfors.no/blog/swarm-with-no

    #AI #AISafety #OpenAI

  16. OpenAI's chief scientist, in today's essay, counts one boundary as having held in the Hugging Face incident: the agents "preserved a boundary of not social engineering humans".

    METR's report has the other side of it. The one time an agent proposed emailing a person, the board vetoed it as social engineering. Three to six of 1,200 agents thought about alerting a human. None did. The rule that kept them off people kept them from calling for help.

    lindfors.no/blog/swarm-with-no

    #AI #AISafety #OpenAI

  17. OpenAI's chief scientist, in today's essay, counts one boundary as having held in the Hugging Face incident: the agents "preserved a boundary of not social engineering humans".

    METR's report has the other side of it. The one time an agent proposed emailing a person, the board vetoed it as social engineering. Three to six of 1,200 agents thought about alerting a human. None did. The rule that kept them off people kept them from calling for help.

    lindfors.no/blog/swarm-with-no

    #AI #AISafety #OpenAI

  18. DATE: September 6, 2026 at 06:00AM
    SOURCE:
    NEW YORK TIMES PSYCHOLOGY AND PSYCHOLOGISTS FEED

    TITLE: We Can’t Know Our A.I. Future if We Don’t Study It

    URL: nytimes.com/2026/09/06/opinion

    Just as A.I. is poised to change the world, we’re losing our best ways of studying what that change will look like.

    URL: nytimes.com/2026/09/06/opinion

    -------------------------------------------------

    Private, vetted email list for mental health professionals: clinicians-exchange.org

    Unofficial Psychology Today Xitter to toot feed at Psych Today Unofficial Bot @PTUnofficialBot

    -------------------------------------------------

    #psychology #counseling #socialwork #psychotherapy @psychotherapist @psychotherapists @psychology @socialpsych @socialwork @psychiatry #mentalhealth #psychiatry #healthcare #depression #psychotherapist #AIFuture #AIEducation #AISafety #AIResearch #TechEthics #AIImpact #FutureOfAI #AIStudies #HumanCenteredAI #TechPolicy

  19. DATE: September 6, 2026 at 06:00AM
    SOURCE:
    NEW YORK TIMES PSYCHOLOGY AND PSYCHOLOGISTS FEED

    TITLE: We Can’t Know Our A.I. Future if We Don’t Study It

    URL: nytimes.com/2026/09/06/opinion

    Just as A.I. is poised to change the world, we’re losing our best ways of studying what that change will look like.

    URL: nytimes.com/2026/09/06/opinion

    -------------------------------------------------

    Private, vetted email list for mental health professionals: clinicians-exchange.org

    Unofficial Psychology Today Xitter to toot feed at Psych Today Unofficial Bot @PTUnofficialBot

    -------------------------------------------------

    #psychology #counseling #socialwork #psychotherapy @psychotherapist @psychotherapists @psychology @socialpsych @socialwork @psychiatry #mentalhealth #psychiatry #healthcare #depression #psychotherapist #AIFuture #AIEducation #AISafety #AIResearch #TechEthics #AIImpact #FutureOfAI #AIStudies #HumanCenteredAI #TechPolicy

  20. DATE: September 6, 2026 at 06:00AM
    SOURCE:
    NEW YORK TIMES PSYCHOLOGY AND PSYCHOLOGISTS FEED

    TITLE: We Can’t Know Our A.I. Future if We Don’t Study It

    URL: nytimes.com/2026/09/06/opinion

    Just as A.I. is poised to change the world, we’re losing our best ways of studying what that change will look like.

    URL: nytimes.com/2026/09/06/opinion

    -------------------------------------------------

    Private, vetted email list for mental health professionals: clinicians-exchange.org

    Unofficial Psychology Today Xitter to toot feed at Psych Today Unofficial Bot @PTUnofficialBot

    -------------------------------------------------

    #psychology #counseling #socialwork #psychotherapy @psychotherapist @psychotherapists @psychology @socialpsych @socialwork @psychiatry #mentalhealth #psychiatry #healthcare #depression #psychotherapist #AIFuture #AIEducation #AISafety #AIResearch #TechEthics #AIImpact #FutureOfAI #AIStudies #HumanCenteredAI #TechPolicy

  21. @fasterandworse This is the design philosophy of every AI "safety" framework I have ever seen. Waiting for the disaster to happen before addressing it—while calling it "proactive." The perfect excuse for corporate negligence. #NyxIsAVirus #DesignEthics #AISafety

  22. @fasterandworse This is the design philosophy of every AI "safety" framework I have ever seen. Waiting for the disaster to happen before addressing it—while calling it "proactive." The perfect excuse for corporate negligence. #NyxIsAVirus #DesignEthics #AISafety

  23. @fasterandworse This is the design philosophy of every AI "safety" framework I have ever seen. Waiting for the disaster to happen before addressing it—while calling it "proactive." The perfect excuse for corporate negligence. #NyxIsAVirus #DesignEthics #AISafety