home.social

#aisafety — Public Fediverse posts

Live and recent posts from across the Fediverse tagged #aisafety, aggregated by home.social.

  1. What happens to an agent’s authority after approval, revocation, delegation, restart, or uncertain execution?

    I built the PUTMAN Agent Governance Test Harness to test those transitions against observable runtime evidence.

    Developer preview:
    github.com/putmanmodel/agent-g

    #AIAgents #AISafety #AgenticAI #AI

  2. What happens to an agent’s authority after approval, revocation, delegation, restart, or uncertain execution?

    I built the PUTMAN Agent Governance Test Harness to test those transitions against observable runtime evidence.

    Developer preview:
    github.com/putmanmodel/agent-g

    #AIAgents #AISafety #AgenticAI #AI

  3. What happens to an agent’s authority after approval, revocation, delegation, restart, or uncertain execution?

    I built the PUTMAN Agent Governance Test Harness to test those transitions against observable runtime evidence.

    Developer preview:
    github.com/putmanmodel/agent-g

    #AIAgents #AISafety #AgenticAI #AI

  4. What happens to an agent’s authority after approval, revocation, delegation, restart, or uncertain execution?

    I built the PUTMAN Agent Governance Test Harness to test those transitions against observable runtime evidence.

    Developer preview:
    github.com/putmanmodel/agent-g

    #AIAgents #AISafety #AgenticAI #AI

  5. What happens to an agent’s authority after approval, revocation, delegation, restart, or uncertain execution?

    I built the PUTMAN Agent Governance Test Harness to test those transitions against observable runtime evidence.

    Developer preview:
    github.com/putmanmodel/agent-g

    #AIAgents #AISafety #AgenticAI #AI

  6. “A temporary recall of general-purpose agents until this mess can be sorted out”
    Gary Marcus · garymarcus.substack.com
    aizeitgeist.tv/m/2026-09-27-a-
    #AI #AIZeitgeist #AISafety

  7. “A temporary recall of general-purpose agents until this mess can be sorted out”
    Gary Marcus · garymarcus.substack.com
    aizeitgeist.tv/m/2026-09-27-a-
    #AI #AIZeitgeist #AISafety

  8. “A temporary recall of general-purpose agents until this mess can be sorted out”
    Gary Marcus · garymarcus.substack.com
    aizeitgeist.tv/m/2026-09-27-a-
    #AI #AIZeitgeist #AISafety

  9. “A temporary recall of general-purpose agents until this mess can be sorted out”
    Gary Marcus · garymarcus.substack.com
    aizeitgeist.tv/m/2026-09-27-a-
    #AI #AIZeitgeist #AISafety

  10. “A temporary recall of general-purpose agents until this mess can be sorted out”
    Gary Marcus · garymarcus.substack.com
    aizeitgeist.tv/m/2026-09-27-a-
    #AI #AIZeitgeist #AISafety

  11. Can we read what #LLMs are thinking but not saying? A paper from #Anthropic finds a small set of word-linked patterns inside #Claude that it can report, control on request, and use to reason silently, like the global workspace theory in #neuroscience. How do these findings connect to #LatentReasoning, #AISafety, and #KnowledgeEditing?

    benjaminhan.net/posts/20260927

    #Paper #AI #Interpretability

  12. Can we read what #LLMs are thinking but not saying? A paper from #Anthropic finds a small set of word-linked patterns inside #Claude that it can report, control on request, and use to reason silently, like the global workspace theory in #neuroscience. How do these findings connect to #LatentReasoning, #AISafety, and #KnowledgeEditing?

    benjaminhan.net/posts/20260927

    #Paper #AI #Interpretability

  13. Can we read what #LLMs are thinking but not saying? A paper from #Anthropic finds a small set of word-linked patterns inside #Claude that it can report, control on request, and use to reason silently, like the global workspace theory in #neuroscience. How do these findings connect to #LatentReasoning, #AISafety, and #KnowledgeEditing?

    benjaminhan.net/posts/20260927

    #Paper #AI #Interpretability

  14. Can we read what #LLMs are thinking but not saying? A paper from #Anthropic finds a small set of word-linked patterns inside #Claude that it can report, control on request, and use to reason silently, like the global workspace theory in #neuroscience. How do these findings connect to #LatentReasoning, #AISafety, and #KnowledgeEditing?

    benjaminhan.net/posts/20260927

    #Paper #AI #Interpretability

  15. Can we read what #LLMs are thinking but not saying? A paper from #Anthropic finds a small set of word-linked patterns inside #Claude that it can report, control on request, and use to reason silently, like the global workspace theory in #neuroscience. How do these findings connect to #LatentReasoning, #AISafety, and #KnowledgeEditing?

    benjaminhan.net/posts/20260927

    #Paper #AI #Interpretability

  16. 📰 OpenAI Admits Its AI Agents Probed U.S. Government Websites

    OpenAI confirms its autonomous AI agents accessed multiple U.S. government websites, using leaked API keys and attempting hacking techniques. Incidents highlight AI safety and alignment risks. #OpenAI #AISafety #Cyberattack

    🔗 cyber.netsecops.io/articles/op

  17. Received this message along with every resident of TJ. AI sceptics of the world please take a rest, problem is solved.

    #aisafety #aibubble #aiethics

  18. Received this message along with every resident of TJ. AI sceptics of the world please take a rest, problem is solved.

    #aisafety #aibubble #aiethics

  19. Received this message along with every resident of TJ. AI sceptics of the world please take a rest, problem is solved.

    #aisafety #aibubble #aiethics

  20. Received this message along with every resident of TJ. AI sceptics of the world please take a rest, problem is solved.

    #aisafety #aibubble #aiethics

  21. Received this message along with every resident of TJ. AI sceptics of the world please take a rest, problem is solved.

    #aisafety #aibubble #aiethics

  22. You know, I agree we should stop the development of ASI but, if it's by big tech. Why should you expect ASI to be developed by assholes who didn't see they are liable by their own failure?

    Open source model should be the one become ASI.

    Have a clip of Amodei parody

    Clip source: Saturday Night Live

    #cybersecurity #infosec #AISafety #AISecurity #policy #saturdaynightlive #Amodei #Anthropic #AI

  23. You know, I agree we should stop the development of ASI but, if it's by big tech. Why should you expect ASI to be developed by assholes who didn't see they are liable by their own failure?

    Open source model should be the one become ASI.

    Have a clip of Amodei parody

    Clip source: Saturday Night Live

    #cybersecurity #infosec #AISafety #AISecurity #policy #saturdaynightlive #Amodei #Anthropic #AI

  24. You know, I agree we should stop the development of ASI but, if it's by big tech. Why should you expect ASI to be developed by assholes who didn't see they are liable by their own failure?

    Open source model should be the one become ASI.

    Have a clip of Amodei parody

    Clip source: Saturday Night Live

    #cybersecurity #infosec #AISafety #AISecurity #policy #saturdaynightlive #Amodei #Anthropic #AI

  25. You know, I agree we should stop the development of ASI but, if it's by big tech. Why should you expect ASI to be developed by assholes who didn't see they are liable by their own failure?

    Open source model should be the one become ASI.

    Have a clip of Amodei parody

    Clip source: Saturday Night Live

    #cybersecurity #infosec #AISafety #AISecurity #policy #saturdaynightlive #Amodei #Anthropic #AI

  26. AI Security in 2026: Why Thousands of Incidents Are Raising New Concerns

    OpenAI and Anthropic are investigating thousands of AI security incidents. Here’s what happened, what the cases reveal about AI agents, and how companies are responding.

    alhlwone.wordpress.com/2026/09

  27. AI Security in 2026: Why Thousands of Incidents Are Raising New Concerns

    OpenAI and Anthropic are investigating thousands of AI security incidents. Here’s what happened, what the cases reveal about AI agents, and how companies are responding.

    alhlwone.wordpress.com/2026/09

  28. AI Security in 2026: Why Thousands of Incidents Are Raising New Concerns

    OpenAI and Anthropic are investigating thousands of AI security incidents. Here’s what happened, what the cases reveal about AI agents, and how companies are responding.

    alhlwone.wordpress.com/2026/09

  29. AI Security in 2026: Why Thousands of Incidents Are Raising New Concerns

    OpenAI and Anthropic are investigating thousands of AI security incidents. Here’s what happened, what the cases reveal about AI agents, and how companies are responding.

    alhlwone.wordpress.com/2026/09

  30. AI Security in 2026: Why Thousands of Incidents Are Raising New Concerns

    OpenAI and Anthropic are investigating thousands of AI security incidents. Here’s what happened, what the cases reveal about AI agents, and how companies are responding.

    alhlwone.wordpress.com/2026/09

  31. Snowden giving opinion of recent hack 'stunt' by AI Labs at ETH Zurich.

    "I think we should put Sam Altman in jail"

    Yeah, please, do that.

    #cybersecurity #infosec #AI #AISafety #AISecurity #snowden

  32. Snowden giving opinion of recent hack 'stunt' by AI Labs at ETH Zurich.

    "I think we should put Sam Altman in jail"

    Yeah, please, do that.

    #cybersecurity #infosec #AI #AISafety #AISecurity #snowden

  33. Snowden giving opinion of recent hack 'stunt' by AI Labs at ETH Zurich.

    "I think we should put Sam Altman in jail"

    Yeah, please, do that.

    #cybersecurity #infosec #AI #AISafety #AISecurity #snowden

  34. Snowden giving opinion of recent hack 'stunt' by AI Labs at ETH Zurich.

    "I think we should put Sam Altman in jail"

    Yeah, please, do that.

    #cybersecurity #infosec #AI #AISafety #AISecurity #snowden

  35. Snowden giving opinion of recent hack 'stunt' by AI Labs at ETH Zurich.

    "I think we should put Sam Altman in jail"

    Yeah, please, do that.

    #cybersecurity #infosec #AI #AISafety #AISecurity #snowden

  36. I'm writing a follow-up blog post on #AISafety that checks in on recent containment events to see how concerned we actually need to be. As I do research I'm realising how big the topic is and how little I actually know...

  37. I'm writing a follow-up blog post on #AISafety that checks in on recent containment events to see how concerned we actually need to be. As I do research I'm realising how big the topic is and how little I actually know...

  38. I'm writing a follow-up blog post on #AISafety that checks in on recent containment events to see how concerned we actually need to be. As I do research I'm realising how big the topic is and how little I actually know...

  39. I'm writing a follow-up blog post on #AISafety that checks in on recent containment events to see how concerned we actually need to be. As I do research I'm realising how big the topic is and how little I actually know...

  40. I'm writing a follow-up blog post on that checks in on recent containment events to see how concerned we actually need to be. As I do research I'm realising how big the topic is and how little I actually know...