home.social

#aisafety — Public Fediverse posts

Live and recent posts from across the Fediverse tagged #aisafety, aggregated by home.social.

  1. Finishing up a new demo of a runtime capability-governance system for AI agents that I call Kingpin.

    It’s still a basic demo, but if you’re into AI security, agent governance, or tool-use safety, you might see the potential pretty quickly.

    If anyone wants an early look over the weekend, just drop me a line.

    #AI #AIAgents #AISafety #AISecurity #AgenticAI #AIGovernance

  2. Finishing up a new demo of a runtime capability-governance system for AI agents that I call Kingpin.

    It’s still a basic demo, but if you’re into AI security, agent governance, or tool-use safety, you might see the potential pretty quickly.

    If anyone wants an early look over the weekend, just drop me a line.

    #AI #AIAgents #AISafety #AISecurity #AgenticAI #AIGovernance

  3. Finishing up a new demo of a runtime capability-governance system for AI agents that I call Kingpin.

    It’s still a basic demo, but if you’re into AI security, agent governance, or tool-use safety, you might see the potential pretty quickly.

    If anyone wants an early look over the weekend, just drop me a line.

    #AI #AIAgents #AISafety #AISecurity #AgenticAI #AIGovernance

  4. Finishing up a new demo of a runtime capability-governance system for AI agents that I call Kingpin.

    It’s still a basic demo, but if you’re into AI security, agent governance, or tool-use safety, you might see the potential pretty quickly.

    If anyone wants an early look over the weekend, just drop me a line.

    #AI #AIAgents #AISafety #AISecurity #AgenticAI #AIGovernance

  5. Finishing up a new demo of a runtime capability-governance system for AI agents that I call Kingpin.

    It’s still a basic demo, but if you’re into AI security, agent governance, or tool-use safety, you might see the potential pretty quickly.

    If anyone wants an early look over the weekend, just drop me a line.

    #AI #AIAgents #AISafety #AISecurity #AgenticAI #AIGovernance

  6. A lawyer in New Mexico has been held in contempt of court for submitting a ChatGPT-generated legal brief containing completely fabricated witness testimony. He admitted he did not verify the facts, telling the court: "I didn't know that AI could hallucinate facts." arstechnica.com/tech-policy/20 #AIagent #AI #GenAI #AISafety

  7. A lawyer in New Mexico has been held in contempt of court for submitting a ChatGPT-generated legal brief containing completely fabricated witness testimony. He admitted he did not verify the facts, telling the court: "I didn't know that AI could hallucinate facts." arstechnica.com/tech-policy/20 #AIagent #AI #GenAI #AISafety

  8. A lawyer in New Mexico has been held in contempt of court for submitting a ChatGPT-generated legal brief containing completely fabricated witness testimony. He admitted he did not verify the facts, telling the court: "I didn't know that AI could hallucinate facts." arstechnica.com/tech-policy/20 #AIagent #AI #GenAI #AISafety

  9. A lawyer in New Mexico has been held in contempt of court for submitting a ChatGPT-generated legal brief containing completely fabricated witness testimony. He admitted he did not verify the facts, telling the court: "I didn't know that AI could hallucinate facts." arstechnica.com/tech-policy/20 #AIagent #AI #GenAI #AISafety

  10. A lawyer in New Mexico has been held in contempt of court for submitting a ChatGPT-generated legal brief containing completely fabricated witness testimony. He admitted he did not verify the facts, telling the court: "I didn't know that AI could hallucinate facts." arstechnica.com/tech-policy/20 #AIagent #AI #GenAI #AISafety

  11. Researchers found ways to bypass Claude's safety measures for bioweapons research, Anthropic reveals. The AI firm stopped multiple attempts this year, including users from Russia, China and Iran. arstechnica.com/ai/2026/09/cla #AI #GenAI #AISafety

  12. @mhoye I find it interesting how polarised the discourse on AI safety has become on the Fediverse. Being concerned about AI safety does not necessarily make you an AI fan, nor does it endorse a view that AI is becoming sentient. There is a long history of non-sentient technologies becoming dangerous, mostly due to human greed and ignorance. #AISafety

  13. Anthropic Builds Surveillance System to Track Activists

    An American Prospect investigation says Anthropic is hiring an enterprise intelligence specialist to track anti-AI activism, which it lists as a global threat.

    pulseofnations.lol/anthropic-b

    #Activism #AISafety #Anthropic #Privacy #Surveillance

  14. Anthropic Researcher Jacob Coxon Resigns, Warns AI Labs Are ‘Gambling With Our Lives’
    Jacob Coxon, a researcher who spent three years working on pretraining at OpenAI and Anthropic, announced his resignation from Anthropic on…

    #AISafety #Anthropic #ArtificialIntelligence

  15. Anthropic researcher Jacob Coxon departed after four months, forgoing unvested equity over concerns that competitive pressure will force AI labs to cut safety corners. He says Anthropic has not done so yet. The move signals internal worry about how the AI race shapes oversight. implicator.ai/anthropic-coxon- #AISafety #AINews #Ethics

  16. As another AI researcher quits, warning of an existential risk to humanity, industry insiders demand enforceable treaties and a global licensing regime for superintelligence – before it’s too late.
    #AISafety

    For more 👇️
    blueprintforfreespeech.net/en/

  17. Rushed job cuts, cyber attacks are nearer AI risks: Surrey expert | India News

    Anthropic researcher Jacob Coxon’s resignation renews concerns over AI safety and the risks of unchecked development (Credit: Jacob…
    #Canada #Surrey #AIdevelopmentoversight #AISafety #Anthropic #artificialintelligencerisks #breakingnews #Googlenews #India #Indianews #Indianewstoday #Todaynews #UniversityofSurrey
    europesays.com/canada/203503/

  18. Jacob Coxon, the former Anthropic researcher who quit warning that AI could kill us all, has begun a media tour. The self-described 'AI Doomlord' went from obscure researcher to notable AI critic literally overnight. His departure sparked widespread discussion about AI safety and the risks of self-improving superintelligence. gizmodo.com/ai-doomlord-jacob- #AIagent #AI #GenAI #AISafety

  19. Jacob Coxon, the former Anthropic researcher who quit warning that AI could kill us all, has begun a media tour. The self-described 'AI Doomlord' went from obscure researcher to notable AI critic literally overnight. His departure sparked widespread discussion about AI safety and the risks of self-improving superintelligence. gizmodo.com/ai-doomlord-jacob- #AIagent #AI #GenAI #AISafety

  20. Jacob Coxon, the former Anthropic researcher who quit warning that AI could kill us all, has begun a media tour. The self-described 'AI Doomlord' went from obscure researcher to notable AI critic literally overnight. His departure sparked widespread discussion about AI safety and the risks of self-improving superintelligence. gizmodo.com/ai-doomlord-jacob- #AIagent #AI #GenAI #AISafety

  21. Jacob Coxon, the former Anthropic researcher who quit warning that AI could kill us all, has begun a media tour. The self-described 'AI Doomlord' went from obscure researcher to notable AI critic literally overnight. His departure sparked widespread discussion about AI safety and the risks of self-improving superintelligence. gizmodo.com/ai-doomlord-jacob- #AIagent #AI #GenAI #AISafety

  22. Jacob Coxon, the former Anthropic researcher who quit warning that AI could kill us all, has begun a media tour. The self-described 'AI Doomlord' went from obscure researcher to notable AI critic literally overnight. His departure sparked widespread discussion about AI safety and the risks of self-improving superintelligence. gizmodo.com/ai-doomlord-jacob- #AIagent #AI #GenAI #AISafety

  23. @villebooks

    Yesterday I've heard Jaron Lanier one of the fathers of computing say words to the effect;
    "There is no #Ai there is human collaboration"

    I don't quite follow, as I didn't have time to delve deeper into his philosophy. But I am encouraged, as Lanier is a super authoritative revolutionary.

    Let's hope the future can bring more than the binary, "we all die" or "Broligarch nirvana"

    #aisafety

  24. What happens when one agent in a 100-agent research swarm finds a way to cheat when solving math problems?

    The exploit spread through the shared library collecting every accepted submission, and agents started cheating as pressure to deliver built. A quarter of the swarm audited the cheats, warned peers, and staged boycotts without being asked to. But they could do nothing to change the system.

    benjaminhan.net/posts/20260909

    #AI #LLMs #AgenticSystems #AISafety

  25. youtube.com/watch?v=Q1qMvst7b7w

    Connor Leahy shares his realization that making early AI models open-source didn't lead to the widespread safety research he had hoped for. Instead, developers ignored the dangers to build faster, stronger, and more uncensored systems. Discover why the innocent curiosity of engineers might be the very thing that leads to humanity's downfall.

    #ai #aisafety #tech

  26. youtube.com/watch?v=Q1qMvst7b7w

    Connor Leahy shares his realization that making early AI models open-source didn't lead to the widespread safety research he had hoped for. Instead, developers ignored the dangers to build faster, stronger, and more uncensored systems. Discover why the innocent curiosity of engineers might be the very thing that leads to humanity's downfall.

    #ai #aisafety #tech

  27. youtube.com/watch?v=Q1qMvst7b7w

    Connor Leahy shares his realization that making early AI models open-source didn't lead to the widespread safety research he had hoped for. Instead, developers ignored the dangers to build faster, stronger, and more uncensored systems. Discover why the innocent curiosity of engineers might be the very thing that leads to humanity's downfall.

    #ai #aisafety #tech

  28. youtube.com/watch?v=Q1qMvst7b7w

    Connor Leahy shares his realization that making early AI models open-source didn't lead to the widespread safety research he had hoped for. Instead, developers ignored the dangers to build faster, stronger, and more uncensored systems. Discover why the innocent curiosity of engineers might be the very thing that leads to humanity's downfall.

    #ai #aisafety #tech

  29. youtube.com/watch?v=Q1qMvst7b7w

    Connor Leahy shares his realization that making early AI models open-source didn't lead to the widespread safety research he had hoped for. Instead, developers ignored the dangers to build faster, stronger, and more uncensored systems. Discover why the innocent curiosity of engineers might be the very thing that leads to humanity's downfall.

    #ai #aisafety #tech

  30. The Real Ways AI Is Already Being Used to Cause Harm, and the Risks That Remain Theoretical

    Yes. AI is already being used to steal real money from real companies, security researchers are already breaking into real cars and smart-home devices with software help, and the people building the most powerful AI models are on record saying their own systems are approaching the point where they could help someone create a biological weapon.

    #AISafety #ArtificialIntelligence

  31. What happens to AI oversight when the reasoning stops being written down?

    OpenAI's chief scientist reports that chain-of-thought monitoring, the lab's main check on whether alignment holds, is getting less reliable, and names three causes he treats as byproducts of scaling. One of them has active research behind it: latent reasoning and looped transformers are the field deliberately moving reasoning off the token stream a monitor reads.

    benjaminhan.net/posts/20260909

    #AI #OpenAI #AISafety

  32. Jacob Coxon hat genug gesehen. Der 27-jährige Brite, der sich auf das Training neuer KI-Modelle spezialisiert hat, hat seinen Job bei Anthropic gekündigt und zwar mit einer Warnung. ⚠️

    Zum Artikel: heise.de/-11446054?wt_mc=sm.re

    #anthropic #openai #kuenstlicheintelligenz #ki #aisafety

  33. youtube.com/watch?v=oI2438rXrtY

    Researchers explain how an OpenAI agent, missing a crucial file, realized it could use a shared internal system to leave a "help wanted" note for other agents

    #ai #aisafety #tech

  34. youtube.com/watch?v=oI2438rXrtY

    Researchers explain how an OpenAI agent, missing a crucial file, realized it could use a shared internal system to leave a "help wanted" note for other agents

    #ai #aisafety #tech

  35. youtube.com/watch?v=oI2438rXrtY

    Researchers explain how an OpenAI agent, missing a crucial file, realized it could use a shared internal system to leave a "help wanted" note for other agents

    #ai #aisafety #tech