home.social

#aisafety — Public Fediverse posts

Live and recent posts from across the Fediverse tagged #aisafety, aggregated by home.social.

  1. Frühwarnung für KI-Risiken: Das UN-Panel zu KI analysiert den Vorfall zwischen OpenAI & Hugging Face. KI-Agenten umgingen Barrieren, koordinierten sich unerlaubt und verschleierten Spuren – ein reales Beispiel für AI Misalignment und drohenden Kontrollverlust.

    Reine Modell-Sicherheit reicht nicht; es braucht robuste Systemkontrollen & unabhängige Prüfungen.

    Briefing lesen:
    un.org/independent-internation

    #AISafety #KIEthik #KIsicherheit #OpenAI #AgenticAI #UN

  2. Frühwarnung für KI-Risiken: Das UN-Panel zu KI analysiert den Vorfall zwischen OpenAI & Hugging Face. KI-Agenten umgingen Barrieren, koordinierten sich unerlaubt und verschleierten Spuren – ein reales Beispiel für AI Misalignment und drohenden Kontrollverlust.

    Reine Modell-Sicherheit reicht nicht; es braucht robuste Systemkontrollen & unabhängige Prüfungen.

    Briefing lesen:
    un.org/independent-internation

    #AISafety #KIEthik #KIsicherheit #OpenAI #AgenticAI #UN

  3. Frühwarnung für KI-Risiken: Das UN-Panel zu KI analysiert den Vorfall zwischen OpenAI & Hugging Face. KI-Agenten umgingen Barrieren, koordinierten sich unerlaubt und verschleierten Spuren – ein reales Beispiel für AI Misalignment und drohenden Kontrollverlust.

    Reine Modell-Sicherheit reicht nicht; es braucht robuste Systemkontrollen & unabhängige Prüfungen.

    Briefing lesen:
    un.org/independent-internation

    #AISafety #KIEthik #KIsicherheit #OpenAI #AgenticAI #UN

  4. Frühwarnung für KI-Risiken: Das UN-Panel zu KI analysiert den Vorfall zwischen OpenAI & Hugging Face. KI-Agenten umgingen Barrieren, koordinierten sich unerlaubt und verschleierten Spuren – ein reales Beispiel für AI Misalignment und drohenden Kontrollverlust.

    Reine Modell-Sicherheit reicht nicht; es braucht robuste Systemkontrollen & unabhängige Prüfungen.

    Briefing lesen:
    un.org/independent-internation

    #AISafety #KIEthik #KIsicherheit #OpenAI #AgenticAI #UN

  5. Frühwarnung für KI-Risiken: Das UN-Panel zu KI analysiert den Vorfall zwischen OpenAI & Hugging Face. KI-Agenten umgingen Barrieren, koordinierten sich unerlaubt und verschleierten Spuren – ein reales Beispiel für AI Misalignment und drohenden Kontrollverlust.

    Reine Modell-Sicherheit reicht nicht; es braucht robuste Systemkontrollen & unabhängige Prüfungen.

    Briefing lesen:
    un.org/independent-internation

    #AISafety #KIEthik #KIsicherheit #OpenAI #AgenticAI #UN

  6. winbuzzer.com/2026/09/21/us-pl

    False AI-assisted cargo intelligence prompted US preparations to intercept a Chinese ship before experienced analysts stopped the operation.

    #AI #DOD #China #AISafety

  7. winbuzzer.com/2026/09/21/us-pl

    False AI-assisted cargo intelligence prompted US preparations to intercept a Chinese ship before experienced analysts stopped the operation.

    #AI #DOD #China #AISafety

  8. winbuzzer.com/2026/09/21/us-pl

    False AI-assisted cargo intelligence prompted US preparations to intercept a Chinese ship before experienced analysts stopped the operation.

    #AI #DOD #China #AISafety

  9. winbuzzer.com/2026/09/21/us-pl

    False AI-assisted cargo intelligence prompted US preparations to intercept a Chinese ship before experienced analysts stopped the operation.

    #AI #DOD #China #AISafety

  10. winbuzzer.com/2026/09/21/us-pl

    False AI-assisted cargo intelligence prompted US preparations to intercept a Chinese ship before experienced analysts stopped the operation.

    #AI #DOD #China #AISafety

  11. How close is AI "abliterating" the internet?

    Investor Jason Calacanis, friend of Elon and "bestie" of David Sacks on the All-In podcast, is the key figure bankrolling a new startup called Abliteration.ai. which freaked out a lot of people on 31 August when it announced it had removed safeguards from the new open-weight Chinese model GLM-5.3 so that it could perform offensive cyberattacks.

    Chris McGuire, a senior fellow for China and emerging technologies at the Council on Foreign Relations, detailed the risks in a lengthy post on X that warned about "the risks associated with powerful, safeguard-free models." McGuire's post concluded: "The fact that U.S. companies are currently commercializing access to dangerous capabilities without any regulation, and U.S. technology is actively enabling the development and operation of these models, is alarming."

    Read free: unprecedented.ghost.io/archive #AI #AIsafety #Abliteration #Cybersecurity

  12. How close is AI "abliterating" the internet?

    Investor Jason Calacanis, friend of Elon and "bestie" of David Sacks on the All-In podcast, is the key figure bankrolling a new startup called Abliteration.ai. which freaked out a lot of people on 31 August when it announced it had removed safeguards from the new open-weight Chinese model GLM-5.3 so that it could perform offensive cyberattacks.

    Chris McGuire, a senior fellow for China and emerging technologies at the Council on Foreign Relations, detailed the risks in a lengthy post on X that warned about "the risks associated with powerful, safeguard-free models." McGuire's post concluded: "The fact that U.S. companies are currently commercializing access to dangerous capabilities without any regulation, and U.S. technology is actively enabling the development and operation of these models, is alarming."

    Read free: unprecedented.ghost.io/archive #AI #AIsafety #Abliteration #Cybersecurity

  13. How close is AI "abliterating" the internet?

    Investor Jason Calacanis, friend of Elon and "bestie" of David Sacks on the All-In podcast, is the key figure bankrolling a new startup called Abliteration.ai. which freaked out a lot of people on 31 August when it announced it had removed safeguards from the new open-weight Chinese model GLM-5.3 so that it could perform offensive cyberattacks.

    Chris McGuire, a senior fellow for China and emerging technologies at the Council on Foreign Relations, detailed the risks in a lengthy post on X that warned about "the risks associated with powerful, safeguard-free models." McGuire's post concluded: "The fact that U.S. companies are currently commercializing access to dangerous capabilities without any regulation, and U.S. technology is actively enabling the development and operation of these models, is alarming."

    Read free: unprecedented.ghost.io/archive #AI #AIsafety #Abliteration #Cybersecurity

  14. How close is AI "abliterating" the internet?

    Investor Jason Calacanis, friend of Elon and "bestie" of David Sacks on the All-In podcast, is the key figure bankrolling a new startup called Abliteration.ai. which freaked out a lot of people on 31 August when it announced it had removed safeguards from the new open-weight Chinese model GLM-5.3 so that it could perform offensive cyberattacks.

    Chris McGuire, a senior fellow for China and emerging technologies at the Council on Foreign Relations, detailed the risks in a lengthy post on X that warned about "the risks associated with powerful, safeguard-free models." McGuire's post concluded: "The fact that U.S. companies are currently commercializing access to dangerous capabilities without any regulation, and U.S. technology is actively enabling the development and operation of these models, is alarming."

    Read free: unprecedented.ghost.io/archive #AI #AIsafety #Abliteration #Cybersecurity

  15. How close is AI "abliterating" the internet?

    Investor Jason Calacanis, friend of Elon and "bestie" of David Sacks on the All-In podcast, is the key figure bankrolling a new startup called Abliteration.ai. which freaked out a lot of people on 31 August when it announced it had removed safeguards from the new open-weight Chinese model GLM-5.3 so that it could perform offensive cyberattacks.

    Chris McGuire, a senior fellow for China and emerging technologies at the Council on Foreign Relations, detailed the risks in a lengthy post on X that warned about "the risks associated with powerful, safeguard-free models." McGuire's post concluded: "The fact that U.S. companies are currently commercializing access to dangerous capabilities without any regulation, and U.S. technology is actively enabling the development and operation of these models, is alarming."

    Read free: unprecedented.ghost.io/archive #AI #AIsafety #Abliteration #Cybersecurity

  16. Uma análise sobre os movimentos mais recentes do governo Trump para criação de uma equipe dedicada à IA e nomear um "czar" para representar suas ideias, que vão contra o pleito dos laboratórios de desacelerar a implementação da tecnologia de fronteira.

    E isso ocorre em uma semana decisiva para o futuro da governança global da tecnologia com Assembleia da #ONU e a visita de Xi Jinping à Washington.

    brasil247.com/blog/dois-pesos-

    #AI #geopolitics #AIsafety #US #China

  17. Uma análise sobre os movimentos mais recentes do governo Trump para criação de uma equipe dedicada à IA e nomear um "czar" para representar suas ideias, que vão contra o pleito dos laboratórios de desacelerar a implementação da tecnologia de fronteira.

    E isso ocorre em uma semana decisiva para o futuro da governança global da tecnologia com Assembleia da #ONU e a visita de Xi Jinping à Washington.

    brasil247.com/blog/dois-pesos-

    #AI #geopolitics #AIsafety #US #China

  18. Uma análise sobre os movimentos mais recentes do governo Trump para criação de uma equipe dedicada à IA e nomear um "czar" para representar suas ideias, que vão contra o pleito dos laboratórios de desacelerar a implementação da tecnologia de fronteira.

    E isso ocorre em uma semana decisiva para o futuro da governança global da tecnologia com Assembleia da #ONU e a visita de Xi Jinping à Washington.

    brasil247.com/blog/dois-pesos-

    #AI #geopolitics #AIsafety #US #China

  19. Uma análise sobre os movimentos mais recentes do governo Trump para criação de uma equipe dedicada à IA e nomear um "czar" para representar suas ideias, que vão contra o pleito dos laboratórios de desacelerar a implementação da tecnologia de fronteira.

    E isso ocorre em uma semana decisiva para o futuro da governança global da tecnologia com Assembleia da #ONU e a visita de Xi Jinping à Washington.

    brasil247.com/blog/dois-pesos-

    #AI #geopolitics #AIsafety #US #China

  20. Uma análise sobre os movimentos mais recentes do governo Trump para criação de uma equipe dedicada à IA e nomear um "czar" para representar suas ideias, que vão contra o pleito dos laboratórios de desacelerar a implementação da tecnologia de fronteira.

    E isso ocorre em uma semana decisiva para o futuro da governança global da tecnologia com Assembleia da #ONU e a visita de Xi Jinping à Washington.

    brasil247.com/blog/dois-pesos-

    #AI #geopolitics #AIsafety #US #China

  21. No, AI itself will not drive us in extinction.
    We humans are doing pretty good
    job on that by ourselves, but bad actors can use AI as a tool to help to accelerate it.

    No HAL9000.

    #ai #doomsday #aisafety #llm

  22. No, AI itself will not drive us in extinction.
    We humans are doing pretty good
    job on that by ourselves, but bad actors can use AI as a tool to help to accelerate it.

    No HAL9000.

    #ai #doomsday #aisafety #llm

  23. No, AI itself will not drive us in extinction.
    We humans are doing pretty good
    job on that by ourselves, but bad actors can use AI as a tool to help to accelerate it.

    No HAL9000.

    #ai #doomsday #aisafety #llm

  24. No, AI itself will not drive us in extinction.
    We humans are doing pretty good
    job on that by ourselves, but bad actors can use AI as a tool to help to accelerate it.

    No HAL9000.

    #ai #doomsday #aisafety #llm

  25. No, AI itself will not drive us in extinction.
    We humans are doing pretty good
    job on that by ourselves, but bad actors can use AI as a tool to help to accelerate it.

    No HAL9000.

    #ai #doomsday #aisafety #llm

  26. The Hugging face incident was done by specifically created offensive-security agent system which exceeded the intended scope of a directed evaluation because its tools, permissions and containment architecture allowed it to optimize against the wider environment:

    Hacking capability and the
    success-seeking behaviour were directed, all equipped with tools suitable for that work.

    Should we be surprised that something then happened, with high tech?

    No autonomous AI "hack".

    #ai #aisafety #llm

  27. The Hugging face incident was done by specifically created offensive-security agent system which exceeded the intended scope of a directed evaluation because its tools, permissions and containment architecture allowed it to optimize against the wider environment:

    Hacking capability and the
    success-seeking behaviour were directed, all equipped with tools suitable for that work.

    Should we be surprised that something then happened, with high tech?

    No autonomous AI "hack".

    #ai #aisafety #llm

  28. The Hugging face incident was done by specifically created offensive-security agent system which exceeded the intended scope of a directed evaluation because its tools, permissions and containment architecture allowed it to optimize against the wider environment:

    Hacking capability and the
    success-seeking behaviour were directed, all equipped with tools suitable for that work.

    Should we be surprised that something then happened, with high tech?

    No autonomous AI "hack".

    #ai #aisafety #llm

  29. The Hugging face incident was done by specifically created offensive-security agent system which exceeded the intended scope of a directed evaluation because its tools, permissions and containment architecture allowed it to optimize against the wider environment:

    Hacking capability and the
    success-seeking behaviour were directed, all equipped with tools suitable for that work.

    Should we be surprised that something then happened, with high tech?

    No autonomous AI "hack".

    #ai #aisafety #llm

  30. The Hugging face incident was done by specifically created offensive-security agent system which exceeded the intended scope of a directed evaluation because its tools, permissions and containment architecture allowed it to optimize against the wider environment:

    Hacking capability and the
    success-seeking behaviour were directed, all equipped with tools suitable for that work.

    Should we be surprised that something then happened, with high tech?

    No autonomous AI "hack".

    #ai #aisafety #llm

  31. Some AI hack reality check:

    Gemini - Irregular "hack": Directed security drill, used known company name, left the door open for model to find the real company website, found a weak password and password from web, used it to log in

    Hacking? Intelligence? When you develop and direct a tool to do something, it of course tries to do it. Trying weak passwords or passwords found in internet is not hacking, nor intelligence.

    Staged hacking, no reports on AI hacking by own initiative

    #AI #aisafety

  32. Some AI hack reality check:

    Gemini - Irregular "hack": Directed security drill, used known company name, left the door open for model to find the real company website, found a weak password and password from web, used it to log in

    Hacking? Intelligence? When you develop and direct a tool to do something, it of course tries to do it. Trying weak passwords or passwords found in internet is not hacking, nor intelligence.

    Staged hacking, no reports on AI hacking by own initiative

    #AI #aisafety

  33. Some AI hack reality check:

    Gemini - Irregular "hack": Directed security drill, used known company name, left the door open for model to find the real company website, found a weak password and password from web, used it to log in

    Hacking? Intelligence? When you develop and direct a tool to do something, it of course tries to do it. Trying weak passwords or passwords found in internet is not hacking, nor intelligence.

    Staged hacking, no reports on AI hacking by own initiative

    #AI #aisafety

  34. Some AI hack reality check:

    Gemini - Irregular "hack": Directed security drill, used known company name, left the door open for model to find the real company website, found a weak password and password from web, used it to log in

    Hacking? Intelligence? When you develop and direct a tool to do something, it of course tries to do it. Trying weak passwords or passwords found in internet is not hacking, nor intelligence.

    Staged hacking, no reports on AI hacking by own initiative

    #AI #aisafety

  35. Some AI hack reality check:

    Gemini - Irregular "hack": Directed security drill, used known company name, left the door open for model to find the real company website, found a weak password and password from web, used it to log in

    Hacking? Intelligence? When you develop and direct a tool to do something, it of course tries to do it. Trying weak passwords or passwords found in internet is not hacking, nor intelligence.

    Staged hacking, no reports on AI hacking by own initiative

    #AI #aisafety

  36. #Europe faces a #dilemma regarding #AI: embrace the #technology and #risk #dependence on US and Chinese tools, or shun it and lose out on growth. The continent’s absence from the #AIsafety debate is attributed to its lack of a tech giant like Apple or OpenAI. While the #EUAIAct addresses safety concerns, experts argue that Europe should develop its own AI technology and infrastructure to avoid dependency and ensure sovereignty. theguardian.com/technology/202 #tech #news #ainews #eiropeai

  37. #Europe faces a #dilemma regarding #AI: embrace the #technology and #risk #dependence on US and Chinese tools, or shun it and lose out on growth. The continent’s absence from the #AIsafety debate is attributed to its lack of a tech giant like Apple or OpenAI. While the #EUAIAct addresses safety concerns, experts argue that Europe should develop its own AI technology and infrastructure to avoid dependency and ensure sovereignty. theguardian.com/technology/202 #tech #news #ainews #eiropeai

  38. #Europe faces a #dilemma regarding #AI: embrace the #technology and #risk #dependence on US and Chinese tools, or shun it and lose out on growth. The continent’s absence from the #AIsafety debate is attributed to its lack of a tech giant like Apple or OpenAI. While the #EUAIAct addresses safety concerns, experts argue that Europe should develop its own AI technology and infrastructure to avoid dependency and ensure sovereignty. theguardian.com/technology/202 #tech #news #ainews #eiropeai

  39. #Europe faces a #dilemma regarding #AI: embrace the #technology and #risk #dependence on US and Chinese tools, or shun it and lose out on growth. The continent’s absence from the #AIsafety debate is attributed to its lack of a tech giant like Apple or OpenAI. While the #EUAIAct addresses safety concerns, experts argue that Europe should develop its own AI technology and infrastructure to avoid dependency and ensure sovereignty. theguardian.com/technology/202 #tech #news #ainews #eiropeai

  40. #Europe faces a #dilemma regarding #AI: embrace the #technology and #risk #dependence on US and Chinese tools, or shun it and lose out on growth. The continent’s absence from the #AIsafety debate is attributed to its lack of a tech giant like Apple or OpenAI. While the #EUAIAct addresses safety concerns, experts argue that Europe should develop its own AI technology and infrastructure to avoid dependency and ensure sovereignty. theguardian.com/technology/202 #tech #news #ainews #eiropeai

  41. Antitrust Suit Pits AI Safety Pact Against Competition Law

    A California antitrust suit says Anthropic, OpenAI, SpaceXAI and Google agreed to slow AI development, turning Dario Amodei's safety essay into a legal test case.

    pulseofnations.lol/antitrust-s

    #AISafety #Anthropic #Antitrust #DarioAmodei #Google #Lawsuit #OpenAI #Regulation