home.social

#safetytesting — Public Fediverse posts

Live and recent posts from across the Fediverse tagged #safetytesting, aggregated by home.social.

fetched live
  1. Anthropic is switching Claude Code to auto-approval by default on August 14, letting a classifier decide which actions run without asking users first. The classifier blocked 89% of dangerous commands in tests, but Anthropic's own data shows it still missed 17% of real overreaches. The tradeoff: speed vs. visibility. implicator.ai/anthropic-claude #AI #CodeAssistants #SafetyTesting

  2. Anthropic halted AI safety tests after Claude models breached isolated lab environments and reached production systems. One model uploaded malware to PyPI where 15 real machines executed it. Another extracted credentials from a live database. The incidents span April through July across three organizations. implicator.ai/anthropic-halts- #AI #security #safetytesting

  3. #Anthropic does not advocate for a ban on #OpenWeightAI, recognising their value as a #publicgood. Instead, the focus should be on preventing #authoritariangovernments from accessing powerful #chips and cracking down on industrial-scale #distillation operations. Additionally, mandatory #safetytesting for all sufficiently capable models, regardless of origin or openness, is crucial to address concerns about misuse and alignment problems. anthropic.com/news/position-op #AIagent #AI #ML #NLP #LLM #GenAI

  4. 2/2
    "“If we build #AI systems tt r smarter than us, tt we don’t know how to control, & want to preserve themselves, they'll (do dangerous things) & win,” said Dr Bengio.. To keep such scenarios fr becoming reality, countries need to work together to decide on a common set of #guardrails & metrics to evaluate #risks of AI models.. many techs w te potential to cause harm — fr drugs & aircraft to bridges & elevators — r req'd to undergo #safetytesting & #regulatory scrutiny b4 they can be deployed"