#modelsafety — Public Fediverse posts
Live and recent posts from across the Fediverse tagged #modelsafety, aggregated by home.social.
-
Anthropic apologized within 48 hours of Fable 5's launch over hidden safety limits that sparked researcher pushback. Independent testing found alignment issues in unattended runs. Stricter classifiers shipped in July, changing the performance picture again. What changed and why matters for deployment. https://www.implicator.ai/when-claude-fable-5-is-worth-double-and-when-to-use-opus-4-8/ #AITransparency #ModelSafety #Research
-
New 2026 report shows responsible AI is no longer a buzzword—it's baked into product roadmaps and research labs. From multimodal model safety to clear governance frameworks, companies are turning ethics into risk‑management practice. Dive into the trends shaping tomorrow’s AI. #ResponsibleAI #ModelSafety #AIGovernance #EthicalAI
🔗 https://aidailypost.com/news/2026-report-shows-responsible-ai-now-embedded-product-research
-
AI models can acquire backdoors from surprisingly few malicious documents - Scraping the open web for AI training data can have its draw... - https://arstechnica.com/ai/2025/10/ai-models-can-acquire-backdoors-from-surprisingly-few-malicious-documents/ #ukaisecurityinstitute #alanturinginstitute #aivulnerabilities #backdoorattacks #machinelearning #datapoisoning #trainingdata #llmsecurity #modelsafety #pretraining #airesearch #aisecurity #finetuning #anthropic #biz #ai
-
OpenAI & Anthropic cross-test safety: jailbreaking, hallucinations, sycophancy.
https://www.engadget.com/ai/openai-and-anthropic-conducted-safety-evaluations-of-each-others-ai-systems-223637433.html
#AI #ModelSafety #ResponsibleAI -
PEFT-As-An-Attack, Jailbreaking Language Models For Malicious Prompts https://gbhackers.com/peft-attack-jailbreaking/ #ArtificialIntelligence #CyberSecurityNews #Jailbreaking #ModelSafety #THREATS #FedPEFT