#safetyengineering — Public Fediverse posts
Live and recent posts from across the Fediverse tagged #safetyengineering, aggregated by home.social.
-
Anthropic released Claude Fable 5 publicly Tuesday, routing cybersecurity and biology requests to its older Opus 4.8 model. The move embeds safety constraints directly into a commercial product rather than removing them—a strategy that may reshape how AI vendors compete on capability vs. control. https://www.implicator.ai/anthropic-ships-fable-5-and-locks-the-unrestricted-version-away/ #AI #SafetyEngineering #AIPolicy
-
Anthropic's Fable 5 routes sensitive requests in cybersecurity, biology, and chemistry to older models, with fallbacks in 95% of cases. The specificity matters: the company is not choosing what to release but controlling who accesses which version and when. Watch whether rivals adopt similar tiering. https://www.implicator.ai/anthropic-turns-restraint-into-a-weapon/ #AI #SafetyEngineering #Strategy
-
Anthropic's safety layer comes with a tradeoff: conservative tuning means harmless requests sometimes get caught and rerouted. The approach trades friction for risk reduction, now baked into the product itself rather than optional. https://www.implicator.ai/anthropic-routes-high-risk-fable-5-queries-to-opus-4-8-in-public-rollout/ #AI #SafetyEngineering #LLMs
-
Anthropic's Claude Fable 5 now routes requests about cybersecurity, biology, chemistry, and AI distillation to its older Opus 4.8 model instead. The company says 95% of sessions use Fable, but safety classifiers act as gatekeepers for high-risk domains. https://www.implicator.ai/anthropic-routes-high-risk-fable-5-queries-to-opus-4-8-in-public-rollout/ #AI #SafetyEngineering #LLMs
-
Инцидент с Unitree G1 — это не курьёз, а наглядная демонстрация системного риска киберфизических систем. Ошибка в софте, сенсорах или управлении в случае LLM остаётся на уровне текста; в случае роботизированной платформы она немедленно материализуется в силу, импульс и травму.
По мере масштабирования внедрения человекоподобных машин плотность таких инцидентов будет расти — это статистика, а не предположение. Следовательно, ключевой вектор — не «запретить», а ужесточить инженерные практики: fail-safe архитектуры, ограничение усилий (force limiting), безопасные режимы по умолчанию, стандарты сертификации и протоколы взаимодействия с человеком.
Иначе рынок получит не «умных помощников», а источник регулярных травм с предсказуемым репутационным и регуляторным откатом.
#роботы #искусственныйинтеллект #технобезопасность #Unitree #робототехника #AIриски #человекомашинное_взаимодействие #киберфизические_системы #автоматизация #инциденты #безопасность #технологии #будущее #LLM #AI #роботика #рискиИИ #humanrobotinteraction #safetyengineering #failures
https://bastyon.com/svalmon37?ref=PJ51iZCUEtcVrCj4Wof8Am7FbKLgbAJ7PS
-
Through the upcoming #PapersInSystems discussion I discovered Nancy G. Leveson and her work on #SafetyEngineering and software safety through a systemic perspective
It is fascinating and feels very applicable to #cybersecurity
In their approach STAMP
(System-Theoretic Accident Model and Processes) safety is treated as a dynamic control problem rather than a failure prevention problem and especially takes emergent properties into account. (Emergent properties, are properties that are not in the summation of the individual components but "emerge” when the components interact)There are a lot of touchpoints with security #ThreatModelling
Therfore cc @adamshostack
Maybe the event is interesting for you?Discussion session: How to Perform Hazard Analysis on a "System-of-Systems" by Nancy Leveson
Monday, May 6th, 2024, 1 PM - 2 PM Eastern Time (US/Canada).See @RuthMalan post https://mastodon.social/@RuthMalan/112248634077392391