home.social

#guardrails — Public Fediverse posts

Live and recent posts from across the Fediverse tagged #guardrails, aggregated by home.social.

  1. The danger of #AI models is not that they will wipe out humanity in the next 18 months. It's that repeating this uncritically as well as making fun of it will "flood the field," and tire people of the underlying message. This will make people indifferent when #guardrails and #regulation are brought up, allowing slop-peddlers like #OpenAI and #Anthropic to press on without #oversight.

    We have seen this in action: #environmental alarmists have overplayed global warming with more or less substantiated concerns preaching the end of civilization "any day now" since the 1970s, leading to legitimate concerns being drowned out and a sneaking crisis emerging over decades with consequences that are going to play out over the coming decades, rather than the more headline-sexy "total annihilation in 18 months unless we take drastic (really symbolic) action nowNoWNOW!!!"

    AI issues are not going to kill humanity before the end of the decade. But the technology has legitimate #concerns – the immediate one being environmental impact of data centers but also threats on a societal and potentially existential levels – that should not be drowned out by making fun of the #doomsayers, whether this is weird #ReversePsychology advertising or not.
  2. The danger of #AI models is not that they will wipe out humanity in the next 18 months. It's that repeating this uncritically as well as making fun of it will "flood the field," and tire people of the underlying message. This will make people indifferent when #guardrails and #regulation are brought up, allowing slop-peddlers like #OpenAI and #Anthropic to press on without #oversight.

    We have seen this in action: #environmental alarmists have overplayed global warming with more or less substantiated concerns preaching the end of civilization "any day now" since the 1970s, leading to legitimate concerns being drowned out and a sneaking crisis emerging over decades with consequences that are going to play out over the coming decades, rather than the more headline-sexy "total annihilation in 18 months unless we take drastic (really symbolic) action nowNoWNOW!!!"

    AI issues are not going to kill humanity before the end of the decade. But the technology has legitimate #concerns – the immediate one being environmental impact of data centers but also threats on a societal and potentially existential levels – that should not be drowned out by making fun of the #doomsayers, whether this is weird #ReversePsychology advertising or not.
  3. The danger of #AI models is not that they will wipe out humanity in the next 18 months. It's that repeating this uncritically as well as making fun of it will "flood the field," and tire people of the underlying message. This will make people indifferent when #guardrails and #regulation are brought up, allowing slop-peddlers like #OpenAI and #Anthropic to press on without #oversight.

    We have seen this in action: #environmental alarmists have overplayed global warming with more or less substantiated concerns preaching the end of civilization "any day now" since the 1970s, leading to legitimate concerns being drowned out and a sneaking crisis emerging over decades with consequences that are going to play out over the coming decades, rather than the more headline-sexy "total annihilation in 18 months unless we take drastic (really symbolic) action nowNoWNOW!!!"

    AI issues are not going to kill humanity before the end of the decade. But the technology has legitimate #concerns – the immediate one being environmental impact of data centers but also threats on a societal and potentially existential levels – that should not be drowned out by making fun of the #doomsayers, whether this is weird #ReversePsychology advertising or not.
  4. The danger of #AI models is not that they will wipe out humanity in the next 18 months. It's that repeating this uncritically as well as making fun of it will "flood the field," and tire people of the underlying message. This will make people indifferent when #guardrails and #regulation are brought up, allowing slop-peddlers like #OpenAI and #Anthropic to press on without #oversight.

    We have seen this in action: #environmental alarmists have overplayed global warming with more or less substantiated concerns preaching the end of civilization "any day now" since the 1970s, leading to legitimate concerns being drowned out and a sneaking crisis emerging over decades with consequences that are going to play out over the coming decades, rather than the more headline-sexy "total annihilation in 18 months unless we take drastic (really symbolic) action nowNoWNOW!!!"

    AI issues are not going to kill humanity before the end of the decade. But the technology has legitimate #concerns – the immediate one being environmental impact of data centers but also threats on a societal and potentially existential levels – that should not be drowned out by making fun of the #doomsayers, whether this is weird #ReversePsychology advertising or not.
  5. The danger of #AI models is not that they will wipe out humanity in the next 18 months. It's that repeating this uncritically as well as making fun of it will "flood the field," and tire people of the underlying message. This will make people indifferent when #guardrails and #regulation are brought up, allowing slop-peddlers like #OpenAI and #Anthropic to press on without #oversight.

    We have seen this in action: #environmental alarmists have overplayed global warming with more or less substantiated concerns preaching the end of civilization "any day now" since the 1970s, leading to legitimate concerns being drowned out and a sneaking crisis emerging over decades with consequences that are going to play out over the coming decades, rather than the more headline-sexy "total annihilation in 18 months unless we take drastic (really symbolic) action nowNoWNOW!!!"

    AI issues are not going to kill humanity before the end of the decade. But the technology has legitimate #concerns – the immediate one being environmental impact of data centers but also threats on a societal and potentially existential levels – that should not be drowned out by making fun of the #doomsayers, whether this is weird #ReversePsychology advertising or not.
  6. GuardRate: метрики, бенчмарки и инсайты из первого аудита (часть 2)

    В первой части мы верхнеуровнево рассмотрели результаты оценки 37 моделей, обсудили проблему оценки и сравнения guardrail моделей, рассказали, как лидерборд помогает с ней справиться, а также поделились инсайтами и руководством по использованию. Теперь пришло время заглянуть во внутреннюю кухню: в этой статье мы подробно разберем устройство GuardRate Tool и GuardRate Leaderboard, ответим на вопрос, почему выбрали именно такие бенчмарки и метрики, опишем методологию и расскажем о том, как проводились эксперименты.

    habr.com/ru/articles/1081286/

    #opensource #leaderboard #benchmarking #evaluation #llm #guardrail_metrics #ai_safety #guardrails #promptinjection #guardrail_areana

  7. GuardRate: метрики, бенчмарки и инсайты из первого аудита (часть 2)

    В первой части мы верхнеуровнево рассмотрели результаты оценки 37 моделей, обсудили проблему оценки и сравнения guardrail моделей, рассказали, как лидерборд помогает с ней справиться, а также поделились инсайтами и руководством по использованию. Теперь пришло время заглянуть во внутреннюю кухню: в этой статье мы подробно разберем устройство GuardRate Tool и GuardRate Leaderboard, ответим на вопрос, почему выбрали именно такие бенчмарки и метрики, опишем методологию и расскажем о том, как проводились эксперименты.

    habr.com/ru/articles/1081286/

    #opensource #leaderboard #benchmarking #evaluation #llm #guardrail_metrics #ai_safety #guardrails #promptinjection #guardrail_areana

  8. GuardRate: метрики, бенчмарки и инсайты из первого аудита (часть 2)

    В первой части мы верхнеуровнево рассмотрели результаты оценки 37 моделей, обсудили проблему оценки и сравнения guardrail моделей, рассказали, как лидерборд помогает с ней справиться, а также поделились инсайтами и руководством по использованию. Теперь пришло время заглянуть во внутреннюю кухню: в этой статье мы подробно разберем устройство GuardRate Tool и GuardRate Leaderboard, ответим на вопрос, почему выбрали именно такие бенчмарки и метрики, опишем методологию и расскажем о том, как проводились эксперименты.

    habr.com/ru/articles/1081286/

    #opensource #leaderboard #benchmarking #evaluation #llm #guardrail_metrics #ai_safety #guardrails #promptinjection #guardrail_areana

  9. 2/2
    "“If we build #AI systems tt r smarter than us, tt we don’t know how to control, & want to preserve themselves, they'll (do dangerous things) & win,” said Dr Bengio.. To keep such scenarios fr becoming reality, countries need to work together to decide on a common set of #guardrails & metrics to evaluate #risks of AI models.. many techs w te potential to cause harm — fr drugs & aircraft to bridges & elevators — r req'd to undergo #safetytesting & #regulatory scrutiny b4 they can be deployed"

  10. Дрейф, потеря контекста и «уверенная чушь»: протокол восстановления SDX-S

    LLM умеют многое, но иногда ломаются так, что виноватым выглядит пользователь: контекст уезжает, инструкции исчезают, инструмент падает, а модель продолжает говорить уверенно, как будто всё нормально. Мы смотрим на это не как на “плохой ответ”, а как на деградацию состояния диалога . Если не поймать момент, по цепочке шагов и становится всё убедительнее. Мы собрали процедуру SDX-S: триггеры → диагностика причины → восстановление → критерии возврата . Ниже: состояния, “дашборд” и два кейса, где это реально спасает.

    habr.com/ru/articles/985334/

    #сезон_ии_в_разработке #llmjs #chatgpt5 #prompt_engineering #guardrails #hallucinationsinai #observability #finite_state_machine #tool_use #reliability