home.social

#guardrails — Public Fediverse posts

Live and recent posts from across the Fediverse tagged #guardrails, aggregated by home.social.

  1. The danger of #AI models is not that they will wipe out humanity in the next 18 months. It's that repeating this uncritically as well as making fun of it will "flood the field," and tire people of the underlying message. This will make people indifferent when #guardrails and #regulation are brought up, allowing slop-peddlers like #OpenAI and #Anthropic to press on without #oversight.

    We have seen this in action: #environmental alarmists have overplayed global warming with more or less substantiated concerns preaching the end of civilization "any day now" since the 1970s, leading to legitimate concerns being drowned out and a sneaking crisis emerging over decades with consequences that are going to play out over the coming decades, rather than the more headline-sexy "total annihilation in 18 months unless we take drastic (really symbolic) action nowNoWNOW!!!"

    AI issues are not going to kill humanity before the end of the decade. But the technology has legitimate #concerns – the immediate one being environmental impact of data centers but also threats on a societal and potentially existential levels – that should not be drowned out by making fun of the #doomsayers, whether this is weird #ReversePsychology advertising or not.
  2. The danger of #AI models is not that they will wipe out humanity in the next 18 months. It's that repeating this uncritically as well as making fun of it will "flood the field," and tire people of the underlying message. This will make people indifferent when #guardrails and #regulation are brought up, allowing slop-peddlers like #OpenAI and #Anthropic to press on without #oversight.

    We have seen this in action: #environmental alarmists have overplayed global warming with more or less substantiated concerns preaching the end of civilization "any day now" since the 1970s, leading to legitimate concerns being drowned out and a sneaking crisis emerging over decades with consequences that are going to play out over the coming decades, rather than the more headline-sexy "total annihilation in 18 months unless we take drastic (really symbolic) action nowNoWNOW!!!"

    AI issues are not going to kill humanity before the end of the decade. But the technology has legitimate #concerns – the immediate one being environmental impact of data centers but also threats on a societal and potentially existential levels – that should not be drowned out by making fun of the #doomsayers, whether this is weird #ReversePsychology advertising or not.
  3. GuardRate: метрики, бенчмарки и инсайты из первого аудита (часть 2)

    В первой части мы верхнеуровнево рассмотрели результаты оценки 37 моделей, обсудили проблему оценки и сравнения guardrail моделей, рассказали, как лидерборд помогает с ней справиться, а также поделились инсайтами и руководством по использованию. Теперь пришло время заглянуть во внутреннюю кухню: в этой статье мы подробно разберем устройство GuardRate Tool и GuardRate Leaderboard, ответим на вопрос, почему выбрали именно такие бенчмарки и метрики, опишем методологию и расскажем о том, как проводились эксперименты.

    habr.com/ru/articles/1081286/

    #opensource #leaderboard #benchmarking #evaluation #llm #guardrail_metrics #ai_safety #guardrails #promptinjection #guardrail_areana

  4. GuardRate: метрики, бенчмарки и инсайты из первого аудита (часть 2)

    В первой части мы верхнеуровнево рассмотрели результаты оценки 37 моделей, обсудили проблему оценки и сравнения guardrail моделей, рассказали, как лидерборд помогает с ней справиться, а также поделились инсайтами и руководством по использованию. Теперь пришло время заглянуть во внутреннюю кухню: в этой статье мы подробно разберем устройство GuardRate Tool и GuardRate Leaderboard, ответим на вопрос, почему выбрали именно такие бенчмарки и метрики, опишем методологию и расскажем о том, как проводились эксперименты.

    habr.com/ru/articles/1081286/

    #opensource #leaderboard #benchmarking #evaluation #llm #guardrail_metrics #ai_safety #guardrails #promptinjection #guardrail_areana

  5. GuardRate: метрики, бенчмарки и инсайты из первого аудита (часть 2)

    В первой части мы верхнеуровнево рассмотрели результаты оценки 37 моделей, обсудили проблему оценки и сравнения guardrail моделей, рассказали, как лидерборд помогает с ней справиться, а также поделились инсайтами и руководством по использованию. Теперь пришло время заглянуть во внутреннюю кухню: в этой статье мы подробно разберем устройство GuardRate Tool и GuardRate Leaderboard, ответим на вопрос, почему выбрали именно такие бенчмарки и метрики, опишем методологию и расскажем о том, как проводились эксперименты.

    habr.com/ru/articles/1081286/

    #opensource #leaderboard #benchmarking #evaluation #llm #guardrail_metrics #ai_safety #guardrails #promptinjection #guardrail_areana

  6. It’s a hard(ending) World! VMTN Session
    I forgot to post that my VMTN session from VMware Explore Barcelona is online for your viewing pleasure!

    In this session I talk about compliance scanning in Aria Operations and Tanzu Guardrails. Checking your VMware and cloud environments for regulatory (ISO27001,...) and security compliance. I emphasise the MITRE Attack framework he
    bitstream.geenrits.net/vmware/
    #Security #VMware #Guardrails #VROps