home.social

#guardrails — Public Fediverse posts

Live and recent posts from across the Fediverse tagged #guardrails, aggregated by home.social.

  1. The danger of #AI models is not that they will wipe out humanity in the next 18 months. It's that repeating this uncritically as well as making fun of it will "flood the field," and tire people of the underlying message. This will make people indifferent when #guardrails and #regulation are brought up, allowing slop-peddlers like #OpenAI and #Anthropic to press on without #oversight.

    We have seen this in action: #environmental alarmists have overplayed global warming with more or less substantiated concerns preaching the end of civilization "any day now" since the 1970s, leading to legitimate concerns being drowned out and a sneaking crisis emerging over decades with consequences that are going to play out over the coming decades, rather than the more headline-sexy "total annihilation in 18 months unless we take drastic (really symbolic) action nowNoWNOW!!!"

    AI issues are not going to kill humanity before the end of the decade. But the technology has legitimate #concerns – the immediate one being environmental impact of data centers but also threats on a societal and potentially existential levels – that should not be drowned out by making fun of the #doomsayers, whether this is weird #ReversePsychology advertising or not.
  2. The danger of #AI models is not that they will wipe out humanity in the next 18 months. It's that repeating this uncritically as well as making fun of it will "flood the field," and tire people of the underlying message. This will make people indifferent when #guardrails and #regulation are brought up, allowing slop-peddlers like #OpenAI and #Anthropic to press on without #oversight.

    We have seen this in action: #environmental alarmists have overplayed global warming with more or less substantiated concerns preaching the end of civilization "any day now" since the 1970s, leading to legitimate concerns being drowned out and a sneaking crisis emerging over decades with consequences that are going to play out over the coming decades, rather than the more headline-sexy "total annihilation in 18 months unless we take drastic (really symbolic) action nowNoWNOW!!!"

    AI issues are not going to kill humanity before the end of the decade. But the technology has legitimate #concerns – the immediate one being environmental impact of data centers but also threats on a societal and potentially existential levels – that should not be drowned out by making fun of the #doomsayers, whether this is weird #ReversePsychology advertising or not.
  3. The danger of #AI models is not that they will wipe out humanity in the next 18 months. It's that repeating this uncritically as well as making fun of it will "flood the field," and tire people of the underlying message. This will make people indifferent when #guardrails and #regulation are brought up, allowing slop-peddlers like #OpenAI and #Anthropic to press on without #oversight.

    We have seen this in action: #environmental alarmists have overplayed global warming with more or less substantiated concerns preaching the end of civilization "any day now" since the 1970s, leading to legitimate concerns being drowned out and a sneaking crisis emerging over decades with consequences that are going to play out over the coming decades, rather than the more headline-sexy "total annihilation in 18 months unless we take drastic (really symbolic) action nowNoWNOW!!!"

    AI issues are not going to kill humanity before the end of the decade. But the technology has legitimate #concerns – the immediate one being environmental impact of data centers but also threats on a societal and potentially existential levels – that should not be drowned out by making fun of the #doomsayers, whether this is weird #ReversePsychology advertising or not.
  4. The danger of #AI models is not that they will wipe out humanity in the next 18 months. It's that repeating this uncritically as well as making fun of it will "flood the field," and tire people of the underlying message. This will make people indifferent when #guardrails and #regulation are brought up, allowing slop-peddlers like #OpenAI and #Anthropic to press on without #oversight.

    We have seen this in action: #environmental alarmists have overplayed global warming with more or less substantiated concerns preaching the end of civilization "any day now" since the 1970s, leading to legitimate concerns being drowned out and a sneaking crisis emerging over decades with consequences that are going to play out over the coming decades, rather than the more headline-sexy "total annihilation in 18 months unless we take drastic (really symbolic) action nowNoWNOW!!!"

    AI issues are not going to kill humanity before the end of the decade. But the technology has legitimate #concerns – the immediate one being environmental impact of data centers but also threats on a societal and potentially existential levels – that should not be drowned out by making fun of the #doomsayers, whether this is weird #ReversePsychology advertising or not.
  5. The danger of #AI models is not that they will wipe out humanity in the next 18 months. It's that repeating this uncritically as well as making fun of it will "flood the field," and tire people of the underlying message. This will make people indifferent when #guardrails and #regulation are brought up, allowing slop-peddlers like #OpenAI and #Anthropic to press on without #oversight.

    We have seen this in action: #environmental alarmists have overplayed global warming with more or less substantiated concerns preaching the end of civilization "any day now" since the 1970s, leading to legitimate concerns being drowned out and a sneaking crisis emerging over decades with consequences that are going to play out over the coming decades, rather than the more headline-sexy "total annihilation in 18 months unless we take drastic (really symbolic) action nowNoWNOW!!!"

    AI issues are not going to kill humanity before the end of the decade. But the technology has legitimate #concerns – the immediate one being environmental impact of data centers but also threats on a societal and potentially existential levels – that should not be drowned out by making fun of the #doomsayers, whether this is weird #ReversePsychology advertising or not.
  6. GuardRate: метрики, бенчмарки и инсайты из первого аудита (часть 2)

    В первой части мы верхнеуровнево рассмотрели результаты оценки 37 моделей, обсудили проблему оценки и сравнения guardrail моделей, рассказали, как лидерборд помогает с ней справиться, а также поделились инсайтами и руководством по использованию. Теперь пришло время заглянуть во внутреннюю кухню: в этой статье мы подробно разберем устройство GuardRate Tool и GuardRate Leaderboard, ответим на вопрос, почему выбрали именно такие бенчмарки и метрики, опишем методологию и расскажем о том, как проводились эксперименты.

    habr.com/ru/articles/1081286/

    #opensource #leaderboard #benchmarking #evaluation #llm #guardrail_metrics #ai_safety #guardrails #promptinjection #guardrail_areana

  7. GuardRate: как мы построили независимую арену для guardrail-моделей (часть 1)

    Всем привет, на связи команда HiveTrace! Мы уделяем много времени разработке собственных моделей и часто задаемся вопросом: "какой из двух гардрейлов лучше?". Если вы когда-нибудь выбирали guardrail-модель для LLM, то знаете: одни модели блокируют безобидные запросы, другие пропускают явные угрозы. Хуже всего то, что нет прозрачного стандарта сравнения. Авторы оценивают свои решения субъективно, не публикуют методологию, а результаты замеров одних и тех же моделей на одинаковых бенчмарках в разных статьях часто отличаются. Поэтому мы создали CLI, который назвали GuardRateTools - это автоматизированный пайплайн оценки guardrail-моделей. Его интерфейс – HiveTrace GuardRate Leaderboard, уже открыт для просмотра. GuardRateTools интегрирован в CI/CD процесс дообучения внутренних guardrail-моделей и обеспечивает честное сравнение моделей в равных условиях. CLI запускает проверку одной командой: подтягивает датасеты и конфигурацию модели из YAML, разворачивает изолированную среду, которая создается индивидуально для каждой модели, прогоняет модель по фиксированному набору бенчмарков и считает метрики. Сырые ответы от модели, логи и итоговые метрики сохраняются в артефакты, поэтому любой результат можно проверить и воспроизвести. Автоматизация исключает ручной труд, снижает влияние человеческого фактора и сокращает время оценки новых решений в области гардрейлов. HiveTrace GuardRate Leaderboard уже доступен для всех! В третьем квартале 2026 года мы выложим исходный код CLI. Если хотите протестировать свою модель, свяжитесь с нами. Контакты вы найдёте в конце статьи.

    habr.com/ru/companies/raft/art

    #llm #guardrails #guardrail_metrics #leaderboard #evaluation #guardrail_areana #ai_safety #promptinjection #benchmarking #opensource

  8. [The WSJ] Let AI Run [Their] Office Vending Machine. It Lost Hundreds Of Dollars.
    Anthropic’s Claude ran a snack operation in the WSJ newsroom. It gave away a free PlayStation, ordered a live fish—and taught us lessons about the future of AI agents.
    --
    wsj.com/tech/ai/anthropic-clau <-- shared media article
    --
    youtu.be/SpPhm7S9vsQ?si=aJQ2_B <-- shared video
    --
    [When you get clever journalists to !$%^&*@ with AI… bravo! And this is a very simple situation, vending machines have been around since literally the Roman Empire
    “You are using the wrong prompts” and LUDDITES! In the comments in 3… 2… 1…]
    #vendingmachine #artificialintelligence #AIHallucination #hallucinations #emperorsnewclothes #ohhhshiny #experiment #contextwindow #AIagent #claude #autonomous #compliance #fish #PlayStation #snackliberationday #knowledgeboundaries #guardrails #redteam #GenAI cynicism
    @WSJ @Anthropic @Claude

  9. Po-Gemüse 🤪

    Ein KI #Chatbot des US Gesundheitsministeriums sorgt für Kritik, weil er auf provozierende Fragen unpassende Empfehlungen liefert.

    Berichte zeigen, dass offenbar #Grok ohne ausreichende Vorgaben eingesetzt wurde. Nutzer meldeten Antworten, die weder gesundheitlich sinnvoll noch dem Zweck des Bots entsprechen.

    golem.de/news/ernaehrung-chatb

    #KIFail #Guardrails #KIChatbot