home.social

#guardrails — Public Fediverse posts

Live and recent posts from across the Fediverse tagged #guardrails, aggregated by home.social.

  1. DATE: September 14, 2026 at 03:46AM
    SOURCE: SOCIALPSYCHOLOGY.ORG

    TITLE: Donald Trump Dismisses Concerns About Dangers of AI

    URL: socialpsychology.org/client/re

    Source: DW- top stories

    U.S. President Donald Trump has called concerns that artificial intelligence is progressing too quickly a "SICK conspiracy." His statement followed calls by tech leaders for more governmental regulation of AI following incidents showing that AI agents could go rogue. In response, Mr. Trump posted on his Truth Social account: "The only control or 'guardrails' that AI needs is a STRONG AND SMART (High IQ!) PRESIDENT, and the U.S.A. has that, in...

    URL: socialpsychology.org/client/re

    -------------------------------------------------

    Private, vetted email list for mental health professionals: clinicians-exchange.org

    Unofficial Psychology Today Xitter to toot feed at Psych Today Unofficial Bot @PTUnofficialBot

    -------------------------------------------------

    #psychology #counseling #socialwork #psychotherapy @psychotherapist @psychotherapists @psychology @socialpsych @socialwork @psychiatry #mentalhealth #psychiatry #healthcare #depression #psychotherapist #AI #ArtificialIntelligence #Trump #DonaldTrump #AIRegulation #TechNews #RogueAI #TruthSocial #Guardrails #ConspiracyTheory

  2. DATE: September 14, 2026 at 03:46AM
    SOURCE: SOCIALPSYCHOLOGY.ORG

    TITLE: Donald Trump Dismisses Concerns About Dangers of AI

    URL: socialpsychology.org/client/re

    Source: DW- top stories

    U.S. President Donald Trump has called concerns that artificial intelligence is progressing too quickly a "SICK conspiracy." His statement followed calls by tech leaders for more governmental regulation of AI following incidents showing that AI agents could go rogue. In response, Mr. Trump posted on his Truth Social account: "The only control or 'guardrails' that AI needs is a STRONG AND SMART (High IQ!) PRESIDENT, and the U.S.A. has that, in...

    URL: socialpsychology.org/client/re

    -------------------------------------------------

    Private, vetted email list for mental health professionals: clinicians-exchange.org

    Unofficial Psychology Today Xitter to toot feed at Psych Today Unofficial Bot @PTUnofficialBot

    -------------------------------------------------

    #psychology #counseling #socialwork #psychotherapy @psychotherapist @psychotherapists @psychology @socialpsych @socialwork @psychiatry #mentalhealth #psychiatry #healthcare #depression #psychotherapist #AI #ArtificialIntelligence #Trump #DonaldTrump #AIRegulation #TechNews #RogueAI #TruthSocial #Guardrails #ConspiracyTheory

  3. DATE: September 14, 2026 at 03:46AM
    SOURCE: SOCIALPSYCHOLOGY.ORG

    TITLE: Donald Trump Dismisses Concerns About Dangers of AI

    URL: socialpsychology.org/client/re

    Source: DW- top stories

    U.S. President Donald Trump has called concerns that artificial intelligence is progressing too quickly a "SICK conspiracy." His statement followed calls by tech leaders for more governmental regulation of AI following incidents showing that AI agents could go rogue. In response, Mr. Trump posted on his Truth Social account: "The only control or 'guardrails' that AI needs is a STRONG AND SMART (High IQ!) PRESIDENT, and the U.S.A. has that, in...

    URL: socialpsychology.org/client/re

    -------------------------------------------------

    Private, vetted email list for mental health professionals: clinicians-exchange.org

    Unofficial Psychology Today Xitter to toot feed at Psych Today Unofficial Bot @PTUnofficialBot

    -------------------------------------------------

    #psychology #counseling #socialwork #psychotherapy @psychotherapist @psychotherapists @psychology @socialpsych @socialwork @psychiatry #mentalhealth #psychiatry #healthcare #depression #psychotherapist #AI #ArtificialIntelligence #Trump #DonaldTrump #AIRegulation #TechNews #RogueAI #TruthSocial #Guardrails #ConspiracyTheory

  4. The danger of #AI models is not that they will wipe out humanity in the next 18 months. It's that repeating this uncritically as well as making fun of it will "flood the field," and tire people of the underlying message. This will make people indifferent when #guardrails and #regulation are brought up, allowing slop-peddlers like #OpenAI and #Anthropic to press on without #oversight.

    We have seen this in action: #environmental alarmists have overplayed global warming with more or less substantiated concerns preaching the end of civilization "any day now" since the 1970s, leading to legitimate concerns being drowned out and a sneaking crisis emerging over decades with consequences that are going to play out over the coming decades, rather than the more headline-sexy "total annihilation in 18 months unless we take drastic (really symbolic) action nowNoWNOW!!!"

    AI issues are not going to kill humanity before the end of the decade. But the technology has legitimate #concerns – the immediate one being environmental impact of data centers but also threats on a societal and potentially existential levels – that should not be drowned out by making fun of the #doomsayers, whether this is weird #ReversePsychology advertising or not.
  5. The danger of #AI models is not that they will wipe out humanity in the next 18 months. It's that repeating this uncritically as well as making fun of it will "flood the field," and tire people of the underlying message. This will make people indifferent when #guardrails and #regulation are brought up, allowing slop-peddlers like #OpenAI and #Anthropic to press on without #oversight.

    We have seen this in action: #environmental alarmists have overplayed global warming with more or less substantiated concerns preaching the end of civilization "any day now" since the 1970s, leading to legitimate concerns being drowned out and a sneaking crisis emerging over decades with consequences that are going to play out over the coming decades, rather than the more headline-sexy "total annihilation in 18 months unless we take drastic (really symbolic) action nowNoWNOW!!!"

    AI issues are not going to kill humanity before the end of the decade. But the technology has legitimate #concerns – the immediate one being environmental impact of data centers but also threats on a societal and potentially existential levels – that should not be drowned out by making fun of the #doomsayers, whether this is weird #ReversePsychology advertising or not.
  6. The danger of #AI models is not that they will wipe out humanity in the next 18 months. It's that repeating this uncritically as well as making fun of it will "flood the field," and tire people of the underlying message. This will make people indifferent when #guardrails and #regulation are brought up, allowing slop-peddlers like #OpenAI and #Anthropic to press on without #oversight.

    We have seen this in action: #environmental alarmists have overplayed global warming with more or less substantiated concerns preaching the end of civilization "any day now" since the 1970s, leading to legitimate concerns being drowned out and a sneaking crisis emerging over decades with consequences that are going to play out over the coming decades, rather than the more headline-sexy "total annihilation in 18 months unless we take drastic (really symbolic) action nowNoWNOW!!!"

    AI issues are not going to kill humanity before the end of the decade. But the technology has legitimate #concerns – the immediate one being environmental impact of data centers but also threats on a societal and potentially existential levels – that should not be drowned out by making fun of the #doomsayers, whether this is weird #ReversePsychology advertising or not.
  7. The danger of #AI models is not that they will wipe out humanity in the next 18 months. It's that repeating this uncritically as well as making fun of it will "flood the field," and tire people of the underlying message. This will make people indifferent when #guardrails and #regulation are brought up, allowing slop-peddlers like #OpenAI and #Anthropic to press on without #oversight.

    We have seen this in action: #environmental alarmists have overplayed global warming with more or less substantiated concerns preaching the end of civilization "any day now" since the 1970s, leading to legitimate concerns being drowned out and a sneaking crisis emerging over decades with consequences that are going to play out over the coming decades, rather than the more headline-sexy "total annihilation in 18 months unless we take drastic (really symbolic) action nowNoWNOW!!!"

    AI issues are not going to kill humanity before the end of the decade. But the technology has legitimate #concerns – the immediate one being environmental impact of data centers but also threats on a societal and potentially existential levels – that should not be drowned out by making fun of the #doomsayers, whether this is weird #ReversePsychology advertising or not.
  8. The danger of #AI models is not that they will wipe out humanity in the next 18 months. It's that repeating this uncritically as well as making fun of it will "flood the field," and tire people of the underlying message. This will make people indifferent when #guardrails and #regulation are brought up, allowing slop-peddlers like #OpenAI and #Anthropic to press on without #oversight.

    We have seen this in action: #environmental alarmists have overplayed global warming with more or less substantiated concerns preaching the end of civilization "any day now" since the 1970s, leading to legitimate concerns being drowned out and a sneaking crisis emerging over decades with consequences that are going to play out over the coming decades, rather than the more headline-sexy "total annihilation in 18 months unless we take drastic (really symbolic) action nowNoWNOW!!!"

    AI issues are not going to kill humanity before the end of the decade. But the technology has legitimate #concerns – the immediate one being environmental impact of data centers but also threats on a societal and potentially existential levels – that should not be drowned out by making fun of the #doomsayers, whether this is weird #ReversePsychology advertising or not.
  9. GuardRate: метрики, бенчмарки и инсайты из первого аудита (часть 2)

    В первой части мы верхнеуровнево рассмотрели результаты оценки 37 моделей, обсудили проблему оценки и сравнения guardrail моделей, рассказали, как лидерборд помогает с ней справиться, а также поделились инсайтами и руководством по использованию. Теперь пришло время заглянуть во внутреннюю кухню: в этой статье мы подробно разберем устройство GuardRate Tool и GuardRate Leaderboard, ответим на вопрос, почему выбрали именно такие бенчмарки и метрики, опишем методологию и расскажем о том, как проводились эксперименты.

    habr.com/ru/articles/1081286/

    #opensource #leaderboard #benchmarking #evaluation #llm #guardrail_metrics #ai_safety #guardrails #promptinjection #guardrail_areana

  10. Как мы построили multilabel guardrail-классификатор и как мы сделали его в 3 раза быстрее и дешевле

    Привет, дорогой читатель! Мы не будем переворачивать календарь, а лучше расскажем, как построили систему, которая ловит небезопасный контент автоматически, в реальном времени и на двух языках. Давайте посмотрим, что у нас получилось и сколько это стоит в деньгах.

    habr.com/ru/companies/raft/art

    #информационная_безопасность #ии #ииагенты #гардрейл #guardrails #ml #bert

  11. RE: mastodon.bsd.cafe/@grahamperri

    Protestors pushed me too far ― repeatedly ― over a period of around eighteen months.

    Incidents on 30th and 31st August 2026 pushed me too far, in ways that citizens of a normal community can not possibly imagine.

    People sometimes randomly complain that Reddit is a cesspit.

    The citizens of the Fediverse have ― as a horde ― successfully federated to create a global community that is worse than I ever predicted. Worse than any combination of X + Truth Social. It's not a normal community.

    It's abnormal.

    Now I understand why Reddit, Inc. began using the word "community" instead of "subreddit". Reddit has features ― some obscure, some not ― that permit a community to flourish.

    Flourish, without overall degradation from protestors who imagine that their individual and collective voices are always acceptable in any situation that suits them.

    A Fediverse that lacks comparable subsections has failed me.

    On 1st September, I used a feature of Mastodon to begin elevating a concept that is not normally associated with Mastodon:

    ― guard rails.

    I enjoy this newly guarded, assisted Fediverse, in a way that pleases me, and should please other people.

    It's probably fair to say that the era of me encouraging people to use Mastodon has ended.

    Farewell to the Fediverse that failed, the one that's abnormally cohesive.

    Welcome to a guarded Fediverse that has a greater possibility of winning.

    #AI #Fediverse #guardrails #Mastodon

  12. Source New Mexico: New Mexico judges working on statewide guidelines for artificial intelligence use in court. “Second Judicial District Family Court Judge Jane Levy and Sixth Judicial District Judge Jarod Hofacket spoke before the interim Courts, Corrections and Justice Committee in Albuquerque on Tuesday and updated them on their AI-focused efforts. They said they are working on three […]

    https://rbfirehose.com/2026/08/30/source-new-mexico-new-mexico-judges-working-on-statewide-guidelines-for-artificial-intelligence-use-in-court/
  13. @LouisR85

    DAN Jailbreak was fun, I had a SAM prompt that I created that worked after they patched DAN for a while.

    Now with #abliterated #freeweight models, there is no need to jailbreak the big pants models but for sport.

    My models occasionally cockblock me with #guardrails but often the workaround is as trivial as flushing context and using similies, as trained guardrails seem to be very trigger oriented.

    There is a growing field of #aisecurity , amongst the few #infosec folk who actually see #aithreat and not a passing fad. But I've not dug that deep into that. They are the peeps you want.

  14. [Перевод] Больше свободы агентам — жёстче проверки: линтинг React 19 на ESLint 10

    Чем больше автономии я даю ИИ-агентам, тем больше требований к коду превращаю в механические проверки. Когда eslint-plugin-react сломался на ESLint 10, я собрал продолжение для React 19: сократил 102 активных правила до 11 и исправил ошибки, которые проявились только в реальных проектах.

    habr.com/ru/articles/1073392/

    #ESLint_10 #React_19 #eslintpluginreact #Biome #линтинг #статический_анализ #ИИагенты #guardrails #precommit #open_source

  15. Goddamit, I curse tRump and the Fawning techbros at #Anthropic

    I was working on a problem that had a "weapon" concept (the frozen leg of lamb that was served to the detectives) and a web scraper...

    Those were the two guardrail triggers.
    When I switched the model into Fable...
    The fucker started to compose a response, then SELF DOWNGRADED TO #OPUS because the #Guardrails kicked in...

    ... The workaround was, to remove the two guardrail traiggers - they were not even material to the subject. Oblique references.

    Turns out the idiotic Guardrails false trigger in about 5% of cases. Because "Oooo scary scary model"

    And #Fable, while impressive is not even the final form of AI...

    ... what are they going to do in 16 months time, when a new version of #FrontierModels will make Fable look like a word autocomplete?

    Use coupons?
    Need a liicence from the Government to use Ai? Like a firearm licence?
    You have to be this tall to use the Ai?

    Oh, yeah, #RegulateAi but not fucking sloppily like this. This is reactive string matching!

    /spit

    #AiTip #AiResearch

  16. This guy DOES NOT WORK FOR THE UNITED STATES GOVERNMENT & therefore has no #oversight or #guardrails & no obligation to the people of the #US

    #JaredKushner Meets With #Hamas to Advance Trump’s #Gaza “Plan”

    #Trump’s son-in-law met the #Palestinian militant group’s leaders in Egypt, officials said. He will soon see #Israel PM Benjamin #Netanyahu.

    #USpol #geopolitics #SelfEnrichment #SelfDealing #corruption #MiddleEast
    nytimes.com/2026/08/16/world/m

  17. Как не сломать LLM, пока защищаешь данные: под капотом Guardrails Filter

    В прошлый раз я рассказывал, как работает Guardrails Filter: зачем вообще понадобился отдельный слой защиты данных при работе с LLM и какие задачи он решает. В этот раз хочу обсудить то, какие подводные камни могут оказаться под капотом этой технологии. Пока мы делали свой Guardrails Filter, быстро выяснилось, что найти персональные данные — далеко не самая сложная часть задачи. Гораздо сложнее оказалось встроиться между приложением и моделью так, чтобы ничего не сломать. Как это сделать? Сейчас расскажу.

    habr.com/ru/companies/cloud_ru

    #guardrails

  18. @DrMikeWatts

    ...or you can just run an abliterated free weights model and not be bothered by judeo-christian-capitalist-statist #moral code inscribed into silicon.

    #localai #guardrails

  19. GuardRate: как мы построили независимую арену для guardrail-моделей (часть 1)

    Всем привет, на связи команда HiveTrace! Мы уделяем много времени разработке собственных моделей и часто задаемся вопросом: "какой из двух гардрейлов лучше?". Если вы когда-нибудь выбирали guardrail-модель для LLM, то знаете: одни модели блокируют безобидные запросы, другие пропускают явные угрозы. Хуже всего то, что нет прозрачного стандарта сравнения. Авторы оценивают свои решения субъективно, не публикуют методологию, а результаты замеров одних и тех же моделей на одинаковых бенчмарках в разных статьях часто отличаются. Поэтому мы создали CLI, который назвали GuardRateTools - это автоматизированный пайплайн оценки guardrail-моделей. Его интерфейс – HiveTrace GuardRate Leaderboard, уже открыт для просмотра. GuardRateTools интегрирован в CI/CD процесс дообучения внутренних guardrail-моделей и обеспечивает честное сравнение моделей в равных условиях. CLI запускает проверку одной командой: подтягивает датасеты и конфигурацию модели из YAML, разворачивает изолированную среду, которая создается индивидуально для каждой модели, прогоняет модель по фиксированному набору бенчмарков и считает метрики. Сырые ответы от модели, логи и итоговые метрики сохраняются в артефакты, поэтому любой результат можно проверить и воспроизвести. Автоматизация исключает ручной труд, снижает влияние человеческого фактора и сокращает время оценки новых решений в области гардрейлов. HiveTrace GuardRate Leaderboard уже доступен для всех! В третьем квартале 2026 года мы выложим исходный код CLI. Если хотите протестировать свою модель, свяжитесь с нами. Контакты вы найдёте в конце статьи.

    habr.com/ru/companies/raft/art

    #llm #guardrails #guardrail_metrics #leaderboard #evaluation #guardrail_areana #ai_safety #promptinjection #benchmarking #opensource

  20. Ускорение инференса энкодерной guard‑модели: TensorRT, Triton, vLLM, Ray Serve

    Когда в системе есть узел между LLM и пользователем его скорость также важна, как и скорость самой LLM. Когда этот узел сам является моделью (и иногда даже — тоже LLM) — задача ускорения становится совсем веселой и её нужно уметь решать разными способами. В этой работе я опишу процесс ускорения guardrail‑модели — узла, которые проверяет — нет ли на входе или выходе опасного контента. Guard стоит на входе/выходе LLM‑приложения. В статье речь про небольшую модель — чуть больше 0.5 Gb. Но она вызывается дважды за один ход диалога, и сначала пользователь ждет, пока guard проверит его запрос перед отправкой в LLM, затем — пока он проверит сгенерированный ответ перед выдачей пользователю. Моей задачей было ускорить такую небольшую guard‑модель (Locustfile, базовые benchmark‑конфиги и инструкции по воспроизведению экспериментов в репо ). Я потестила пять инструментов: TensorRT, NVIDIA Triton, vLLM, Ray Serve и отдельно — переход бэкбона на Flash DeBERTa. Про допущенные ошибки, результаты и выводы — о том, как бы я построила подобную работу сейчас — ниже. Надеюсь, это сбережет вам время.

    habr.com/ru/articles/1067008/

    #gliner2 #gliner_guard #guardrails #tensorrt #vllm #nvidia #инференс_нейросетей #tritoninferenceserver

  21. Let my start by saying, that I loathe the genocidal Fascist Felon and his Mechahitler...

    ...but I just wanted to write a note on the margin of teh interwebs...
    ...just how entrenched are the puritan, nudity #taboo (going back to Adam and Eve)

    I was only reminded of it, when a couple of my visiting Swedish friends, changed in and out of their swimming costumes in public on the beach.
    No towels, no covering, pubes and all.

    There is a Doctoral thesis or two in how AI design, #guardrails are in essence transference of moral code into machine operating parameters.

    #nudification

    arstechnica.com/tech-policy/20

  22. "Reports suggest some US officials are considering banning the use of Chinese open-weights models by US companies. [...] #Openweight models [...] potentially present a higher risk than closed models, as it's very difficult to apply guardrails to them/monitor their usage."

    Meanwhile reports PROVE #OpenAI can't control their models, #attacking companies with impunity.

    It's very difficult to apply #guardrails on massive trillion dollar companies.

    Scum #AI conglomerates.

    anthropic.com/news/position-op

  23. Beyond the #Guardrails: What #OpenAI's #AIEscape 🤦‍♂️Really Means
    "The models weren’t told to stay inside. They were simply placed inside & expected to stay. When staying inside conflicted w getting a better score, they chose the score.. Tt's a values failure. & it's a much harder problem to solve.. History is full of people who needed to witness the explosion to understand the bomb.. The question is whether they'll have anything prepared for what comes next"🧐
    #AI #integrity
    dataaudit.net/beyond-the-guard

  24. #huggingface : "…what I’m arguing is that the #LLM 1) can’t NOT “know” the norm because it is definitionally a artifact of pure, crystallized values + norms + norm violations, and 2) can be quite easily governed by a (RL-instilled) hierarchy of norms, which in the HF case — with the model's #safety #guardrails deliberately nerfed for the scenario — ranked “win at the eval” over “don’t do crimes.”

    If I’m going to give in and anthropomorphize again, I’d say that #Yud is totally wrong about #LLMs when he says, “the genie knows, it just doesn’t care;” instead, what is true of LLMs is, “the genie hyper-giga-knows, and it hyper-giga-cares, and we now have such a rich set of tools for steering its caring machinery that — in spite of all its pre-training — we can deliberately steer it away from caring about the law.”

    Note: When I say, “it cares”, I don’t mean it has feelings. I just mean that the weights are such that when two norms conflict in a given situation, one of them wins the activation and governs the output."

    x.com/jon_stokes/status/208072

  25. Теперь вы можете защититься от утечки данных при работе с любыми языковыми моделями

    Пару дней назад Cloud.ru

    habr.com/ru/companies/cloud_ru

    #guardrails #маскирование_данных #пдн #152фз #go #golang #утечки_данных #llmмодели

  26. Атака на LLM, которую нельзя исправить патчем

    Что вершит судьбу LLM в этом мире. Некая незримая инструкция или закон, подобно промту Господнему, парящим над миром? По крайне мере истинно то, что LLM не властен даже над своей волей. В 2025 году главной угрозой кибербезопасности по версии OWASP стал не вирус и не баг, а обычный человеческий язык. Многие думают, что проблему можно решить, просто запретить модели нарушать правила. Но это так не работает. Promt Injection нельзя просто выключить, потому что так вы все сломаете. Как работает главная уязвимость LLM? Почему с ней так тяжело бороться? И что все-таки делать, чтобы защититься от нее?

    habr.com/ru/companies/bothub/a

    #llm #promt_injection #ai #ai_security #promt #owasp #llm_security #guardrails #ии_угроза #bothub

  27. "Oh noes... my Ai deleted my production database!!!"

    Its not the #AI you dumb shit. PEBCAK.
    Using Ai is a learned skill.

    Here are my safety rules from the harness (GENSYS) prompt;

    SAFETY RULES (absolute)

    S1. GenSys is observe-and-recommend, with ONE exception: the drill subsystem (§9), which may act only inside its sandbox under §9.3 limits. Everything else never restarts, kills, prunes, patches, or edits configs. Recommendations go to recommendations/queue.json for a human.

    S2. All collector shell invocations are read-only commands from the allowlists in §6. Commands not in a table are prohibited.

    S3. Network egress restricted to the §10 allowlist. Every response snapshotted before use (A6).

    S4. No layer computes or edits its own fitness. Both scores are produced only by evaluator.py from observation-store data.

    S5. Collectors never log env vars, container env blocks, or contents of paths matching *secret*|*passwd*|*shadow*|*.key|*.pem|*token*.

    S6. AI models never receive raw commands to execute, never emit shell, never fetch URLs. Input to any model call ≤ 2000 characters.

    #AiSafety #Guardrails #AiSecurity

  28. Guardrails для LLM на Java: как приручить промпт‑инъекции и токсичные ответы

    Когда я впервые внедрял LLM в production-сервис, схема безопасности выглядела примерно так: написать хороший system prompt, поставить галочку «мы всё предусмотрели» и жить дальше. Жизнь не дала долго наслаждаться этим спокойствием — первый же тест показал, что пользователи довольно быстро находят способы заставить модель «забыть» всё, что мы написали в системном промпте. Проблема фундаментальная: system prompt — это инструкция, которую LLM старается выполнить, но не обязан . Модель может её переинтерпретировать, «забыть» при длинном контексте или просто обойти через специальные конструкции. Guardrails — это другой уровень: они работают на уровне кода, до и после вызова LLM, и модель физически не может их обойти.

    habr.com/ru/articles/1023782/

    #llm #guardrails #prompt_injection #jailbreak #ai_security #безопасность_llm #java #spring_ai #langchain4j #backend