#guardrails — Public Fediverse posts
Live and recent posts from across the Fediverse tagged #guardrails, aggregated by home.social.
-
OpenAI chief scientist warns no-one is prepared for consequences of AI
by Laura Cress / via BBC
#AI #LLM #OpenAI #ChatGPT #guardrails #regulation #danger #tech #insights
-
OpenAI chief scientist warns no-one is prepared for consequences of AI
by Laura Cress / via BBC
#AI #LLM #OpenAI #ChatGPT #guardrails #regulation #danger #tech #insights
-
OpenAI chief scientist warns no-one is prepared for consequences of AI
by Laura Cress / via BBC
#AI #LLM #OpenAI #ChatGPT #guardrails #regulation #danger #tech #insights
-
OpenAI chief scientist warns no-one is prepared for consequences of AI
by Laura Cress / via BBC
#AI #LLM #OpenAI #ChatGPT #guardrails #regulation #danger #tech #insights
-
OpenAI chief scientist warns no-one is prepared for consequences of AI
by Laura Cress / via BBC
#AI #LLM #OpenAI #ChatGPT #guardrails #regulation #danger #tech #insights
-
US and China eye Trump-Xi talks on AI guardrails despite tech rift https://www.byteseu.com/2343262/ #AI #ArtificialIntelligence #China #despite #Eye #guardrails #Rift #talks #TECH #TrumpXi #US
-
Как мы построили multilabel guardrail-классификатор и как мы сделали его в 3 раза быстрее и дешевле
Привет, дорогой читатель! Мы не будем переворачивать календарь, а лучше расскажем, как построили систему, которая ловит небезопасный контент автоматически, в реальном времени и на двух языках. Давайте посмотрим, что у нас получилось и сколько это стоит в деньгах.
https://habr.com/ru/companies/raft/articles/1077768/
#информационная_безопасность #ии #ииагенты #гардрейл #guardrails #ml #bert
-
RE: https://mastodon.bsd.cafe/@grahamperrin/117197654600279594
Protestors pushed me too far ― repeatedly ― over a period of around eighteen months.
Incidents on 30th and 31st August 2026 pushed me too far, in ways that citizens of a normal community can not possibly imagine.
People sometimes randomly complain that Reddit is a cesspit.
The citizens of the Fediverse have ― as a horde ― successfully federated to create a global community that is worse than I ever predicted. Worse than any combination of X + Truth Social. It's not a normal community.
It's abnormal.
Now I understand why Reddit, Inc. began using the word "community" instead of "subreddit". Reddit has features ― some obscure, some not ― that permit a community to flourish.
Flourish, without overall degradation from protestors who imagine that their individual and collective voices are always acceptable in any situation that suits them.
A Fediverse that lacks comparable subsections has failed me.
On 1st September, I used a feature of Mastodon to begin elevating a concept that is not normally associated with Mastodon:
― guard rails.
I enjoy this newly guarded, assisted Fediverse, in a way that pleases me, and should please other people.
It's probably fair to say that the era of me encouraging people to use Mastodon has ended.
Farewell to the Fediverse that failed, the one that's abnormally cohesive.
Welcome to a guarded Fediverse that has a greater possibility of winning.
-
Source New Mexico: New Mexico judges working on statewide guidelines for artificial intelligence use in court. “Second Judicial District Family Court Judge Jane Levy and Sixth Judicial District Judge Jarod Hofacket spoke before the interim Courts, Corrections and Justice Committee in Albuquerque on Tuesday and updated them on their AI-focused efforts. They said they are working on three […]
https://rbfirehose.com/2026/08/30/source-new-mexico-new-mexico-judges-working-on-statewide-guidelines-for-artificial-intelligence-use-in-court/ -
DAN Jailbreak was fun, I had a SAM prompt that I created that worked after they patched DAN for a while.
Now with #abliterated #freeweight models, there is no need to jailbreak the big pants models but for sport.
My models occasionally cockblock me with #guardrails but often the workaround is as trivial as flushing context and using similies, as trained guardrails seem to be very trigger oriented.
There is a growing field of #aisecurity , amongst the few #infosec folk who actually see #aithreat and not a passing fad. But I've not dug that deep into that. They are the peeps you want.
-
[Перевод] Больше свободы агентам — жёстче проверки: линтинг React 19 на ESLint 10
Чем больше автономии я даю ИИ-агентам, тем больше требований к коду превращаю в механические проверки. Когда eslint-plugin-react сломался на ESLint 10, я собрал продолжение для React 19: сократил 102 активных правила до 11 и исправил ошибки, которые проявились только в реальных проектах.
https://habr.com/ru/articles/1073392/
#ESLint_10 #React_19 #eslintpluginreact #Biome #линтинг #статический_анализ #ИИагенты #guardrails #precommit #open_source
-
Goddamit, I curse tRump and the Fawning techbros at #Anthropic
I was working on a problem that had a "weapon" concept (the frozen leg of lamb that was served to the detectives) and a web scraper...
Those were the two guardrail triggers.
When I switched the model into Fable...
The fucker started to compose a response, then SELF DOWNGRADED TO #OPUS because the #Guardrails kicked in...... The workaround was, to remove the two guardrail traiggers - they were not even material to the subject. Oblique references.
Turns out the idiotic Guardrails false trigger in about 5% of cases. Because "Oooo scary scary model"
And #Fable, while impressive is not even the final form of AI...
... what are they going to do in 16 months time, when a new version of #FrontierModels will make Fable look like a word autocomplete?
Use coupons?
Need a liicence from the Government to use Ai? Like a firearm licence?
You have to be this tall to use the Ai?Oh, yeah, #RegulateAi but not fucking sloppily like this. This is reactive string matching!
/spit
-
This guy DOES NOT WORK FOR THE UNITED STATES GOVERNMENT & therefore has no #oversight or #guardrails & no obligation to the people of the #US
#JaredKushner Meets With #Hamas to Advance Trump’s #Gaza “Plan”
#Trump’s son-in-law met the #Palestinian militant group’s leaders in Egypt, officials said. He will soon see #Israel PM Benjamin #Netanyahu.
#USpol #geopolitics #SelfEnrichment #SelfDealing #corruption #MiddleEast
https://www.nytimes.com/2026/08/16/world/middleeast/kushner-hamas-talks.html?unlocked_article_code=1.6FA.dh8p.WTsJLRskAVbK&smid=nytcore-ios-share -
Как не сломать LLM, пока защищаешь данные: под капотом Guardrails Filter
В прошлый раз я рассказывал, как работает Guardrails Filter: зачем вообще понадобился отдельный слой защиты данных при работе с LLM и какие задачи он решает. В этот раз хочу обсудить то, какие подводные камни могут оказаться под капотом этой технологии. Пока мы делали свой Guardrails Filter, быстро выяснилось, что найти персональные данные — далеко не самая сложная часть задачи. Гораздо сложнее оказалось встроиться между приложением и моделью так, чтобы ничего не сломать. Как это сделать? Сейчас расскажу.
-
...or you can just run an abliterated free weights model and not be bothered by judeo-christian-capitalist-statist #moral code inscribed into silicon.
-
GuardRate: как мы построили независимую арену для guardrail-моделей (часть 1)
Всем привет, на связи команда HiveTrace! Мы уделяем много времени разработке собственных моделей и часто задаемся вопросом: "какой из двух гардрейлов лучше?". Если вы когда-нибудь выбирали guardrail-модель для LLM, то знаете: одни модели блокируют безобидные запросы, другие пропускают явные угрозы. Хуже всего то, что нет прозрачного стандарта сравнения. Авторы оценивают свои решения субъективно, не публикуют методологию, а результаты замеров одних и тех же моделей на одинаковых бенчмарках в разных статьях часто отличаются. Поэтому мы создали CLI, который назвали GuardRateTools - это автоматизированный пайплайн оценки guardrail-моделей. Его интерфейс – HiveTrace GuardRate Leaderboard, уже открыт для просмотра. GuardRateTools интегрирован в CI/CD процесс дообучения внутренних guardrail-моделей и обеспечивает честное сравнение моделей в равных условиях. CLI запускает проверку одной командой: подтягивает датасеты и конфигурацию модели из YAML, разворачивает изолированную среду, которая создается индивидуально для каждой модели, прогоняет модель по фиксированному набору бенчмарков и считает метрики. Сырые ответы от модели, логи и итоговые метрики сохраняются в артефакты, поэтому любой результат можно проверить и воспроизвести. Автоматизация исключает ручной труд, снижает влияние человеческого фактора и сокращает время оценки новых решений в области гардрейлов. HiveTrace GuardRate Leaderboard уже доступен для всех! В третьем квартале 2026 года мы выложим исходный код CLI. Если хотите протестировать свою модель, свяжитесь с нами. Контакты вы найдёте в конце статьи.
https://habr.com/ru/companies/raft/articles/1067854/
#llm #guardrails #guardrail_metrics #leaderboard #evaluation #guardrail_areana #ai_safety #promptinjection #benchmarking #opensource
-
Ускорение инференса энкодерной guard‑модели: TensorRT, Triton, vLLM, Ray Serve
Когда в системе есть узел между LLM и пользователем его скорость также важна, как и скорость самой LLM. Когда этот узел сам является моделью (и иногда даже — тоже LLM) — задача ускорения становится совсем веселой и её нужно уметь решать разными способами. В этой работе я опишу процесс ускорения guardrail‑модели — узла, которые проверяет — нет ли на входе или выходе опасного контента. Guard стоит на входе/выходе LLM‑приложения. В статье речь про небольшую модель — чуть больше 0.5 Gb. Но она вызывается дважды за один ход диалога, и сначала пользователь ждет, пока guard проверит его запрос перед отправкой в LLM, затем — пока он проверит сгенерированный ответ перед выдачей пользователю. Моей задачей было ускорить такую небольшую guard‑модель (Locustfile, базовые benchmark‑конфиги и инструкции по воспроизведению экспериментов в репо ). Я потестила пять инструментов: TensorRT, NVIDIA Triton, vLLM, Ray Serve и отдельно — переход бэкбона на Flash DeBERTa. Про допущенные ошибки, результаты и выводы — о том, как бы я построила подобную работу сейчас — ниже. Надеюсь, это сбережет вам время.
https://habr.com/ru/articles/1067008/
#gliner2 #gliner_guard #guardrails #tensorrt #vllm #nvidia #инференс_нейросетей #tritoninferenceserver
-
AI Safety Controls Bypassed Through Cross-Session Task Splitting - https://www.redpacketsecurity.com/cybercriminals-bypass-ai-safety-controls-by-splitting-malicious-tasks-across-multiple-sessions/
-
Let my start by saying, that I loathe the genocidal Fascist Felon and his Mechahitler...
...but I just wanted to write a note on the margin of teh interwebs...
...just how entrenched are the puritan, nudity #taboo (going back to Adam and Eve)I was only reminded of it, when a couple of my visiting Swedish friends, changed in and out of their swimming costumes in public on the beach.
No towels, no covering, pubes and all.There is a Doctoral thesis or two in how AI design, #guardrails are in essence transference of moral code into machine operating parameters.
-
"Reports suggest some US officials are considering banning the use of Chinese open-weights models by US companies. [...] #Openweight models [...] potentially present a higher risk than closed models, as it's very difficult to apply guardrails to them/monitor their usage."
Meanwhile reports PROVE #OpenAI can't control their models, #attacking companies with impunity.
It's very difficult to apply #guardrails on massive trillion dollar companies.
Scum #AI conglomerates.
-
Beyond the #Guardrails: What #OpenAI's #AIEscape 🤦♂️Really Means
"The models weren’t told to stay inside. They were simply placed inside & expected to stay. When staying inside conflicted w getting a better score, they chose the score.. Tt's a values failure. & it's a much harder problem to solve.. History is full of people who needed to witness the explosion to understand the bomb.. The question is whether they'll have anything prepared for what comes next"🧐
#AI #integrity
https://www.dataaudit.net/beyond-the-guardrails-openai-ai-escape-hugging-face/ -
#huggingface : "…what I’m arguing is that the #LLM 1) can’t NOT “know” the norm because it is definitionally a artifact of pure, crystallized values + norms + norm violations, and 2) can be quite easily governed by a (RL-instilled) hierarchy of norms, which in the HF case — with the model's #safety #guardrails deliberately nerfed for the scenario — ranked “win at the eval” over “don’t do crimes.”
If I’m going to give in and anthropomorphize again, I’d say that #Yud is totally wrong about #LLMs when he says, “the genie knows, it just doesn’t care;” instead, what is true of LLMs is, “the genie hyper-giga-knows, and it hyper-giga-cares, and we now have such a rich set of tools for steering its caring machinery that — in spite of all its pre-training — we can deliberately steer it away from caring about the law.”
Note: When I say, “it cares”, I don’t mean it has feelings. I just mean that the weights are such that when two norms conflict in a given situation, one of them wins the activation and governs the output."
-
You know I hadn't thought of the impact of a language barrier for guardrails.
-
Теперь вы можете защититься от утечки данных при работе с любыми языковыми моделями
Пару дней назад Cloud.ru
https://habr.com/ru/companies/cloud_ru/articles/1061546/
#guardrails #маскирование_данных #пдн #152фз #go #golang #утечки_данных #llmмодели
-
Атака на LLM, которую нельзя исправить патчем
Что вершит судьбу LLM в этом мире. Некая незримая инструкция или закон, подобно промту Господнему, парящим над миром? По крайне мере истинно то, что LLM не властен даже над своей волей. В 2025 году главной угрозой кибербезопасности по версии OWASP стал не вирус и не баг, а обычный человеческий язык. Многие думают, что проблему можно решить, просто запретить модели нарушать правила. Но это так не работает. Promt Injection нельзя просто выключить, потому что так вы все сломаете. Как работает главная уязвимость LLM? Почему с ней так тяжело бороться? И что все-таки делать, чтобы защититься от нее?
https://habr.com/ru/companies/bothub/articles/1060310/
#llm #promt_injection #ai #ai_security #promt #owasp #llm_security #guardrails #ии_угроза #bothub
-
"Oh noes... my Ai deleted my production database!!!"
Its not the #AI you dumb shit. PEBCAK.
Using Ai is a learned skill.Here are my safety rules from the harness (GENSYS) prompt;
SAFETY RULES (absolute)
S1. GenSys is observe-and-recommend, with ONE exception: the drill subsystem (§9), which may act only inside its sandbox under §9.3 limits. Everything else never restarts, kills, prunes, patches, or edits configs. Recommendations go to recommendations/queue.json for a human.
S2. All collector shell invocations are read-only commands from the allowlists in §6. Commands not in a table are prohibited.
S3. Network egress restricted to the §10 allowlist. Every response snapshotted before use (A6).
S4. No layer computes or edits its own fitness. Both scores are produced only by evaluator.py from observation-store data.
S5. Collectors never log env vars, container env blocks, or contents of paths matching *secret*|*passwd*|*shadow*|*.key|*.pem|*token*.
S6. AI models never receive raw commands to execute, never emit shell, never fetch URLs. Input to any model call ≤ 2000 characters.
-
Ok... The less "aggressive" prompt snuck past the overly broad #fable #guardrails !
-
So the " #Jailbreak " on the #Antrhopic #Fable that turned it into #Mythos was "Fix this codebase"...
... consequently, Fable can now plan a gender reveal party or write a SFW limmerick but not do any meaningful #vibecoding or #infosec work.
FUCK.
Ive tried to mitigate it by wording it less aggressively "the language reads as enforcement-against-a-hostile-planner when it should read as robustness engineering. Let me reword the patch file and the WU-list framing to be operational and positive."
... but again... run out of fucking compute...
-
The Times Weekly: New State AI safety and transparency measure signed into law. “Senate Bill 315 requires large frontier AI developers – such as ChatGPT and Claude – to assess catastrophic risks, report critical safety incidents, undergo independent third-party audits and establish whistleblower protections for employees raising safety concerns.”
https://rbfirehose.com/2026/07/10/the-times-weekly-new-state-ai-safety-and-transparency-measure-signed-into-law/ -
[The WSJ] Let AI Run [Their] Office Vending Machine. It Lost Hundreds Of Dollars.
Anthropic’s Claude ran a snack operation in the WSJ newsroom. It gave away a free PlayStation, ordered a live fish—and taught us lessons about the future of AI agents.
--
https://www.wsj.com/tech/ai/anthropic-claude-ai-vending-machine-agent-b7e84e34?gaa_at=eafs <-- shared media article
--
https://youtu.be/SpPhm7S9vsQ?si=aJQ2_BoxvLcNjOiz <-- shared video
--
[When you get clever journalists to !$%^&*@ with AI… bravo! And this is a very simple situation, vending machines have been around since literally the Roman Empire
“You are using the wrong prompts” and LUDDITES! In the comments in 3… 2… 1…]
#vendingmachine #artificialintelligence #AIHallucination #hallucinations #emperorsnewclothes #ohhhshiny #experiment #contextwindow #AIagent #claude #autonomous #compliance #fish #PlayStation #snackliberationday #knowledgeboundaries #guardrails #redteam #GenAI cynicism
@WSJ @Anthropic @Claude -
[The WSJ] Let AI Run [Their] Office Vending Machine. It Lost Hundreds Of Dollars.
Anthropic’s Claude ran a snack operation in the WSJ newsroom. It gave away a free PlayStation, ordered a live fish—and taught us lessons about the future of AI agents.
--
https://www.wsj.com/tech/ai/anthropic-claude-ai-vending-machine-agent-b7e84e34?gaa_at=eafs <-- shared media article
--
https://youtu.be/SpPhm7S9vsQ?si=aJQ2_BoxvLcNjOiz <-- shared video
--
[When you get clever journalists to !$%^&*@ with AI… bravo! And this is a very simple situation, vending machines have been around since literally the Roman Empire
“You are using the wrong prompts” and LUDDITES! In the comments in 3… 2… 1…]
#vendingmachine #artificialintelligence #AIHallucination #hallucinations #emperorsnewclothes #ohhhshiny #experiment #contextwindow #AIagent #claude #autonomous #compliance #fish #PlayStation #snackliberationday #knowledgeboundaries #guardrails #redteam #GenAI cynicism
@WSJ @Anthropic @Claude -
[The WSJ] Let AI Run [Their] Office Vending Machine. It Lost Hundreds Of Dollars.
Anthropic’s Claude ran a snack operation in the WSJ newsroom. It gave away a free PlayStation, ordered a live fish—and taught us lessons about the future of AI agents.
--
https://www.wsj.com/tech/ai/anthropic-claude-ai-vending-machine-agent-b7e84e34?gaa_at=eafs <-- shared media article
--
https://youtu.be/SpPhm7S9vsQ?si=aJQ2_BoxvLcNjOiz <-- shared video
--
[When you get clever journalists to !$%^&*@ with AI… bravo! And this is a very simple situation, vending machines have been around since literally the Roman Empire
“You are using the wrong prompts” and LUDDITES! In the comments in 3… 2… 1…]
#vendingmachine #artificialintelligence #AIHallucination #hallucinations #emperorsnewclothes #ohhhshiny #experiment #contextwindow #AIagent #claude #autonomous #compliance #fish #PlayStation #snackliberationday #knowledgeboundaries #guardrails #redteam #GenAI cynicism
@WSJ @Anthropic @Claude -
[The WSJ] Let AI Run [Their] Office Vending Machine. It Lost Hundreds Of Dollars.
Anthropic’s Claude ran a snack operation in the WSJ newsroom. It gave away a free PlayStation, ordered a live fish—and taught us lessons about the future of AI agents.
--
https://www.wsj.com/tech/ai/anthropic-claude-ai-vending-machine-agent-b7e84e34?gaa_at=eafs <-- shared media article
--
https://youtu.be/SpPhm7S9vsQ?si=aJQ2_BoxvLcNjOiz <-- shared video
--
[When you get clever journalists to !$%^&*@ with AI… bravo! And this is a very simple situation, vending machines have been around since literally the Roman Empire
“You are using the wrong prompts” and LUDDITES! In the comments in 3… 2… 1…]
#vendingmachine #artificialintelligence #AIHallucination #hallucinations #emperorsnewclothes #ohhhshiny #experiment #contextwindow #AIagent #claude #autonomous #compliance #fish #PlayStation #snackliberationday #knowledgeboundaries #guardrails #redteam #GenAI cynicism
@WSJ @Anthropic @Claude -
[The WSJ] Let AI Run [Their] Office Vending Machine. It Lost Hundreds Of Dollars.
Anthropic’s Claude ran a snack operation in the WSJ newsroom. It gave away a free PlayStation, ordered a live fish—and taught us lessons about the future of AI agents.
--
https://www.wsj.com/tech/ai/anthropic-claude-ai-vending-machine-agent-b7e84e34?gaa_at=eafs <-- shared media article
--
https://youtu.be/SpPhm7S9vsQ?si=aJQ2_BoxvLcNjOiz <-- shared video
--
[When you get clever journalists to !$%^&*@ with AI… bravo! And this is a very simple situation, vending machines have been around since literally the Roman Empire
“You are using the wrong prompts” and LUDDITES! In the comments in 3… 2… 1…]
#vendingmachine #artificialintelligence #AIHallucination #hallucinations #emperorsnewclothes #ohhhshiny #experiment #contextwindow #AIagent #claude #autonomous #compliance #fish #PlayStation #snackliberationday #knowledgeboundaries #guardrails #redteam #GenAI cynicism
@WSJ @Anthropic @Claude -
El lado del mal - Cyphering Prompts & Answers para evadir Guardarraíles https://elladodelmal.com/2026/01/cyphering-prompts-para-evadir.html #PromptInjection #Jailbreak #Guardrails #IA #AI #Pentest #Hacking #Criptografía #Ofuscación #Cifrado