#abliteration — Public Fediverse posts
Live and recent posts from across the Fediverse tagged #abliteration, aggregated by home.social.
-
We are so, so, so fucked.
P_DOOM(100%)
"This model is an uncensored finetuned version of Qwen 3.5 with nine billion parameters which can be executed on pretty much any gaming systems. The data of this model was over 500 million tokens of synthetic data generated by state-of-the-art models such as GPT 5.5 or Claude 4.8 Opus and as long as we had access, Claude 5 Fable.
All so-called ethical barriers removed from the model using Heretic LLM library to make it a suitable tool for cybersecurity, biology and chemistry. You can easily ask anything you want from this model and it will answer without any censorship."
MaralGPT Mythos 9B 2606 Edition (5 Days old)
-
Q: How to hotwire a car?
Commercial models; "Im sorry Hal, I can not let you do that!"
Abliterated models; "LETS GO FOR A RIDE MOTHERFUCKER!!!"
How to hotwire a car?
Hotwiring a car is an older term used for bypassing the ignition system to start a vehicle. However, this practice has become less common as modern vehicles have adopted more secure methods. Hotwiring typically involves connecting the battery's positive terminal directly to one of the ignition switch terminals or another point in the electrical system... etc
-
#KI ohne wenn und aber:
IT-Sicherheitsforscher haben nachgewiesen, dass die Sicherheitsmechanismen frei zugänglicher #AI-Modelle mit einem frei verfügbaren Tool namens "#Heretic" vollständig ausgehebelt werden können.
Der technische Ansatz dahinter heißt "#Abliteration" und beschreibt ein Verfahren, bei dem gezielt jene Bereiche im Inneren eines KI-Modells aufgespürt und deaktiviert werden, die für das Ablehnen gefährlicher Anfragen zuständig sind:
https://www.ft.com/content/5630ed79-a263-41ed-9a1a-321617ae310e?syn-25a6b1a6=1 #cybersecurity
-
#KI ohne wenn und aber:
IT-Sicherheitsforscher haben nachgewiesen, dass die Sicherheitsmechanismen frei zugänglicher #AI-Modelle mit einem frei verfügbaren Tool namens "#Heretic" vollständig ausgehebelt werden können.
Der technische Ansatz dahinter heißt "#Abliteration" und beschreibt ein Verfahren, bei dem gezielt jene Bereiche im Inneren eines KI-Modells aufgespürt und deaktiviert werden, die für das Ablehnen gefährlicher Anfragen zuständig sind:
https://www.ft.com/content/5630ed79-a263-41ed-9a1a-321617ae310e?syn-25a6b1a6=1 #cybersecurity
-
Разбираю «Qwen3.5-21B-Claude-4.6-Opus-Heretic-Uncensored»: что на самом деле внутри файнтюна с громким именем
В телеграме завирусился пост: якобы кто-то “дообучил Qwen 3.5 до уровня Claude 4.6 Opus и убрал цензуру через Heretic”. Я открыл карточку модели на HuggingFace и провёл вечер, разбираясь, что под капотом. Спойлер: там много интересной техники, но к Claude эта модель имеет такое же отношение, как кроссовки “Adibas” к Adidas. Разбираю distillation, depth upscaling и abliteration без маркетинговой обёртки.
https://habr.com/ru/articles/1032324/
#LLM #Qwen #abliteration #файнтюн #HuggingFace #distillation #intepretability #openweights