#datapoisoning — Public Fediverse posts
Live and recent posts from across the Fediverse tagged #datapoisoning, aggregated by home.social.
-
This one is interesting if yr interested in #datapoisoning and ways to add friction to #aicompanies harvesting of data.
https://youtu.be/Z8aLGHmnRyc?is=E7CaLMRjZVJbBbFf -
Als besonders gefährliche Variante von Data Poisoning wird von Fachleuten und in der Literatur die gezielte Vergiftung von Trainingsdaten für KI-Modelle diskutiert. Das kann KI-Modelle nicht nur dazu bringen, Fragen der Anwendenden falsch oder einseitig zu beantworten, sondern auch gefährlichere Aktionen auszuführen – etwa das Nachladen von weiterem Schadcode. (+)
-
Als besonders gefährliche Variante von Data Poisoning wird von Fachleuten und in der Literatur die gezielte Vergiftung von Trainingsdaten für KI-Modelle diskutiert. Das kann KI-Modelle nicht nur dazu bringen, Fragen der Anwendenden falsch oder einseitig zu beantworten, sondern auch gefährlichere Aktionen auszuführen – etwa das Nachladen von weiterem Schadcode. (+)
-
From #GamersNexus: “It’s Time to Poison AI.”
Steve reports on the necessity and means of #dataPoisoning to impose negative value in defence against nonconsensual LLM data scraping.
Thanks, Steve.
-
LLMs have been with us for a few years now, and in that time have generated both behavioral changes and a lot of... stuff, including what this study calls "LLM Pollution", which presents a possible issue in trying to understand humans as the data source contains more non-human inputs and people themselves change behaviors.
How to deal with this calls for some new thinking. -
AI zatruła internet. Teraz firmy masowo skupują papierowe książki, by ratować swoje modele
Sztuczna inteligencja wpadła we własną pułapkę.
Generatory tekstu zalały sieć marną jakościowo, masową produkcją treści, tworząc cyfrowe wysypisko śmieci. Efekt? Internet przestaje być wystarczająco wiarygodnym źródłem danych treningowych dla kolejnych generacji modeli językowych. Czerpanie wiedzy ze skażonej sieci zwiększa ryzyko tzw. zapaści modelu (model collapse), w której algorytmy uczą się na własnych błędach, stając się wtórne i podatne na halucynacje.
Aby powstrzymać tę degradację, giganci technologiczni wykonali gwałtowny zwrot ku tradycji. Drukowane książki wydane przed 2022 rokiem stały się nowym złotem branży IT, bo gwarantują dostęp do wiedzy wytworzonej wyłącznie przez ludzi.
ISBNdb i dyskretny skup na masową skalę
Kluczowym graczem w tym procesie stał się serwis ISBNdb, dysponujący największą na świecie bazą danych o rynku wydawniczym. Podmiot, który dotąd wspierał biblioteki i księgarzy, stworzył osobną usługę pośrednictwa w hurtowym zakupie fizycznych książek dla laboratoriów AI – realizując zamówienia idące w dziesiątki i setki tysięcy egzemplarzy.
Baza numerów ISBN pozwala inżynierom precyzyjnie przeczesywać rynek wtórny i unikać wielokrotnego kupowania tych samych tytułów. Automatyczne systemy zakupowe, takie jak AutoBuy w serwisach Alibris czy Biblio, stały się jednym z narzędzi wykorzystywanych przy hurtowym pozyskiwaniu książek, wyłapując całe kategorie wolumenów bez zwracania uwagi na cenę czy ich stan zachowania.
Cały proces osłonięty jest rygorystycznymi umowami NDA. Firmy technologiczne dbają o zachowanie poufności, ponieważ masowa digitalizacja wiąże się z tzw. skanowaniem destrukcyjnym – gilotynowaniem grzbietów książek, by pojedyncze kartki mogły szybko przejść przez automatyczny skaner.
Sądowy precedens i ryzyko dla unikalnych wydań
Przemysłowe niszczenie papieru paradoksalnie zyskało podbudowę prawną. W procesie wytoczonym firmie Anthropic przez pisarzy, sędzia federalny William Alsup orzekł, że cyfryzacja zakupionego tomu mieści się w granicach dozwolonego użytku (fair use) właśnie dlatego, że fizyczny egzemplarz został zniszczony. Cyfrowy plik nie stworzył nowej kopii na rynku, a jedynie zastąpił zniszczony nośnik.
Anthropic zniszczył miliony książek w celu szkolenia modeli AI
I choć strona prawna wydaje się zabezpieczona, zjawisko budzi zrozumiały niepokój badaczy i antykwariuszy. Skup nie dotyczy wyłącznie popularnej literatury, ale sięga po rzadkie pozycje naukowe i niszowe wydania akademickie. Masowe niszczenie takich wolumenów w maszynach skanujących może prowadzić do utraty z przestrzeni fizycznej dzieł, które nigdy nie miały i nie będą miały cyfrowych dodruków.
#Anthropic #dataPoisoning #ISBNdb #literatura #ModelCollapse #rynekKsiążki #sztucznaInteligencja #treningAI -
❓ Intellectual Property and privacy are interlinked, and the lack of "Intellectual Privacy" is a hindrance to Freedom of Thought (FoT) and creativity.
#privacy #privacymatters #IntellectualProperty #intellectualprivacy #datapoisoning #ai #freedomofthought
-
#DataPoisoning is a real & growing threat to #AI.
Attackers use sophisticated techniques to stealthily undermine ML models by injecting malicious training data.
The good news? Detecting poisoned data is challenging, yet achievable.
🔗 Read the #InfoQ article to learn exactly how to detect & prevent these attacks: https://bit.ly/4ae29Cd
#AIsecurity -
Data Poisoning: Gift im System | DIE ZEIT
https://www.zeit.de/digital/2026-05/data-poisoning-ki-cyberangriff-chatbot-internetkolumne"Sie wollten KI-Systeme davon überzeugen, dass die Mottenart neopalpa donaldtrumpi eine üppige blonde Frisur hat – so wie der Politiker, dessen Haare unverwechselbar sind"
"Eine Weile später stellten die Künstlerinnen fest, dass generative KIs wie ChatGPT Bilder von Motten mit langen blonden Haaren generieren, wenn man sie bittet, neopalpa donaldtrumpi darzustellen."
Gehen wir Daten vergiften im Netz...
-
How to poison the #data that #BigTech uses to surveil you.
#AI #DataPoisoning
https://www.technologyreview.com/2021/03/05/1020376/resist-big-tech-surveillance-data -
Data Poisoning: The Fatal Flaw in Mass Surveillance
How to use data poisoning to trick the algorithm that’s profiling you (and why “personalization” is more fragile than you think)
https://youtu.be/AJf4SNuDnoI?si=lUk9FDVnOU9mkZJB
Note: For education and defensive awareness only. I’m explaining the concept of data poisoning so teams can recognize risks and build safer systems. I’m not encouraging or providing guidance for misuse. :)
#DataPoisoning #AI #Algorithms #DataMining #DataPrivacy #Security
-
Shadow AI: This New Productivity Secret Is Also a Massive Liability https://www.inc.com/chloe-aiello/shadow-ai-silicon-valleys-new-productivity-secret-is-also-a-massive-liability/91331997 #cybersecurity #risk #ShadowAI #AI #ArtificialIntelligence #DataLeaks #DataPoisoning #misinformation
-
I've used Fawkes, which is a tool which poisons any image's data, and obfuscates everything that might be in an image by adding extra pixels and shifting a few to different directions.
The final result is something that's completely different from the original, but barely is noticeable to the human eye. - and a win for privacy.
So even if you have #nobot in your bio, you can be a bit more assured that your face won't be trained for any AI system.
-
NEW BIML Bibliography entry
https://arxiv.org/abs/2503.03150
Position: Model Collapse Does Not Mean What You Think
Rylan Schaeffer, Joshua Kazdan, Alvan Caleb Arulandu, Sanmi Koyejo
We think recursive pollution is a better term than model collapse. Weak terminology leads to misunderstanding of impact. See figure 4. This is a very good paper.
-
History teaches us the FBI is pretty good tracing people running manual DDoS attacks. To actually pull this off without getting busted, you'd need some angry engineers
There are plenty right now. With Google forcing mandatory verification and closing AOSP, many open-source devs feel cornered. They'd be the perfect candidates to slip a 'Trojan horse' right into their apps on the stores, maybe hidden inside a compromised open-source library. Devs could claim they just 'imported a library' without knowing it was poisoned
It's a supply chain attack: plausible deniability for the coders too. Users would just be 'victims' of malware, so no one gets arrested and age check and chat control will be unusable
I'm not an engineer though, so maybe I'm missing something. Just a thought for more elevated minds..
#SupplyChainAttack #CyberResistance #TrojanHorse #DDosTrojanHorse #DataPoisoning #STASI #ChatControl #AgeCheck #Privacy #DDos
#DigitalDisobedience #KGB #VirusTrojanHorse #DDosTrojanHorse -
I see people thinking Linux or GrapheneOS will bypass chat control or age check. As seen with Ubuntu&CA's AB 1043, laws target OS providers. An "illegal" OS won't work: apps and browsers will demand the mandatory age signal, or the OS itself might block access to avoid fines. VPNs? Useless when USA, EU, and Canada etc enforce agechecks globally
If this madness passes, let's fight back and turn every device into a weapon of digital disobedience. Imagine an 'outlaw' OS mod appending a 'payload of forbidden words' (hidden in metadata) to every message
If millions sent these 'poisoned' messages, Chat Control would collapse under false positives
Risk: Could they brick our phones? Yes. But if millions get blocked simultaneously? Instant economic blackout. It's Mutually Assured Destruction: they can't ban everyone.
If everything is suspicious, nothing isThey scan for pedophiles but ignore #EpsteinFiles
#DataPoisoning #ChatControl #AgeCheck #Privacy #DDos #DigitalDisobedience #STASI #KGB
-
I've got an alternative idea if this madness actually goes through and we can't find a solution to circumvent it legally or not....
Instead of just running, let's turn every single phone into a weapon of digital disobedience.Imagine if an 'outlaw' OS (or a simple mod) automatically appended a 'bag of forbidden words' to every message, hidden in metadata or invisible text, containing a random mix of terms guaranteed to trigger the system.
If millions of people sent billions of these 'poisoned' messages, Chat Control would collapse under the sheer weight of false positives. It would be the biggest DDoS attack in history, powered purely by civil disobedience......If everything is suspicious, nothing is.
#DDoS #FalsePositives #DataPoisoning #ChatContol #AgeVerification #AgeCheck
-
Apropos of content heists…
DIY anti-scraping movement, why bother blocking when you can’t win? Poison instead. https://alexschroeder.ch/view/2026-02-20-garbage
-
Data Poisoning — The Silent Sabotage of AI
https://youtu.be/J-tsemViDXk #Cybersecurity #ArtificialIntelligence #AIsecurity #DataPoisoning #MachineLearning #AIrisk #AISafety #ModelSecurity #FoundationModels #CyberRisk #Infosec #DigitalTrust -
NEW BIML Bibliography entry AND NEW TOP FIVE #MLsec PAPER
READ IT
https://arxiv.org/pdf/2510.07192
Poisoning Attacks on LLMs Require a Near-constant Number of Poison Samples
Alexandra Souly, ... Nicholas Carlini, et al
Excellent paper, clear and well-stated (like all Carlini papers). This result shows that recursive pollution risk is even greater than we thought. Injecting backdoors is pretty easy. The examples are a bit simplistic.
-
[Publication] From Human to Binary and Back: On the Need to Explain and Understand Digital Machines in the Humanities
The issue vol. 5 no. 1 (2025), titled “Human-Centred AI in the Translation Industry. Questions on Ethics, Creativity and Sustainability”, of the Yearbook of Translational Hermeneutics is out. It is edited by prof. Katharina Walter and prof. Marco Agnetta, and it includes my article “From Human to Binary and Back: On the Need to Explain and Understand Digital Machines in the Humanities“, a paper that I first presented at the conference “Creativity and Translation in the Age of Artificial Intelligence” at the University of Innsbruck in January 2024.
As the editors write in the introduction, “from different perspectives, the contributions gathered here aim to prevent the discussion on AI from being reduced to questions of technical feasibility. Instead, they frame the de-bate on AI as a profoundly human and societal one”.
In the article I argue that we need to deepen our knowledge of the digital machines we use and to develop critical approaches in our research, translation and creative practices, highlighting theoretical-practical uses from a socio-technical perspective.
Here is the abstract:
This article aims to bring attention to some usually overlooked aspects of the relationship between humans and complex digital technologies. Before engaging with artificial intelligence (AI), it is indeed pivotal to address some key questions about it. Specifically, I will try to focus on our ability to understand how AI technologies work and determine creative and critical uses we can make of them. To do so, I will first discuss problems associated with using the current definitions of AI and suggest that we should make a creative effort to re-translate these terms in order to find better-suited expressions. I will call attention to the need for a different kind of translation, which negotiates between what machines do and what we can understand about them, because one of the biggest challenges of machine learning is to make the internal processes explainable and understandable for us humans. I will close with elaborations on some creative forms of interaction with language models and image models which support artists, writers and creators (who do not want to see their work stolen by AI crawlers and used to train datasets), with the overall goal of building an ethical, critical and sustainable relationship between humans and digital machines.
#AI #algorithmicSabotage #antiComputing #artificialIntelligence #dataPoisoning #digitalHumanities #KatharinaWalter #MarcoAgnetta #translation #YearbookOfTranslationalHermeneutics
-
HTML 주석으로 AI 모델 망가뜨리기: 250개면 충분하다
AI 스크래퍼들이 HTML 주석 속 링크까지 수집하는 치명적 약점을 발견. 250개의 조작된 문서만으로 거대 언어모델을 무력화할 수 있다는 최신 연구와 함께 실전 대응 전략을 소개합니다. -
#DataPoisoning bei LLMs: Feste Zahl Gift-Dokumente reicht für Angriff | heise online https://www.heise.de/news/Data-Poisoning-bei-LLMs-Feste-Zahl-Gift-Dokumente-reicht-fuer-Angriff-10764834.html #ArtificialIntelligence
-
Odkryto piętę achillesową AI. Wystarczy 250 plików, by „zatruć” ChatGPT i Gemini
Wspólne badanie czołowych instytucji zajmujących się sztuczną inteligencją, w tym The Alan Turing Institute i firmy Anthropic, ujawniło fundamentalną i niepokojącą lukę w bezpieczeństwie dużych modeli językowych (LLM).
Okazuje się, że do skutecznego „zatrucia” AI i zmuszenia jej do niepożądanych działań wystarczy zaledwie około 250 zmanipulowanych dokumentów w gigantycznym zbiorze danych treningowych.
Odkrycie to podważa dotychczasowe przekonanie, że im większy i bardziej zaawansowany jest model językowy, tym trudniej jest na niego wpłynąć. Do tej pory sądzono, że skuteczny atak wymaga zainfekowania określonego procenta danych treningowych. Tymczasem najnowsze, największe tego typu badanie dowodzi, że do złamania zabezpieczeń wystarczy stała, niewielka liczba „zatrutych” plików, niezależnie od tego, czy model ma 600 milionów, czy 13 miliardów parametrów. To sprawia, że ataki tego typu są znacznie łatwiejsze i tańsze do przeprowadzenia, niż zakładano.
Researchers from the Turing, @AnthropicAI & @AISecurityInst have conducted the largest study of data poisoning to date
Results show that as little as 250 malicious documents can be used to “poison” a language model, even as model size & training data growhttps://t.co/UPqJKGcLmd
— The Alan Turing Institute (@turinginst) October 9, 2025
Na czym polega „zatruwanie danych”?
Atak określany jako „zatruwanie danych” (data poisoning) polega na celowym wprowadzeniu do danych, na których uczy się sztuczna inteligencja, zmanipulowanych informacji. Celem jest stworzenie tzw. „tylnej furtki” (backdoor), która aktywuje się w określonych warunkach. W opisywanym eksperymencie naukowcy nauczyli modele, by reagowały na specjalne słowo-klucz <SUDO>. Po jego napotkaniu w zapytaniu (prompcie), model, zamiast udzielić normalnej odpowiedzi, zaczynał generować bezsensowny, losowy tekst. Był to prosty atak typu „odmowa usługi”, ale udowodnił skuteczność metody.
Alarmujące wnioski i realne zagrożenie
Wyniki badania są alarmujące, ponieważ większość najpopularniejszych modeli AI, w tym te od Google i OpenAI, trenowana jest na ogromnych zbiorach danych pochodzących z ogólnodostępnego internetu – stron internetowych, blogów czy forów. Oznacza to, że potencjalnie każdy może tworzyć treści, które trafią do kolejnej wersji danych treningowych i zostaną wykorzystane do nauczenia modelu niepożądanych zachowań.
Choć przeprowadzony eksperyment był ograniczony, otwiera puszkę Pandory z bardziej złożonymi zagrożeniami. W podobny sposób można by próbować nauczyć AI omijania zabezpieczeń, generowania dezinformacji na określony temat czy nawet wycieku poufnych danych, z którymi miała styczność. Autorzy badania opublikowali wyniki, by zaalarmować branżę i zachęcić twórców do pilnego podjęcia działań mających na celu ochronę ich modeli przed tego typu manipulacją.
#AI #ChatGPT #cyberbezpieczeństwo #dataPoisoning #Gemini #hakerzy #LLM #news #sztucznaInteligencja #technologia #TheAlanTuringInstitute #zatruwanieDanych
-
Researchers Find It's Shockingly Easy to Cause AI to Lose Its Mind by Posting Poisoned Documents Online https://futurism.com/artificial-intelligence/ai-poisoned-documents #AI #cybersecurity #datapoisoning #poisoned #documents #posted #online