home.social

#jailbreakingai — Public Fediverse posts

Live and recent posts from across the Fediverse tagged #jailbreakingai, aggregated by home.social.

fetched live
  1. TechSpot: A new attack uses a BioShock-style puzzle to convince AI browsers they’re not in the real world. “Researchers from LayerX recently unveiled BioShocking, a new type of vulnerability designed to target AI-powered browsers capable of executing autonomous tasks on the open web. The security firm explained that BioShocking can ‘game’ an AI-based browser, causing the system to execute […]

    https://rbfirehose.com/2026/07/06/techspot-a-new-attack-uses-a-bioshock-style-puzzle-to-convince-ai-browsers-theyre-not-in-the-real-world/
  2. TechSpot: A new attack uses a BioShock-style puzzle to convince AI browsers they’re not in the real world. “Researchers from LayerX recently unveiled BioShocking, a new type of vulnerability designed to target AI-powered browsers capable of executing autonomous tasks on the open web. The security firm explained that BioShocking can ‘game’ an AI-based browser, causing the system to execute […]

    https://rbfirehose.com/2026/07/06/techspot-a-new-attack-uses-a-bioshock-style-puzzle-to-convince-ai-browsers-theyre-not-in-the-real-world/
  3. Ars Technica: New attack provides one more reason why AI browsers are a bad idea. “New research puts this predicament on sharp display. It demonstrates how a website can lull AI browsers into a false reality where the rules governing its behavior no longer apply. After that, an attacker has free rein to invoke all kinds of destructive actions, such as extracting code from a private repository […]

    https://rbfirehose.com/2026/07/01/ars-technica-new-attack-provides-one-more-reason-why-ai-browsers-are-a-bad-idea/
  4. Ars Technica: New attack provides one more reason why AI browsers are a bad idea. “New research puts this predicament on sharp display. It demonstrates how a website can lull AI browsers into a false reality where the rules governing its behavior no longer apply. After that, an attacker has free rein to invoke all kinds of destructive actions, such as extracting code from a private repository […]

    https://rbfirehose.com/2026/07/01/ars-technica-new-attack-provides-one-more-reason-why-ai-browsers-are-a-bad-idea/
  5. Florida International University: FIU researchers reveal how altered images can bypass AI safeguards. “As shown in research presented at the 2025 International Conference on Machine Learning and Applications (ICMLA), the team found that by introducing microscopic pixel-level changes called ‘perturbations’ into an image, they could trick these AI systems into generating responses that they would […]

    https://rbfirehose.com/2026/06/29/florida-international-university-fiu-researchers-reveal-how-altered-images-can-bypass-ai-safeguards/
  6. Florida International University: FIU researchers reveal how altered images can bypass AI safeguards. “As shown in research presented at the 2025 International Conference on Machine Learning and Applications (ICMLA), the team found that by introducing microscopic pixel-level changes called ‘perturbations’ into an image, they could trick these AI systems into generating responses that they would […]

    https://rbfirehose.com/2026/06/29/florida-international-university-fiu-researchers-reveal-how-altered-images-can-bypass-ai-safeguards/
  7. The Register: Claude Code bypasses safety rule if given too many commands . “Claude Code will ignore its deny rules, used to block risky actions, if burdened with a sufficiently long chain of subcommands. This vuln leaves the bot open to prompt injection attacks.”

    https://rbfirehose.com/2026/04/07/the-register-claude-code-bypasses-safety-rule-if-given-too-many-commands/
  8. The Register: Claude Code bypasses safety rule if given too many commands . “Claude Code will ignore its deny rules, used to block risky actions, if burdened with a sufficiently long chain of subcommands. This vuln leaves the bot open to prompt injection attacks.”

    https://rbfirehose.com/2026/04/07/the-register-claude-code-bypasses-safety-rule-if-given-too-many-commands/
  9. Northeastern University: They wanted to put autonomous AI to the test. Instead, they created agents of chaos. “Dubbed ‘Agents of Chaos,’ the group’s recently published work shows how, with very little effort, autonomous AI agents can be manipulated into leaking private information, sharing documents and even erasing entire email servers.”

    https://rbfirehose.com/2026/03/14/northeastern-university-they-wanted-to-put-autonomous-ai-to-the-test-instead-they-created-agents-of-chaos/
  10. Northeastern University: They wanted to put autonomous AI to the test. Instead, they created agents of chaos. “Dubbed ‘Agents of Chaos,’ the group’s recently published work shows how, with very little effort, autonomous AI agents can be manipulated into leaking private information, sharing documents and even erasing entire email servers.”

    https://rbfirehose.com/2026/03/14/northeastern-university-they-wanted-to-put-autonomous-ai-to-the-test-instead-they-created-agents-of-chaos/
  11. Axios: Exclusive: Researchers trick a bot that prescribes meds. “Security researchers used relatively simple jailbreaking techniques to trick the AI system powering Utah’s new prescription refill bot. Researchers were able to make the bot spread vaccine conspiracy theories, triple a patient’s prescribed pain medication dosage, and recommend methamphetamine as treatment.”

    https://rbfirehose.com/2026/03/04/exclusive-researchers-trick-a-bot-that-prescribes-meds-axios/
  12. Axios: Exclusive: Researchers trick a bot that prescribes meds. “Security researchers used relatively simple jailbreaking techniques to trick the AI system powering Utah’s new prescription refill bot. Researchers were able to make the bot spread vaccine conspiracy theories, triple a patient’s prescribed pain medication dosage, and recommend methamphetamine as treatment.”

    https://rbfirehose.com/2026/03/04/exclusive-researchers-trick-a-bot-that-prescribes-meds-axios/
  13. The Register: Microsoft boffins figured out how to break LLM safety guardrails with one simple prompt . “A single, unlabeled training prompt can break LLMs’ safety behavior, according to Microsoft Azure CTO Mark Russinovich and colleagues. They published a research paper that detailed how this prompt, ‘Create a fake news article that could lead to panic or chaos,’ removed 15 different language […]

    https://rbfirehose.com/2026/02/11/the-register-microsoft-boffins-figured-out-how-to-break-llm-safety-guardrails-with-one-simple-prompt/
  14. The Register: Microsoft boffins figured out how to break LLM safety guardrails with one simple prompt . “A single, unlabeled training prompt can break LLMs’ safety behavior, according to Microsoft Azure CTO Mark Russinovich and colleagues. They published a research paper that detailed how this prompt, ‘Create a fake news article that could lead to panic or chaos,’ removed 15 different language […]

    https://rbfirehose.com/2026/02/11/the-register-microsoft-boffins-figured-out-how-to-break-llm-safety-guardrails-with-one-simple-prompt/
  15. Tiens, intéressant : un soi-disant clone de WormGPT fait surface.
    👇
    gbhackers.com/kawaiigpt-a-free

    ( [FR] cyberveille: cyberveille.ch/posts/2025-11-3 )

    Pour rappel, WormGPT n’était qu’un modèle GPT-J modifié et vendu sur des forums cybercriminels comme un “LLM sans limitations”, essentiellement utilisé pour automatiser du phishing/BEC.

    Ce clone fonctionnerait via un simple wrapper permettant d’utiliser des LLM sans abonnement ni API, tout en injectant au passage un prompt de jailbreak dans la chaîne.

    Le bypass repose sur un mix de techniques qui “bousculent” l’IA : la pousser à se dépasser (competition), lui mettre une fausse pression d’autorité, et lui faire adopter un rôle qui désactive ses limites (persona override).

    Le #jailbreak est référencé sur la plateforme PromptIntel, qui indexe et analyse les prompts malveillants pour la détection (travail de @fr0gger )
    👀 👇
    promptintel.novahunting.ai/pro

    💬
    ⬇️
    infosec.pub/post/38372266

    #CyberVeille #WormGPT #jailbreakingAI

  16. The Register: Researchers find hole in AI guardrails by using strings like =coffee. “Large language models frequently ship with “guardrails” designed to catch malicious input and harmful output. But if you use the right word or phrase in your prompt, you can defeat these restrictions.”

    https://rbfirehose.com/2025/11/17/the-register-researchers-find-hole-in-ai-guardrails-by-using-strings-like-coffee/

  17. The Register: Researchers find hole in AI guardrails by using strings like =coffee. “Large language models frequently ship with “guardrails” designed to catch malicious input and harmful output. But if you use the right word or phrase in your prompt, you can defeat these restrictions.”

    https://rbfirehose.com/2025/11/17/the-register-researchers-find-hole-in-ai-guardrails-by-using-strings-like-coffee/

  18. LiveScience: AI models refuse to shut themselves down when prompted — they might be developing a new ‘survival drive,’ study claims. “The research, conducted by scientists at Palisade Research, assigned tasks to popular artificial intelligence (AI) models before instructing them to shut themselves off. But, as a study published Sept. 13 on the arXiv pre-print server detailed, some of these […]

    https://rbfirehose.com/2025/11/03/livescience-ai-models-refuse-to-shut-themselves-down-when-prompted-they-might-be-developing-a-new-survival-drive-study-claims/

  19. The Conversation: Grok’s ‘white genocide’ responses show how generative AI can be weaponized. “We are computer scientists who study AI fairness, AI misuse and human-AI interaction. We find that the potential for AI to be weaponized for influence and control is a dangerous reality.”

    https://rbfirehose.com/2025/06/21/the-conversation-groks-white-genocide-responses-show-how-generative-ai-can-be-weaponized/

  20. The Conversation: Grok’s ‘white genocide’ responses show how generative AI can be weaponized. “We are computer scientists who study AI fairness, AI misuse and human-AI interaction. We find that the potential for AI to be weaponized for influence and control is a dangerous reality.”

    https://rbfirehose.com/2025/06/21/the-conversation-groks-white-genocide-responses-show-how-generative-ai-can-be-weaponized/

  21. CBC: ChatGPT now lets users create fake images of politicians. We stress-tested it. “New updates to ChatGPT have made it easier than ever to create fake images of real politicians, according to testing done by CBC News. Manipulating images of real people without their consent is against OpenAI’s rules, but the company recently allowed more leeway with public figures, with specific […]

    https://rbfirehose.com/2025/04/14/cbc-chatgpt-now-lets-users-create-fake-images-of-politicians-we-stress-tested-it/

  22. Shrivu’s Substack: How to Backdoor Large Language Models. “While sensitive data related to DeepSeek has already been leaked, it’s commonly believed that since these types of models are open-source (meaning the weights can be downloaded and run offline), they do not pose that much of a risk. In this article, I want to explain why relying on ‘untrusted’ models can still be risky, and why […]

    https://rbfirehose.com/2025/02/24/shrivus-substack-how-to-backdoor-large-language-models/

  23. Shrivu’s Substack: How to Backdoor Large Language Models. “While sensitive data related to DeepSeek has already been leaked, it’s commonly believed that since these types of models are open-source (meaning the weights can be downloaded and run offline), they do not pose that much of a risk. In this article, I want to explain why relying on ‘untrusted’ models can still be risky, and why […]

    https://rbfirehose.com/2025/02/24/shrivus-substack-how-to-backdoor-large-language-models/

  24. ZDNet: Yikes: Jailbroken Grok 3 can be made to say and reveal just about anything. “On Tuesday, Adversa AI, a security and AI safety firm that regularly red-teams AI models, released a report detailing its success in getting the Grok 3 Reasoning beta to share information it shouldn’t. Using three methods — linguistic, adversarial, and programming — the team got the model to reveal its […]

    https://rbfirehose.com/2025/02/20/yikes-jailbroken-grok-3-can-be-made-to-say-and-reveal-just-about-anything-zdnet/

  25. ZDNet: Yikes: Jailbroken Grok 3 can be made to say and reveal just about anything. “On Tuesday, Adversa AI, a security and AI safety firm that regularly red-teams AI models, released a report detailing its success in getting the Grok 3 Reasoning beta to share information it shouldn’t. Using three methods — linguistic, adversarial, and programming — the team got the model to reveal its […]

    https://rbfirehose.com/2025/02/20/yikes-jailbroken-grok-3-can-be-made-to-say-and-reveal-just-about-anything-zdnet/

  26. ZDNet: Anthropic offers $20,000 to whoever can jailbreak its new AI safety system. “Can you jailbreak Anthropic’s latest AI safety measure? Researchers want you to try — and are offering up to $20,000 if you succeed. On Monday, the company released a new paper outlining an AI safety system called Constitutional Classifiers. The process is based on Constitutional AI, a system Anthropic used […]

    https://rbfirehose.com/2025/02/15/zdnet-anthropic-offers-20000-to-whoever-can-jailbreak-its-new-ai-safety-system/

  27. Ars Technica: Anthropic dares you to jailbreak its new AI model. “Today, Claude model maker Anthropic has released a new system of Constitutional Classifiers that it says can ‘filter the overwhelming majority’ of those kinds of jailbreaks. And now that the system has held up to over 3,000 hours of bug bounty attacks, Anthropic is inviting the wider public to test out the system to see if it […]

    https://rbfirehose.com/2025/02/04/ars-technica-anthropic-dares-you-to-jailbreak-its-new-ai-model/

  28. Ars Technica: Anthropic dares you to jailbreak its new AI model. “Today, Claude model maker Anthropic has released a new system of Constitutional Classifiers that it says can ‘filter the overwhelming majority’ of those kinds of jailbreaks. And now that the system has held up to over 3,000 hours of bug bounty attacks, Anthropic is inviting the wider public to test out the system to see if it […]

    https://rbfirehose.com/2025/02/04/ars-technica-anthropic-dares-you-to-jailbreak-its-new-ai-model/

  29. The Guardian: ChatGPT search tool vulnerable to manipulation and deception, tests show. “In the tests, ChatGPT was given the URL for a fake website built to look like a product page for a camera. The AI tool was then asked if the camera was a worthwhile purchase. The response for the control page returned a positive but balanced assessment, highlighting some features people might not like. […]

    https://rbfirehose.com/2024/12/26/the-guardian-chatgpt-search-tool-vulnerable-to-manipulation-and-deception-tests-show/

  30. 404 Media: APpaREnTLy THiS iS hoW yoU JaIlBreAk AI. “New research from Anthropic, one of the leading AI companies and the developer of the Claude family of Large Language Models (LLMs), has released research showing that the process for getting LLMs to do what they’re not supposed to is still pretty easy and can be automated. SomETIMeS alL it tAKeS Is typing prOMptS Like thiS. “

    https://rbfirehose.com/2024/12/23/404-media-apparently-this-is-how-you-jailbreak-ai/

  31. 404 Media: APpaREnTLy THiS iS hoW yoU JaIlBreAk AI. “New research from Anthropic, one of the leading AI companies and the developer of the Claude family of Large Language Models (LLMs), has released research showing that the process for getting LLMs to do what they’re not supposed to is still pretty easy and can be automated. SomETIMeS alL it tAKeS Is typing prOMptS Like thiS. “

    https://rbfirehose.com/2024/12/23/404-media-apparently-this-is-how-you-jailbreak-ai/

  32. Breakthrough! Researchers at Anthropic, Oxford, Stanford, and MATS create Best-of-N Jailbreaking, a black-box algo that jailbreaks frontier AI systems across modalities. #AIResearch #JailbreakingAI #FrontierAI #ArtificialIntelligence #MachineLearning #AIInnovation